• Welcome to AppraisersForum.com, the premier online  community for the discussion of real estate appraisal. Register a free account to be able to post and unlock additional forums and features.

Appraisal Statistics: Regression - Do Your Really Know How Much Work It Is?

Status
Not open for further replies.
How many data points per item do you have, typically, for your regression analysis?

That's a good question. My MLS data downloaded from MLSListings has over 400 fields. But, these fields cover different property types, such as income producing apartments. They have a fairly large number of highly covariant vales, such as "Total Baths", "Half-Baths", "Quarter-Baths", "Full Baths" and so on. But they also have fields such as Bathroom and Kitchen Description, which list numerous features that should be used to generate numerous new Boolean fields for analysis. After generating any temporary fields and then eliminating highly covariant fields, you might experiment around feeding 50+ fields into your regression to discover any that might be important contributors to value (assuming you are using something as powerful as MARS). You might also just start out with the fields in the URAR grid and maybe one or two others you know might be important based on you experience with the subject neighborhood. So, your regression could be using input of 5 to 50+ values. But, typically, my final regression might be on 15-20 variables, and of these most usually only 5-10 are of any significant importance. But occasionally, "interactions" may become important - but they have to be added onto the grid as a separate attribute. For example, a large house on a relatively small lot, affording a small front and back yard could very well be a negative factor in price and so you could make an adjustment for GLA/LotSize Ratio.
 
Bert says, "5. Then you calculate the difference between the model prediction and the actual sale price."

That doesn't sound like your finding comparables to estimate the value of the subject. That sounds like your finding comparables to fit a model which will support sales price.
 
I like to have between 30 and 40 values if I'm going to do a regression. Interestingly, occasionally the outliers are the comps that I have already chosen and are in my grid!
But I'm mostly using a regression to back up my land and GLA adjustments (and remove outlier data). It just depends on the situation, the more homogeneous the neighborhood, the better use a regression will be.
I do enjoy reading everyone's different methods on this forum....
 
I like to have between 30 and 40 values if I'm going to do a regression. Interestingly, occasionally the outliers are the comps that I have already chosen and are in my grid!
Excellent observation and, eventually, you will also understand the reasons why leading to much better appraisals! You are on to something important, I think.
 
Bert says, "5. Then you calculate the difference between the model prediction and the actual sale price."

That doesn't sound like your finding comparables to estimate the value of the subject. That sounds like your finding comparables to fit a model which will support sales price.
The comp sale price not the subject. He's describing testing the regression model to see how closely it replicates sale prices used for input.

To clarify:
Sale 1, 2, 3, 4 & 5 were input into the regression model with sale price as the dependent variable. The model spits out predicted prices based on the features analyzed. Then compare actual sale price of sale 1 to predicted sale price...how close is it? Repeat for 2, 3, 4 & 5 and you get an idea of how closely your regression model emulates actual prices.
 
Last edited:
<SNIP>...Typically I would run regression on all the neighborhood sales within a certain GLA sq. ft. range of the subject to extract patterns and the impact of various features on price... <SNIP>In any case, I think you can see, it can be a lot of work to add in the variable values that may be missing from your MLS data and then tweak the model by seeing if you can improve your supplied (most likely subjective) values. Also, you should be able to understand why AVMs are so far off - as they don't know the values of all the important contributors to value that are typically not in the MLS or at least cannot be estimated very well from the data in the MLS. [ But interestingly, if the AVM companies can get some of that data, they can use it to incrementally improve their models - thus the interest in getting their hands on appraiser data.]
It's not my intent to belittle or demean anyone here, but for those of us objecting to Bert's method: as far as I can tell from a superficial reading of this and Bert's other comments as well as my experience with many (too many) MLS multi-variable regressions and grad-level econometrics courses, it is close to - if not - textbook. The appraisal INDUSTRY has to realize that what Bert calls "appraisal statistics" and what I consider "real estate valuation econometrics" is very, very deep and not within the grasp of even the most experienced appraisers that haven't had the training at graduate level (or equivalent). He hasn't even addressed the confidence testing math he alludes to because it would be way over the vast majority of our heads. Some of you know this, but others of you "do not know what you do not know." Even those who routinely run regressions without understanding the math behind them - as well as all the pitfalls (autocorrelation, heteroskedasticity, multicolinearity, etc.) - are making the big mistake thinking that it's just a matter of running the software.

I think it's great that Bert's taking the initiative to bring this topic to the forefront, but unless you have some capacity to understand the concepts and math involved take this opportunity to learn some of the concepts. I urge you to ask questions and let him explain - he may eventually refer you to some of the links to articles explaining the formulas and concepts necessary to obtain "good" results. Hopefully, many of you will go on to learn more about the issues so that you can successfully argue AVM in the AVM language, not from a perspective of ignorance and misunderstanding (sorry, we are all ignorant about some things).

Bert, you are going to face some difficulty obtaining a satisfying and productive discussion of this here. I hope you find enough educated minds, or at least bring some of the concepts to appraisers faced with AVM-obsolescence! I don't do residential anymore and have no profit in the subject, but it would be great if appraisers could present their objection to AVMs from an educated AND experienced perspective (not just educated OR experienced).
 
...Interestingly, occasionally the outliers are the comps that I have already chosen and are in my grid!...
Overnight after reading through this thread I realized that this issue is most likely related to one of the pitfalls of multi-linear regression: autocorrelation, heteroskedasticity, multicolinearity. If you aren't already familiar with these concepts, try to find something online that explains them in English (i.e. no math) and see if one or more aren't happening in your model. For example, trying to fit a model where one or more parameters lie in the outer fringes of a regression where heteroskedasticity results in a break-down in confidence, can skew the regression so that an "outlier" that is actually a very good comparable from a practical perspective falls outside.

My statistics-application experience is waaay out of date. Does your software detect and/or mitigate autocorrelation, heteroskedasticity and/or multicolinearity. Are there additional (more modern) issues you have to watch for?

Here's the first hit I found explaining these:
http://www.academia.edu/6782732/Pro...sticity_Autocorrelation_and_Multicollinearity
 
What Bert is saying in that snipped quote is basically saying what I have been saying about a lot of the data for important contributors to value not even being collected.

It's not even a matter of getting their hands on appraiser data. Most of the differences that are not collected are considered when you select the comparables. So it is not like they can collect that info from appraiser data.
 
What Bert is saying in that snipped quote is basically saying what I have been saying about a lot of the data for important contributors to value not even being collected.

It's not even a matter of getting their hands on appraiser data. Most of the differences that are not collected are considered when you select the comparables. So it is not like they can collect that info from appraiser data.

Agreed, because that data comes from the appraiser's experience and knowledge of the area, and what's project to change in the near future and where.

Not all properties in a neighborhood are equally impacted by changes, or expected changes.

.
 
Agreed, because that data comes from the appraiser's experience and knowledge of the area, and what's project to change in the near future and where.

Not all properties in a neighborhood are equally impacted by changes, or expected changes.

.

So many differences.
 
Status
Not open for further replies.
Find a Real Estate Appraiser - Enter Zip Code

Copyright © 2000-, AppraisersForum.com, All Rights Reserved
AppraisersForum.com is proudly hosted by the folks at
AppraiserSites.com
Back
Top