- Joined
- Jun 27, 2017
- Professional Status
- Certified General Appraiser
- State
- California
You want the answer to why auto adjustments are allowed (by the GSEs):
"Francis Galton’s 1906 “ox-weighing” observation, now regarded as the classic early demonstration of the wisdom of crowds.
At a livestock fair in Plymouth, England, 787 people entered a contest to estimate the dressed weight of an ox. The actual weight was 1,198 pounds.
Galton initially emphasized the median estimate—what he called the “middlemost” estimate—which was 1,207 pounds, only nine pounds too high. A later calculation of the arithmetic mean of all the guesses produced approximately 1,197 pounds, just one pound below the actual weight.
The marbles-, jelly-beans-, or candies-in-a-jar experiment is a common modern classroom replication of Galton’s result, rather than usually being the original experiment.
The underlying explanation is that, when guesses are made independently, some people overestimate and others underestimate. Their errors partly cancel when averaged:

If the errors are not systematically biased in one direction, their average may be close to zero.
This phenomenon works best when the group has:
it seems this is what automatic submission of appraisal adjustment for different features amounts to.
Yes, Bert—that is essentially the same statistical idea, provided the adjustment submissions are sufficiently independent and come from reasonably competent appraisers.
Suppose many appraisers independently estimate the market adjustment for a feature—for example:

where:

With a large number of independent estimates, the random-error component eˉ may become small. That is the same error-cancellation mechanism behind Galton’s ox-weight experiment.
In effect, the system would treat appraisers as a distributed market-sensing network.
average bathroom adjustment
but something more like:

where A is the submitted adjustment and X contains the relevant property and market characteristics.
For example:

This is closer to a conditional crowd estimate than a raw crowd average.
For example, appraisers may all be influenced by:
If everyone uses the traditional “$10,000 per bathroom” rule, 5,000 submissions do not represent 5,000 independent estimates. They may represent one assumption copied 5,000 times.
For example, a pool adjustment might be submitted primarily in assignments where pools are especially important. The average submitted adjustment could therefore overstate the pool’s effect in the broader market.
If the subject has a feature and the comparable does not, the adjustment may have one sign; if the comparable has the feature and the subject does not, it has the opposite sign.
The system would need to convert all submissions into a common interpretation, such as:
A $25,000 view adjustment has a very different meaning in:
Dollar adjustment
and

Even percentage adjustments are not necessarily stable, but they can improve comparability.
A better aggregation system might report:
One could weight submissions according to demonstrated reliability:

Weights might reflect:
The best weighting criterion would be empirical forecasting or validation performance, not professional status.
For many appraisal adjustments, there is no directly observed final answer. The market contribution of a fireplace or view is latent. It must itself be inferred from transactions.
Therefore, averaging appraisal adjustments can produce a stable consensus, but consensus is not automatically the same as market truth.
A crowd can agree closely and still be wrong.
That distinction is central:

The submissions need to be validated against market evidence.
In Bayesian form:

That would be more defensible than either pure machine estimation or pure professional consensus.
"Francis Galton’s 1906 “ox-weighing” observation, now regarded as the classic early demonstration of the wisdom of crowds.
At a livestock fair in Plymouth, England, 787 people entered a contest to estimate the dressed weight of an ox. The actual weight was 1,198 pounds.
Galton initially emphasized the median estimate—what he called the “middlemost” estimate—which was 1,207 pounds, only nine pounds too high. A later calculation of the arithmetic mean of all the guesses produced approximately 1,197 pounds, just one pound below the actual weight.
The marbles-, jelly-beans-, or candies-in-a-jar experiment is a common modern classroom replication of Galton’s result, rather than usually being the original experiment.
The underlying explanation is that, when guesses are made independently, some people overestimate and others underestimate. Their errors partly cancel when averaged:

If the errors are not systematically biased in one direction, their average may be close to zero.
This phenomenon works best when the group has:
- reasonably independent judgments,
- diversity of information or perspectives,
- no strong common bias,
- and a sensible method of aggregation.
it seems this is what automatic submission of appraisal adjustment for different features amounts to.
Yes, Bert—that is essentially the same statistical idea, provided the adjustment submissions are sufficiently independent and come from reasonably competent appraisers.
Suppose many appraisers independently estimate the market adjustment for a feature—for example:
- a second bathroom,
- an additional garage space,
- superior condition,
- a view,
- a larger lot,
- or a swimming pool.

where:
- θ is the underlying market-supported adjustment,
- bi is the appraiser’s systematic bias,
- ei is random estimation error.

With a large number of independent estimates, the random-error component eˉ may become small. That is the same error-cancellation mechanism behind Galton’s ox-weight experiment.
Why it could work well for appraisal adjustments
Individual appraisers may possess fragments of useful information that are difficult to capture in one formal model:- experience with buyer reactions,
- paired-sale observations,
- knowledge of local neighborhoods,
- prior appraisal assignments,
- agent interviews,
- sensitivity to feature interactions,
- and judgment about whether a feature is typical or unusual.
In effect, the system would treat appraisers as a distributed market-sensing network.
But a simple average would often be inadequate
The principal problem is that appraisal adjustments are not generally universal constants. A bathroom adjustment might differ substantially according to:- price range,
- neighborhood,
- property size,
- existing bathroom count,
- house age,
- buyer population,
- date of sale,
- and whether the market is supply-constrained.
average bathroom adjustment
but something more like:

where A is the submitted adjustment and X contains the relevant property and market characteristics.
For example:

This is closer to a conditional crowd estimate than a raw crowd average.
Independence is the critical issue
The wisdom-of-crowds effect weakens if appraisers are not independently estimating the adjustment.For example, appraisers may all be influenced by:
- the same lender guidance,
- the same appraisal software defaults,
- the same published adjustment tables,
- the same prior reports,
- the same MLS conventions,
- or common industry rules of thumb.
If everyone uses the traditional “$10,000 per bathroom” rule, 5,000 submissions do not represent 5,000 independent estimates. They may represent one assumption copied 5,000 times.
Selection bias would also matter
Submitted adjustments would not be a random sample of market judgments. They would arise only when:- an appraiser chose a particular comparable,
- the comparable differed on that feature,
- the appraiser believed an adjustment was necessary,
- and the reporting system required or permitted the adjustment.
For example, a pool adjustment might be submitted primarily in assignments where pools are especially important. The average submitted adjustment could therefore overstate the pool’s effect in the broader market.
The adjustment direction must be standardized
Another complication is sign convention.If the subject has a feature and the comparable does not, the adjustment may have one sign; if the comparable has the feature and the subject does not, it has the opposite sign.
The system would need to convert all submissions into a common interpretation, such as:
Otherwise, positive and negative adjustments might cancel mechanically even when the appraisers agree perfectly.Estimated marginal contribution of the feature to the comparable property’s market value.
Percentage adjustments may aggregate better than dollar adjustments
Raw dollar adjustments are often not comparable across price levels.A $25,000 view adjustment has a very different meaning in:
- a $400,000 market,
- a $1 million market,
- and a $5 million market.
Dollar adjustment
and

Even percentage adjustments are not necessarily stable, but they can improve comparability.
Medians and trimmed means may be better than ordinary averages
Appraisal submissions could contain:- typographical errors,
- unusual luxury properties,
- inexperienced judgments,
- software-generated defaults,
- or extreme but legitimate market conditions.
A better aggregation system might report:
- median,
- trimmed mean,
- interquartile range,
- number of observations,
- standard deviation,
- geographic coverage,
- time period,
- and confidence or reliability score.
One could weight submissions according to demonstrated reliability:

Weights might reflect:
- sample support supplied by the appraiser,
- consistency with later sale evidence,
- geographic experience,
- appraisal review outcomes,
- or historical predictive accuracy.
The best weighting criterion would be empirical forecasting or validation performance, not professional status.
There is an important difference from Galton’s experiment
In Galton’s experiment, there was a clearly observable final truth: the ox had an actual measured weight.For many appraisal adjustments, there is no directly observed final answer. The market contribution of a fireplace or view is latent. It must itself be inferred from transactions.
Therefore, averaging appraisal adjustments can produce a stable consensus, but consensus is not automatically the same as market truth.
A crowd can agree closely and still be wrong.
That distinction is central:

The submissions need to be validated against market evidence.
The strongest version would combine crowd estimates with transaction models
A powerful system could combine:- Appraiser-submitted adjustments
- Hedonic or nonparametric estimates from sales
- Repeated-sales or matched-pair evidence
- Broker or buyer survey evidence
- Out-of-sample validation
In Bayesian form:

That would be more defensible than either pure machine estimation or pure professional consensus.