All insights

The house with no twins

A thin market should make a forecast more candid, not more confident.

The nearest recent sale is a neat three-bedroom house. The property being studied used to be a school. It has a generous plot, a room that still looks like a classroom, and a heating system that needs its own explanation. This is an invented address, but the analyst's problem is ordinary: a database can return a dozen nearby sales without returning one close comparison.

There is a buyer for this property somewhere. Perhaps several. Their willingness to pay and their ability to close by a particular date are not printed in a completed-sales record. The owner wants to know what usable cash a sale might produce before winter. The analyst opens a worksheet of comparable sales and pauses at the first empty row. An automated valuation offers a number to the nearest dollar. The precision belongs to the display, not necessarily to the evidence.

How far away is a comparable?

The analyst has choices, none harmless. Use the house across the road because it is close. Use a converted building across the county because its layout and buyer appeal are closer. Use a sale from three years ago and adjust for a market that has changed. Or admit that each observation illuminates only part of the question. Distance on a map is one kind of similarity. Condition, legal use, financing, maintenance and the likely buyer pool can matter more.

Fannie Mae's comparable-sales guidance recognizes this difficulty in appraisal work. When genuinely similar sales are scarce, appraisers may need to use older, farther or less similar transactions, provided they explain their choices and differences. This is guidance for US mortgage appraisals, not a recipe for a property-liquidity model. It makes one principle plain: selecting imperfect evidence is a judgment that should be visible, not an invisible search-radius setting.

Suppose the nearby three-bedroom home sold within a month. It gives a signal about local demand and perhaps about the seller's timing. It does not show how many buyers can use the old school building or finance its repairs. The converted property farther away tells a different story about buyer appetite for unusual space, but it lives in another local market. A past sale from the same street may be closer in both respects, yet its contract was made under old borrowing costs. If a model silently blends the three into one smooth answer, the result can look more certain than any of its ingredients.

The local price index is useful, too, within its limits. FHFA research found that more localized house-price indexes can improve fit where appreciation varies among submarkets and there are enough transactions to construct them. That last condition matters here. The thinner the local evidence, the more tempting it is to use a very broad index. A broad index can provide context while missing what makes this address difficult to compare.

There is another pile of evidence on the desk: listings that did not sell. An asking price is not a transaction price. Still, a similar unusual property that sat on the market for months may say something important about how long a seller could wait. Its failure might reflect an ambitious price, a defect or a private decision to withdraw. Without the full listing history, it should neither be ignored nor treated as a direct estimate of this home's fate.

THE EVIDENCE GETS THINNER

Similarity is not a distance on a map.

NEARBY

Close in space, but perhaps wrong in condition, use, or buyer pool.

OLDER

Closer in character, but recorded under a different market.

UNSOLD

Proof of an attempt, even without a closing price.

Each record answers part of the question. None becomes a twin because a model needs one.

The uncertainty with two names

Even with excellent data, a future sale is uncertain. A buyer may appear next week or not until spring. An inspection may go smoothly or uncover a problem. Those are different possible paths through a market. The analyst can draw a range of outcomes because the future itself has not happened.

But there is also uncertainty about the analyst's knowledge. Perhaps the true pool of buyers for former schools is larger than the records suggest. Perhaps the sale histories have missed withdrawals and relistings. Perhaps the adjustment from a distant converted building is poorly grounded. Running the same model a million times will describe its assumed buyer pool more precisely; it will not prove that the assumption was right. A good forecast should separate a wide range caused by ordinary variation from a wide range caused by scant evidence, at least in its explanation to the reader.

That distinction protects the property from an unfair conclusion. A lack of close sales does not mean that a home is unsellable. It means the estimate deserves less confidence. Calling the house itself risky when the real problem is an empty dataset turns an evidence gap into a verdict about an asset. Conversely, an unusual property should not be assigned the typical neighborhood sale time merely because no better records are available.

What might a responsible model do? It could borrow some information from a wider set of properties after stating why they belong together. It could test the answer against several plausible buyer pools and sale-time assumptions. It could widen the range and lower its confidence when those choices change the result. It could say that a particular deadline probability cannot be supported from the available records. That last answer may frustrate a buyer of forecasts. It is much better than a precise percentage made from a weak comparison.

Human local knowledge can help, but it needs the same scrutiny. An agent may know which features buyers praise and which lenders question. A surveyor may see a repair that never appears in sale data. These observations are useful inputs if they are recorded and checked. A confident anecdote from one sale should not become a probability distribution without more work. The model is not more rigorous simply because the anecdote has been turned into a parameter.

The owner of our imaginary schoolhouse may decide to commission an inspection, test buyer interest earlier, or keep a larger cash reserve against delay. Those actions have costs. Their value depends on whether the missing fact could change the decision. If the owner can wait a year, learning the exact chance of a sale by winter may matter less. If funds must arrive by then, the weak evidence is itself important information about how carefully to plan.

At the end, the analyst has not found a twin for the house. That is not a failure to be hidden with another decimal place. The once-empty worksheet now holds an older sale, a distant conversion and an unsold listing. Each has a note about what it can and cannot tell us. Beside the cash range sits a sentence the automated figure could not produce: the chance of a sale before winter is not well supported by these records. The owner can plan around that warning. A crisp, unjustified number would have given her nothing to plan with.

Source notes