All insights

Put the forecast on trial

A convincing simulation earns trust from the outcomes it did not train on.

The forecast is six months old. It is on the screen in a meeting room, unchanged since the day it was issued. Across the table sits a list of what actually happened to the properties. The people and meeting are hypothetical. The uncomfortable question is real: how do you tell whether a simulation deserved to be believed before you knew the endings?

One property closed later than expected. Another sold quickly for less cash than the model suggested. A third is still listed. It would be easy to circle those cases in red and pronounce the forecast wrong. It would be just as easy to point to a few close predictions and pronounce it accurate. Both judgments can miss what a probabilistic forecast is for. It was never promising that every house would follow its middle path.

A prediction is not an alibi

Imagine that a model assigns a 70 percent chance of sale by day ninety to a group of comparable listings. The number is illustrative, not an Oftu result. One of those listings can remain unsold without disproving the forecast. Across a sufficiently large and relevant group, however, outcomes should bear a defensible relationship to the probabilities assigned. If listings given high chances routinely miss the deadline, the model has a problem that a few lucky closings cannot excuse. Even a well-calibrated model can be useless if it gives every property the same market-average chance. It must also distinguish properties that are more likely to produce timely cash from those that are less likely.

Timing is only the first test. A model that predicts sale by day ninety might still overstate what the seller takes home after debt, fees and carrying costs. Another may estimate proceeds well among successful sales while quietly dropping the listings that failed to close. A useful review must test the whole question: how much usable cash, by when, and with what chance of shortfall? Otherwise, accuracy becomes a label that can be earned by answering the easiest part.

Even that combined test can hide a fault. Suppose the model performs well overall because it is used mostly for conventional homes in active markets. It may fail on unusual properties, thin markets or places with changing insurance conditions. An average can let a weak segment borrow credibility from a strong one. The meeting should ask who was wronged by the errors, not only how the grand total looked.

The test data must also be honest about time. A model trained on sales through June should be judged on cases it had not seen in July and beyond. If information from a later closing leaks into the earlier forecast, the trial has been rigged without anyone meaning to cheat. The Federal Reserve's current model-risk guidance describes out-of-sample and out-of-time testing, comparison of outputs with real outcomes, and continuing monitoring. That guidance addresses banking organizations. It is not a certificate for Oftu. Its discipline is useful to anyone asking people to act on a model.

THE FORECAST ON TRIAL

Did the future land inside the range?

WHEN

Did sales close by the dates the model assigned?

HOW MUCH

Did usable cash fall within its predicted ranges?

WHO WAS MISSED

Which property types and markets repeatedly broke the forecast?

These are tests to conduct, not results Oftu claims to have achieved.

The cases that make the room go quiet

The most revealing part of the meeting may be the cases no one initially wants to discuss. A house listed twice under different identifiers can appear to have sold quickly the second time, although the owner spent months waiting. An offer that fell through may vanish from a dataset of completed transactions. A property sold just after the deadline may count as a successful sale in one report and a shortfall in the owner's life. Before testing the model, the team has to decide what an outcome is and make sure its records can see it.

Then come the surprises. A cluster of late closings in one district might point to a local bottleneck. A pattern of overstated net proceeds could reveal a missing cost. If low-probability paths appear too often under a new credit regime, the relationship between financing and buyer completion may have changed. The objective is not to invent a story for every miss. It is to check whether errors gather around the same assumptions and to change the model when evidence warrants it.

Some tests take time. A six-month cash forecast cannot be fully evaluated after two weeks. But early signals can still be watched: inquiry rates, offer failures, withdrawal patterns and closing delays. The same Federal Reserve guidance emphasizes monitoring as data and conditions change. Those signals are not substitutes for eventual outcomes. They are warnings that a model may be leaving the market it was built to describe.

There should be a second person in the room willing to ask the rude question. What if the model's apparent improvement comes from changing the test set? What if a wider predicted range covers more outcomes only because it has become too vague to guide a choice? What if the model was accurate for sellers who could wait but harmful for owners with fixed deadlines? A forecast can be technically well calibrated for one purpose and misused for another. Its intended decision needs to be stated before its score is celebrated.

Revising the model is not shameful. Hiding the revision is. A new rule for insurance costs, a better way to link relisted homes, or a different sale-time model may improve later forecasts. Keep the old version beside the new one for a while. Compare both on the same unseen cases. Record where each wins and loses. If the improvement disappears outside the development sample, the original certainty was premature.

This approach also gives a clearer answer to marketing claims. A percentage called “accuracy” is meaningless without the target, time horizon, population and error measure. An exact sale-price hit rate is a different claim from a well-calibrated chance of receiving cash by a date. The latter may be harder to explain in a slogan. It is the one a deadline-bound owner can use.

The meeting ends without a verdict that fits on a badge. There are strengths, blind spots and questions the current records cannot settle. The forecast on the screen remains the one issued six months ago. Its value now is not that it can be made to look inevitable after the fact. It is that the team can learn, case by case, whether the range it showed was honest enough to shape a decision before the ending was known.

Source notes