← The Honest Read
Data-quality detective

How a data vendor turned a homebuilder into the "best stock in the S&P 500" — and how we caught it

Pillar: Data-quality detective · Educational, not financial advice

If you build a stock model, the most dangerous bug isn't in your math. It's in the data your math trusts.

Here's a real one we caught, because it's a clean example of why a model needs to be skeptical of its own inputs — and why "the number looks amazing" should sometimes make you more suspicious, not less.

The setup

Divergia values companies on point-in-time fundamentals: it asks "given what was actually knowable on this date, what is this business worth?" At one historical rebalance (Jan 2025), one name came out on top across every single configuration of the model. Not close — the top pick, everywhere, by a wide margin.

The company was a homebuilder.

A homebuilder topping a quality-and-cash-flow screen is not impossible, but it should raise an eyebrow. Homebuilding is cyclical, capital-intensive, and rarely screens as a free-cash-flow machine. So we looked at why the model loved it.

The tell

The model loved it because its data provider reported the company generating operating cash flow at roughly 4 to 6 times its real level — at a scale comparable to the entire cost of the homes it sold. Pushed through the model, that produced a free-cash-flow figure so large it implied the company was a once-in-a-generation compounder.

There was one detail that gave it away instantly: the reported operating cash flow was larger than the company's total assets.

That's not a great business. That's an impossible one. A company cannot, in a normal year, generate more cash from operations than the sum of everything it owns. When you see it, you are not looking at a great company. You are looking at a corrupted data row.

The fix isn't "blacklist the stock"

The naive fix is to manually delete the bad name and move on. We didn't, because manual blacklists don't generalize — the next corrupted row will be a different ticker.

Instead the model now applies a structural sanity check that is independent of the company, the sector, and the price: if normalized free cash flow exceeds a sane multiple of positive net income, the figure is incoherent and the name is quarantined. Net income and cash flow can diverge, but not by the margin a data error produces. The rule doesn't need to know anything about homebuilders. It just knows that this shape of number can't be real.

When we ran the corrected model over history, the homebuilder disappeared from the top of the list everywhere it had falsely appeared — and the same check caught a second corrupted row, on a different company, in a different year, that wasn't even on our radar. Honest performance over the full window actually improved slightly, because the model was no longer betting on a number that didn't exist.

Why we're telling you this

Most stock tools will never show you this. They ingest a data feed, run their formula, and hand you a ranking. If the feed is wrong, the ranking is wrong, and you'd never know.

The reason we built Divergia to be suspicious of its own inputs — and the reason we'd rather tell you "this number can't be real, we're setting it aside" than serve you a confident, beautiful, wrong answer — is that in investing, the expensive mistakes don't come from the names a model rejects. They come from the ones it loves for the wrong reason.

A model that can't catch a homebuilder out-earning its own balance sheet has no business telling you what to buy.


Educational/informational content generated by a quantitative model — not financial advice, not a recommendation to buy or sell, and not personalised to your situation. Company specifics are described in general terms; intrinsic values and data flags are model outputs, not facts. Do your own research.

Want to see the honest read on a ticker you care about — including when the model sets a name aside? Run it yourself: https://divergia.ai/