CRE Deal Vision — scorecard methodology
How the public scorecard works
Companion to the public scorecard. Written in plain language on purpose: if a rule cannot be explained simply, it is probably hiding something.
1. Predictions are frozen before reality arrives
When an underwrite is finalized on the platform, its inputs, its predicted outputs (year-one NOI, IRR, DSCR, exit value, and so on) and the provenance of every data feed it used are written into an immutable, append-only ledger — a snapshot pinned to the finalized underwriting run. That snapshot is the prediction. It cannot be edited, replaced, or quietly re-run later. Whatever we said before the outcome was known is what gets graded.
2. What counts as an outcome
An outcome is a realized reading recorded against a frozen snapshot: what the deal actually did, under the same metric names the snapshot predicted. Two kinds matter:
- Realized — the deal ran its course and real numbers exist. Each metric is compared against its prediction.
- Written off — the deal failed. A written-off deal with a prediction and no realized number is scored as a miss, full stop. Dropping failed deals from a track record is survivorship bias — the losers vanish and the scorecard flatters itself. Ours keeps them.
Deals that are still running are excluded from the rates and reported as such — they are open calls, not evidence in either direction.
3. What counts as a hit
A hit is a prediction that landed inside its tolerance band. The bands are set in the scoring engine, not chosen per deal:
- Per-metric scorecard (the page you came from): realized within ±10% of predicted counts as a hit (10% relative tolerance, boundary inclusive). If the prediction was exactly zero, only an exact-zero outcome counts — a relative band around zero has zero width, so we never fake one.
- Backtest IRR calibration: actual IRR within ±2 percentage points of predicted IRR. The model is considered production-ready only when at least 70% of realized deals hit that band AND the average signed IRR error stays under 1.5 percentage points (no systematic optimism).
Metrics present on only one side (predicted but never measured, or measured but never predicted) are counted and disclosed as unscored — never silently dropped, never counted as hits.
4. Why every rate carries a Wilson 95% confidence interval
A hit rate from a handful of outcomes is mostly luck, in either direction. So every rate we publish carries a Wilson 95% confidence interval — a standard statistical range meaning: given this many observations, the true long-run rate plausibly lies anywhere inside it. Small samples produce wide intervals, and the interval is the honest claim. “7 hits out of 10” sounds like 70%, but the Wilson interval says the truth could plausibly be anywhere from roughly 40% to 89% — so that is what we tell you. We use the Wilson interval specifically because it stays honest at small sample sizes, where the naive formula overstates precision. The public page uses the exact same interval math as our internal publication gate.
5. The two thresholds: 10 and 30
- 10 realized DEALS — the founder decision that makes the scorecard public at all. Below it the page shows the current count and nothing else: nothing is published early, nothing is curated. We count deals, not gradings. A single deal can be graded several times as reality arrives, and re-grading the same deal does not move us closer to publishing — otherwise one deal graded quarterly for ten quarters would trip this bar on its own.
- 30 realized deals — the conventional statistical floor for citing a proportion. Between 10 and 29 the numbers are live but flagged prominently as a small sample, and the confidence interval does the talking. Separately, the backtest engine refuses to mark its calibration claim “publishable” until the sample reaches 30 and the lower bound of the Wilson interval itself clears 70% — a small lucky streak can never be promoted into a claim. That gate is now evaluated on the public page too, and when it fails the page says so in terms rather than quietly presenting the rate as a claim.
- Confidence intervals are clustered by deal. One deal contributes several metric comparisons, and can be graded more than once. Treating each comparison as an independent trial would make the interval look narrower than the evidence supports, so we take the conservative bound: each deal contributes at most one independent observation. The headline interval is the clustered one; the unclustered figure is shown beside it for reference and is always the narrower of the two.
- We state how many organizations the numbers come from.An aggregate over one operator’s self-recorded outcomes is not a cross-customer statistic, and the page will not describe it as one.
6. What we don't do
- No retroactive edits. Snapshots live in an append-only ledger. A prediction, once frozen, is never revised to match reality.
- No cherry-picking.Every recorded outcome is in the aggregate — hits, misses, and write-offs. There is no “representative selection.”
- No survivorship laundering. Written-off deals are misses, not omissions.
- No per-deal disclosure. The public page shows platform-wide aggregates only — counts, rates, intervals, and error buckets. No deal names, no client names, no per-deal rows, ever. The public data path is physically unable to fetch those fields.
- No point estimates without intervals. If the interval is wide, you see a wide interval.