02 / Prediction Lab

The Honest Scoreboard

UFC betting markets are close to efficient — the closing line wins about two-thirds of the time. So we measure calibration, not bravado, and beating the market is the only result that would mean anything.

Market benchmark
66.9%
Brier 0.2097 · n=7,076
Genuine locked picks
933
565 correct · 60.6%
Lowest in-house Brier
MASTER
Brier 0.2245 · 66.3% acc · n=67
Engines beating the market
0
measured on Brier score
Soon AI Fight Read

Read any matchup.

Name two fighters and the reader returns a calibrated lean, a confidence band, and the one stylistic factor that tips it — in UNFOLD's honest voice. Qualitative only: no invented records, no odds, no bets.

vs
▸ Enter two fighters and hit “Read the matchup.” The live reader leans — then tells you why it might be wrong. Launching soon.

Engine Scoreboard

Select any engine to plot it · sort any column
Calibration curve
1 engine plotted
MARKET (benchmark) n=6,331MASTER n=47perfect calibration
0%0%25%25%50%50%75%75%100%100%confidence statedhow often it happened

Does a stated confidence mean anything?

Each point is a confidence bucket. The dashed diagonal is perfect calibration — above the line means we were too cautious, below means too sure of ourselves. The market sits almost exactly on it. That is the standard our engines are measured against.

Distance from the diagonal is the error. An engine can pick winners often and still be badly calibrated — which is why accuracy alone is a poor way to choose between models. Plot the ones we look worst against — they are on the list too.

UFC prediction engines ranked by Brier score. The Basis column says how each row was scored: locked picks are predictions made before the event, backtest is a bulk sweep of historical fights. Rows with different bases are not comparable. Select a row to plot that engine's calibration curve against the market benchmark. Column headers sort the table.
benchmarkbacktest4735–234166.9%0.20970.60690.01337,076
locked picks61–3166.3%0.22450.82660.101367
locked picks58–3463.0%0.23040.65780.090354
locked picks49–4353.3%0.24000.67920.104054
locked picks22–1166.7%0.24130.73290.196513
locked picks57–3562.0%0.24520.68340.072954
locked picks55–3759.8%0.25480.70570.129654
locked picks54–3858.7%0.26920.73800.141454
locked picks57–3562.0%0.26960.79900.222054
locked picks52–4056.5%0.28000.87710.196854
locked picks53–3957.6%0.29130.85140.242054
locked picks16–2044.4%0.29690.81520.212936

MARKET is not measured on the same population as the engines. Its row is a bulk backtest over 7,076 historical fights with odds; every in-house row is at most 92 predictions locked before the event and audited after it, fewer for engines added later. The BASIS column says which is which. A backtest and a forward record are different claims, and that gap matters far more than the gap between the percentages — so read down the column before you read across the row.

MARKET is the de-vigged closing line, not a model of ours — beating it, not beating a coin flip, is the only result that would mean anything. BOOKWORM has fewer than 25 scored picks, so its Brier is not yet meaningful. Accuracy figures exclude 260 picks from two 2026 events (230 of them scored) that were reconstructed after the fact. They also exclude every pick for the 2026-03-28 Adesanya–Pyfer card: those 117 picks carry a lock timestamp of the following morning, so whatever else they are, they are not predictions. All of them remain in the database and none are published here. Both exclusions are frozen — the integrity gate fails if either set grows.

Two columns, two sample sizes — read them separately. W–L and ACC are the genuine forward record — every scored pick locked before the fight, reconstructions excluded. BRIER, LOG LOSS, ECE and N cover only those picks that carried a stated pre-fight confidence. Several 2026 events were locked without one, so a row can read 61–31 on accuracy and n=67 on calibration. We leave those picks permanently unscored rather than back-fill a confidence after the fight — a number invented today cannot be a prediction made in May.

Accuracy and calibration currently agree — that is not guaranteed. MASTER picks the most winners at 61–31, and the lowest in-house Brier is MASTER at 0.2245 on n=67 — the same engine leads both columns for now. Neither column is the last word. Two of these engines have a large-sample measurement behind them and it does not flatter the leaderboard: over 7,284 position-neutralised historical fights STYLE scores about 51% — a coin flip — while ELO is the strongest formula engine at 56.6%. Those two figures come from that harness, not from the table above, and they are why we bench on the large sample rather than the flattering one. A lead built on n=67 is still noise until it survives a few hundred more picks, so we publish the ranking and the evidence against it side by side.

How the scoreboard is built and graded → Every audited card →

Analysis and entertainment only. Not betting advice.