The Honest Scoreboard
UFC betting markets are close to efficient — the closing line wins about two-thirds of the time. So we measure calibration, not bravado, and beating the market is the only result that would mean anything.
Read any matchup.
Name two fighters and the reader returns a calibrated lean, a confidence band, and the one stylistic factor that tips it — in UNFOLD's honest voice. Qualitative only: no invented records, no odds, no bets.
Engine Scoreboard
Select any engine to plot it · sort any columnDoes a stated confidence mean anything?
Each point is a confidence bucket. The dashed diagonal is perfect calibration — above the line means we were too cautious, below means too sure of ourselves. The market sits almost exactly on it. That is the standard our engines are measured against.
Distance from the diagonal is the error. An engine can pick winners often and still be badly calibrated — which is why accuracy alone is a poor way to choose between models. Plot the ones we look worst against — they are on the list too.
| benchmark | backtest | 4725–2331 | 67.0% | 0.2103 | 0.6081 | 0.0156 | 6,331 |
| locked picks | 52–28 | 65.0% | 0.2233 | 0.6420 | 0.0701 | 34 | |
| locked picks | 53–27 | 66.3% | 0.2341 | 0.9324 | 0.1052 | 47 | |
| locked picks | 22–11 | 66.7% | 0.2413 | 0.7329 | 0.1965 | 13 | |
| locked picks | 51–29 | 63.7% | 0.2441 | 0.6813 | 0.0685 | 34 | |
| locked picks | 48–32 | 60.0% | 0.2455 | 0.6836 | 0.1573 | 34 | |
| locked picks | 43–37 | 53.8% | 0.2540 | 0.7249 | 0.1369 | 34 | |
| locked picks | 48–32 | 60.0% | 0.2581 | 0.7125 | 0.1674 | 34 | |
| locked picks | 50–30 | 62.5% | 0.2638 | 0.7597 | 0.2321 | 34 | |
| locked picks | 9–15 | 37.5% | 0.2745 | 0.7452 | 0.2804 | 16 | |
| locked picks | 44–36 | 55.0% | 0.2803 | 0.8830 | 0.1742 | 34 | |
| locked picks | 47–33 | 58.8% | 0.3008 | 0.8620 | 0.2791 | 34 |
MARKET is not measured on the same population as the engines. Its row is a bulk backtest over 6,331 historical fights with odds; every in-house row is at most 80 predictions locked before the event and audited after it, fewer for engines added later. The BASIS column says which is which. A backtest and a forward record are different claims, and that gap matters far more than the gap between the percentages — so read down the column before you read across the row.
MARKET is the de-vigged closing line, not a model of ours — beating it, not beating a coin flip, is the only result that would mean anything. BOOKWORM, GLICKO2_SHADOW have fewer than 25 scored picks, so their Brier is not yet meaningful. Accuracy figures exclude 260 picks from two 2026 events (230 of them scored) that were reconstructed after the fact. They also exclude every pick for the 2026-03-28 Adesanya–Pyfer card: those 117 picks carry a lock timestamp of the following morning, so whatever else they are, they are not predictions. All of them remain in the database and none are published here. Both exclusions are frozen — the integrity gate fails if either set grows.
Two columns, two sample sizes — read them separately. W–L and ACC are the genuine forward record — every scored pick locked before the fight, reconstructions excluded. BRIER, LOG LOSS, ECE and N cover only those picks that carried a stated pre-fight confidence. Several 2026 events were locked without one, so a row can read 53–27 on accuracy and n=47 on calibration. We leave those picks permanently unscored rather than back-fill a confidence after the fight — a number invented today cannot be a prediction made in May.
Accuracy and calibration disagree about who is best, and that is the point. MASTER picks the most winners at 53–27, and the lowest in-house Brier is ML at 0.2233 on n=34 — a different engine, on a different question. Neither column is the last word. Two of these engines have a large-sample measurement behind them and it does not flatter the leaderboard: over 7,284 position-neutralised historical fights STYLE scores about 51% — a coin flip — while ELO is the strongest formula engine at 56.6%. Those two figures come from that harness, not from the table above, and they are why we bench on the large sample rather than the flattering one. A lead built on n=34 is still noise until it survives a few hundred more picks, so we publish the ranking and the evidence against it side by side.
How the scoreboard is built and graded → Every audited card →