A record of picks tells you what won. Calibration tells you whether the probabilities meant anything: when we say 60%, does it happen 60% of the time? Below, every settled prediction, scored two ways — the sharp market (devigged) and our own model — with the uncomfortable comparison in plain sight.
Our model (Elo + form): 0.611
Our model (Elo + form): 1.020
Our model (Elo + form): 1.8pp
Each point is a 10-point probability band (min. 30 observations). The closer to the diagonal, the better calibrated. Every match contributes its three outcomes.
| Band | Sharp market (devigged) — Predicted (avg) | Observed | Obs. | Our model (Elo + form) — Predicted (avg) | Observed | Obs. |
|---|---|---|---|---|---|---|
| 0–10% | 7.4% | 6.3% | 79 | — | — | — |
| 10–20% | 16.2% | 16.9% | 508 | 16.6% | 13.5% | 303 |
| 20–30% | 25.2% | 23.8% | 1528 | 26% | 24.9% | 1857 |
| 30–40% | 34.3% | 33.5% | 862 | 34.6% | 33.8% | 786 |
| 40–50% | 45% | 46.9% | 480 | 44.8% | 48.6% | 642 |
| 50–60% | 54.8% | 57.8% | 325 | 54.2% | 57.3% | 309 |
| 60–70% | 64.5% | 64.8% | 182 | 64% | 62.4% | 101 |
| 70–80% | 74.4% | 80% | 70 | — | — | — |
Lower Brier and log-loss are better; calibration error is the observation-weighted average gap between predicted and observed. The honest headline: the devigged sharp market is better calibrated than our model — which is precisely why this site prices everything against the market instead of selling you model predictions. Methodology: probabilities as displayed at logging time, no retro-fitting; outcomes from official results.
Full methodology →