A probabilistic forecast is judged, correctly, by a property that most casual readers, and a fair share of professional bettors, never actually check: whether stated probabilities match how often things go on to happen. This property is called calibration, and it is measurable at scale wherever forecasts are issued repeatedly, football fixtures, election contests, prediction-market contracts. A large-scale audit of more than a decade of published forecasts across sports and politics found that calibration holds up remarkably well in aggregate, and that the interesting story lives in the exceptions.
This post works through the underlying mathematics: what calibration formally requires, why it is a necessary but not sufficient condition for a useful forecast, and where research has found real, structural deviations, specifically among heavy favourites in one-sided matches, and among the correlated, multi-leg markets that dominate prediction-market design. Football fixtures provide the running example throughout. The conclusion is not that forecasting is broken. It is that being right on average and being useful are two different claims, and conflating them is the single most common error made by people who bet or trade on outcomes.
§ 1 Defining Calibration Formally
Consider a large set of events, each assigned a stated probability p by some
forecasting process, football match outcomes, election results, prediction-market settlement
conditions. Group events into buckets by their stated probability, rounding to the nearest
convenient interval. Within each bucket, let n_p denote the number of events
forecast at probability p, and k_p the number of those events that
actually occurred.
Calibration is, by construction, a large-sample property. With a small n_p,
the ratio k_p/n_p will wander around p from binomial variance alone,
not from any defect in the model. The relevant yardstick is the standard error of an observed
frequency under the assumption that the forecast is correctly calibrated.
A one-point gap against a standard error of 0.6 points is unremarkable. A six-point gap against a standard error of 1.8 points is not. The same nominal-looking deviation can be noise in one bucket and a real signal in another, purely as a function of sample size. This is the first thing any rigorous audit of a forecasting record has to establish before drawing conclusions from a calibration chart.
§ 2 Empirical Calibration at Scale
One useful benchmark, an audit covering sports and political forecasts issued over more than a decade, is instructive mainly because of its scale. Of 5,589 events assigned roughly a 70 percent chance, the outcome occurred in 71 percent of cases, a one-point gap that Equation 2 shows is well inside a single standard error, indistinguishable from noise. At the other end of the scale, of 55,853 events assigned roughly a 5 percent chance, the outcome occurred about 4 percent of the time. Across the full range of stated probabilities, a plot of predicted probability against observed frequency sits close to the 45-degree line that defines perfect calibration.
A calibration curve does not tell you whether a forecast is skilful, only whether it is honest about its own uncertainty. A curve sitting on the diagonal is a necessary condition for a good forecaster. It is not sufficient, which is the subject of §3.
§ 3 Calibration Is Not the Same as Skill
A forecaster who assigns every one of the 64 teams in a knockout football competition an equal 1-in-64 chance of lifting the trophy is, by definition, perfectly calibrated. Across many such tournaments, a team given a 1.6 percent chance would indeed win about 1.6 percent of the time. But such a forecast has no resolution: it does not separate the genuine contenders from the makeweights, and it is exactly as accurate as a forecast built from the tournament bracket alone, with no football knowledge at all. Calibration without resolution has no value to anyone pricing outcomes.
The standard formal tool for separating these two properties is the decomposition of the Brier score, the mean squared error between stated probabilities and outcomes.
A market-priced forecast that simply reproduces the market's own implied probabilities will score well on reliability and contribute nothing on resolution beyond what the market already knew. An edge, in the sense a bettor or a market maker actually cares about, requires resolution that the market does not already have, delivered without sacrificing reliability. Chasing calibration alone, without ever asking whether a forecast's probabilities differ meaningfully from consensus, is a way of confirming honesty while learning nothing about skill.
§ 4 Where Real Forecasts Drift: Heavy Favourites in One-Sided Matches
Team-sport forecasting research has repeatedly turned up the same soft spot: models tend to run a little hot specifically on heavy favourites in one-sided matches. Probability buckets in the 80 to 90 percent range often resolve a few points lower than stated. The mechanism is not mysterious: a dominant scoreline compounds a string of small, real edges into something that looks inevitable after the fact, but was never as fully locked in beforehand as the final margin suggests.
As an illustrative worked example, applying the same statistical logic to a football setting: suppose a sample of 383 fixtures in which the model-favourite was given an 85 percent chance of victory, drawn from research on this exact pattern in team-sport forecasting. If those favourites actually won only 79 percent of the time, the gap is 6 percentage points against a standard error of 1.8 points (Equation 2), a deviation of more than three standard errors. That is a real, structural signal, not sampling noise, and it is precisely the kind of gap that should prompt a forecaster to revisit how a model handles blowout-prone fixtures.
| Bucket (illustrative) | Predicted | Observed | n | SE | Gap / SE |
|---|---|---|---|---|---|
| Broad forecast average | 70% | 71% | 5,589 | 0.6 pts | 1.6 |
| Heavy favourites, one-sided fixtures | 85% | 79% | 383 | 1.8 pts | 3.3 |
For anyone pricing markets around heavy favourites, the practical reading is straightforward: the shortest prices on the board are exactly where a model, and public sentiment alongside it, are most likely to be overstating certainty.
§ 5 Correlation and the Mispricing of Linked Markets
A second, opposite pattern shows up wherever outcomes are correlated rather than independent. Election forecasting is the clean example: candidates given a 75 percent chance have, in past studies, gone on to win closer to 83 percent of the time, underconfidence rather than overconfidence. The mechanism is correlation between linked contests, many state or regional races moving together, which a model that treats them as independent will systematically misjudge.
When outcomes are positively correlated, the covariance terms are positive, and the true variance of the joint outcome is larger than a model assuming independence would compute. That inflated variance fattens both tails at once: the front-runner sweeping every linked contest becomes more likely than the independent case implies, and so does a clean upset across the board. Prediction markets built on multi-leg or accumulator-style products carry exactly this exposure. A market can look properly priced leg by leg and still be structurally mispriced on the joint outcome, precisely because each leg was evaluated as if the others did not exist.
§ 6 Full Comparison
Across the market structures most relevant to a bettor or a market maker:
| Market Type | Typical Bias | Root Cause | Sample Needed |
|---|---|---|---|
| Broad, single-event forecasts | Close to calibrated | Large, independent sample | Thousands |
| Heavy favourites, one-sided fixtures | Overconfident | Blowouts compound small edges | Hundreds |
| Correlated multi-leg / accumulator markets | Underconfident on joint outcome | Correlation treated as independence | Depends on ρ |
| Rounded probability buckets near 0% or 100% | Apparent bias may be a labelling artefact | Asymmetric truncation at the bounds | — |
| Uniform, zero-information forecasts | Perfectly calibrated | No resolution, no value | Any |
The uniform forecast at the bottom of the table is the reminder worth keeping closest at hand. It scores well on the one column, calibration, that is easiest to measure and hardest to fake badly. It scores at zero on the column, resolution, that actually determines whether a forecast is worth anything to a bettor or a market maker.
§ 7 Practical Safeguards
Two disciplines follow directly from the mathematics above, and both are cheap to apply.
Check the standard error before trusting a gap. A deviation is only meaningful relative to the sample size it was measured on. Equation 2 takes seconds to apply and prevents the common error of treating a small-sample wobble as a structural finding, or dismissing a large-sample deviation as noise because the percentage gap looks small in isolation.
Be wary of rounded bucket labels near the boundaries. A bucket labelled "5 percent" near the edge of the scale can only draw from probabilities between roughly 2.5 and 7.5, and if the true forecasts inside it skew low, which they often do near 0 and 100 since the distribution has nowhere to go but up, the honest average of that bucket can sit below the rounded label even with zero model error. Reading a rounded label as exact ground truth can manufacture the appearance of miscalibration that is not really there.
Individual bettors will rarely accumulate the thousands of same-probability outcomes this kind of audit requires. The practical substitute is closing line value: the market's closing price already aggregates the calibration of every other participant, at a scale no individual record can match. Consistently beating the close does, for one person, much of the statistical work that a large-sample audit does for an entire forecasting operation.
None of this is an argument against building or trusting probabilistic models. It is an argument for asking two separate questions rather than one: is this forecast honest about its own uncertainty, and does it know something the market does not. A forecast can pass the first test completely and fail the second one entirely, and it is the second question, not the first, that determines whether a probability is worth acting on.
References
Silver, Nate. “When We Say 70 Percent, It Really Means 70 Percent.” FiveThirtyEight, April 4, 2019.
Murphy, Allan H. “A New Vector Partition of the Probability Score.” Journal of Applied Meteorology, vol. 12, no. 4, 1973.