How to Read a Calibration Curve

A projection product can publish anything it likes about its own skill. A calibration curve is harder to dress up: it plots what the model claimed against what actually happened, one probability bucket at a time.

By SlatelinePublished
A plotted curve tracking a diagonal reference line across probability buckets

There is exactly one chart that can embarrass a projection system in public, and most products that sell projections do not publish it. It takes every probability the model has ever stated, sorts them into buckets, and asks a question with no room to negotiate: of everything you called 70 percent, how often did it actually happen?

That chart is a calibration curve, sometimes called a reliability diagram. It is not difficult to read once you know what each axis is doing, and learning to read it changes how you evaluate every model you encounter, including the ones you build yourself. This article is about the chart itself: its construction, its failure modes, and the specific things it does and does not prove.

What the chart plots

The horizontal axis is what the model said. The vertical axis is what happened. Both run from zero to one hundred percent. Every point on the chart summarizes a group of past predictions that shared a similar stated probability, and its vertical position is the fraction of that group that came true.

Construction is mechanical. Take every graded prediction. Sort it into a bucket by stated probability, for example everything between 55 and 60 percent. Within each bucket, compute two numbers: the average stated probability, which gives the horizontal coordinate, and the observed frequency of the event occurring, which gives the vertical one. Plot the point. Repeat for every bucket. That is the whole procedure, and its plainness is the point. Calibration) is a comparison between claims and outcomes, computed without any input from the party being evaluated beyond the claims themselves.

The diagonal line running corner to corner is the reference. On it, stated probability equals observed frequency everywhere: the 30 percent bucket happened 30 percent of the time, the 80 percent bucket 80 percent of the time. A perfectly calibrated model traces the diagonal. Nothing real traces it exactly, and a curve that sits precisely on it at every bucket over a small sample is more likely to be a sign of thin data than of excellence.

Buckets, and why volume decides whether a point means anything

The bucket is where most misreadings begin. Each point on the curve is an average of many predictions, and the reliability of that point depends entirely on how many predictions went into it. A bucket holding four hundred graded predictions produces an observed frequency you can lean on. A bucket holding nine produces a number that can only take ten possible values, and moving one outcome from a miss to a hit swings it by eleven points.

So a calibration curve should always be read alongside the count behind each point, and any presentation that hides those counts is asking you to trust the shape more than the shape has earned. Wider buckets give steadier points but blur real structure. Narrower buckets show structure but become noisy at the ends of the range, where predictions are rarer.

  • Read the count first, the position second. A point built on a handful of observations is decoration.
  • Expect the extremes to be sparse. Most sports models make far more predictions near the middle of the range than near 5 percent or 95 percent, so those buckets wobble for reasons that have nothing to do with quality.
  • Beware of curves that pool everything. A model can look calibrated overall while being overconfident in one sport and underconfident in another, the two errors cancelling into a flattering average.
  • Treat a curve computed on the same data used to fit the model as a description, not a test. Honest evaluation grades predictions made before the outcomes were known, which is the entire discipline covered in the backtesting article.
Example: Reading a single bucket

Suppose an imaginary model's 65 to 70 percent bucket contains 240 graded predictions with an average stated probability of 67 percent, and suppose 154 of them came true. The observed frequency is 154 divided by 240, which is about 64 percent. So this point sits slightly below the diagonal: the model claimed 67 and delivered 64, an overstatement of roughly 3 points. With 240 observations behind it, a gap that small is well inside the range you would expect from chance alone, so the honest reading is that this bucket looks fine. Now suppose the same 3 point gap appeared in a bucket holding 18 predictions. That gap would carry no information at all. Same shape on the chart, completely different conclusion, decided by the count. Every number here is invented.

What the shapes mean

Once you can trust the points, the geometry becomes readable, and the common patterns each carry a specific diagnosis.

  • The curve sits below the diagonal on the high side. The model claims more than it delivers: things called 80 percent happen 72 percent of the time. This is overconfidence, and it is the most common defect in sports models, usually caused by treating correlated inputs as independent or by fitting recent form too aggressively.
  • The curve sits above the diagonal on the high side. The model understates what it knows: things called 70 percent happen 78 percent of the time. Underconfidence is rarer and often deliberate, a consequence of heavy shrinkage toward a baseline in a system built to avoid embarrassing itself.
  • An S shape crossing the diagonal, flatter than it in the middle and pushing past it at the ends. The model is compressing its opinions toward the base rate, saying 55 when it should say 62 and 45 when it should say 38. It is directionally right but is not committing, which shows up as too little separation between its confident and unconfident calls.
  • An inverted S, steeper than the diagonal. The model is exaggerating in both directions, pushing probabilities toward zero and one harder than the evidence supports. This is the shape that produces spectacular winning stretches followed by spectacular losing ones.
  • A curve that tracks the diagonal in the middle and wanders wildly at both ends. Almost always thin tail buckets rather than a real defect. Check counts before diagnosing anything.

Calibration is not the same as being useful

Here is the property that surprises people. A model can be perfectly calibrated and completely worthless. Imagine a system that answers 50 percent to every question in a market where events happen half the time. Its curve is a single point sitting exactly on the diagonal. It is flawlessly honest and it tells you nothing, because it never distinguishes one situation from another.

The missing property is discrimination, sometimes called sharpness: the willingness and ability to push probabilities away from the base rate when the evidence justifies it, and to be right about which direction. Discrimination is what makes a model informative. Calibration is what makes its numbers mean what they say. You need both, and they trade against each other in practice, because the easiest way to improve calibration is to stop committing and the easiest way to improve apparent sharpness is to overstate confidence.

This is why the two most common summaries of model quality both fail alone. A hit rate measures neither property cleanly, since it depends on which offers were selected; that argument is the subject of hit rate versus projection quality. And a calibration curve alone cannot tell you whether the model is saying anything worth hearing. Look at the curve to see whether the numbers are honest, then look at the spread of the predictions themselves to see whether the model ever commits.

Brier score, and the limits of one number

When you want a single figure that respects both properties, the standard choice is the Brier score: the average squared difference between each stated probability and the outcome, coded as one or zero. Lower is better. It rewards being confident and right, punishes being confident and wrong more than being cautious and wrong, and cannot be gamed by simply hedging everything toward the middle.

Its limitation is that it is not comparable across contexts. A Brier score on a market where events happen half the time is not comparable with one on a market where they happen a tenth of the time, because the achievable range differs. So Brier scores are for comparing versions of the same system on the same events, and calibration curves are for understanding what a system is doing wrong. Neither replaces the other, and any product that shows one without the other is showing you half the picture. The interpretation of stated probabilities generally, including how much error one honestly carries, is covered in the probability reading article.

Why this is the most honest thing a projection product can publish

Marketing claims about a model cannot be checked. Screenshots of winning slips cannot be checked, and are selected by definition. A hit rate can be true and still uninformative, because the person publishing it chose which offers to count. A calibration curve is different in kind, because it is computed from the model's own stated probabilities against results the model does not control, across every prediction rather than a chosen subset.

It also fails loudly. A model that quietly degrades after a rules change or a data source shift will show it in the curve before it shows it anywhere else, usually as growing overconfidence in the buckets where it was previously reliable. That makes the curve an operational instrument, not just a public one. Slateline publishes a calibration curve per sport in the Model Room, alongside the graded record and the counts behind each bucket, and where a sport has too few decided predictions to say anything, we show the emptiness rather than borrowing another sport's numbers to fill it.

Read ours the way you should read anyone's. Check the counts, check whether the buckets are per sport or pooled, check whether the predictions were graded at the lines actually posted, and treat any curve built on a few weeks of results as a preliminary sketch. Honest calibration is a slow accumulation, and no chart converts a probabilistic process into a predictable one. Set limits before a slate, keep the activity recreational, and if it stops being recreational, help is available at 1 800 GAMBLER and through the National Council on Problem Gambling.

References

See the research in practice

Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.

Open the Model Room

Keep researching

How to Read a Calibration Curve · Slateline