Why Two Models Disagree, and What to Do About It

You check one projection source and it likes the over. You check a second and it likes the under, on the same player, the same stat, the same night. The instinct is to decide which one is broken. Usually neither is, and the disagreement is telling you something more useful than either number was.

By SlatelinePublished
Two projection curves for the same player drawn side by side with their centers offset

You open one projection source and it likes the over on a strikeout line. You open a second, built by people who clearly know what they are doing, and it likes the under on the same pitcher, the same night, the same number. Both sources show their work. Neither looks careless. The natural next move is to decide which one is broken, pick it, and move on.

That move throws away the most informative thing on the screen. When two competent systems land in different places, the gap is almost never a bug. It is the visible residue of a modeling choice one of them made and the other did not, and if you can name that choice you learn more about the prop than either projection told you on its own.

Disagreement is the normal state, not a defect

A projection is not a measurement. It is the output of a process that had to make dozens of decisions no data set answers for it: how much history counts, how hard to pull small samples toward a league average, which contextual adjustments are real enough to apply, what to do when a lineup has not been posted. Every one of those decisions is defensible in more than one direction. Two teams making defensible choices independently will produce different numbers, and the surprise would be if they did not.

The size of the gap matters more than its existence. Two projections separated by a tenth of a strikeout are agreeing, for practical purposes, and the difference is noise from simulation counts and rounding. Two projections separated by a full point on a total that sits near five are making genuinely different claims about the world, and one of those claims is going to age badly. The interesting cases live at the second end.

The five places the gap actually comes from

Most disagreements between serious systems reduce to a short list. Working through it in order is faster than arguing about which model is smarter, because the first item explains the majority of large gaps by itself.

  1. Opportunity assumptions. How many plate appearances, minutes, snaps, possessions, or rounds does each model think the player gets? This is usually the single largest driver of a big gap, because volume multiplies straight through into every counting stat.
  2. Sample handling. How aggressively does each system shrink a small sample toward a baseline? A hot twelve game stretch can be taken nearly at face value by one model and pulled almost all the way back to a career rate by another.
  3. Context adjustments. Opponent quality, venue, weather, surface, pace, and role each get applied by some systems and skipped by others, and even when both apply them the magnitudes differ.
  4. Data vintage. One model ran before the lineup was posted or the availability report landed; the other ran after. Nothing about the methods differs at all, and yet the outputs do.
  5. Output type. A point estimate and a full distribution are different objects. Two systems can project the same average and still disagree about the probability of clearing a line, because the probability depends on the spread, not the center.

That last item catches people out constantly. If one source publishes an expected value of 6.2 strikeouts and another publishes 6.2, they look identical, but the one that simulates a wide right tail and the one that assumes a tight distribution will report meaningfully different probabilities of clearing 6.5. As how projections work lays out, the conversion from a projection to a probability happens through the shape of the distribution, and two models can share a center while disagreeing about everything that matters at the line.

Example: Same average, different answer

Suppose a made up starting pitcher draws a 5.5 strikeout line. Model A projects 6.0 strikeouts with a fairly narrow distribution and reports the over at about 63 percent. Model B also projects 6.0, but its simulation produces more short outings, so its distribution is lopsided toward low counts with a long thin upper tail, and it reports the over at about 55 percent. The headline projections are identical. The disagreement is entirely about how often the pitcher gets pulled early, and that is a question about opportunity, not about strikeout skill.

Diagnose the input rather than crowning a winner

The useful reframing is to stop asking which model is right and start asking which input they disagree about. That question has an answer you can often check, while the first one does not.

Start with opportunity, because it dominates. If one system has a player at twenty eight minutes and another has him at thirty three, you do not need to adjudicate their statistical machinery; you need to find out what the rotation actually looks like tonight. If one has a hitter batting second and the other has him sixth, the plate appearance difference alone can account for the entire gap. Opportunity questions are frequently resolvable with a lineup sheet, an availability report, or a look at how the coach has actually deployed the player recently.

If opportunity matches, move to vintage. A stale model is not a competing opinion, it is an old one, and the fix is to discard it rather than average it. Any system that ran before a meaningful piece of news is answering a question nobody is asking anymore. Check the timestamps before you check anything else about the methods.

If both systems have current information and the same opportunity view, the remaining candidates are shrinkage and context, and now the disagreement genuinely is philosophical. One model believes a recent change in a player's rate is signal; the other believes it is noise. Neither of you can settle that tonight. What you have learned is that the prop's honest uncertainty is wider than either source's confidence interval implies, which is a real and actionable finding.

When averaging helps and when it quietly hurts

Splitting the difference between two projections has a respectable pedigree. Combining several independent estimators is the core idea behind ensemble methods, and it works for a specific reason: if two models make errors that are not correlated, the errors partially cancel, and the average is more accurate than either input on average. That is a real effect and it is why so many forecasting fields default to consensus.

The conditions attached to that result are what people skip. Averaging helps when the inputs are genuinely independent and roughly comparable in quality. It stops helping, and starts actively hurting, in three common situations.

  • One input is stale. Averaging a current model with an outdated one does not diversify anything. It just drags a good estimate partway back toward yesterday's information.
  • One input is systematically biased in a known direction. Blending an unbiased estimate with a tilted one imports part of the tilt, and the result inherits a bias the better model did not have.
  • The inputs are not independent. Two sources that consume the same third party projections and apply thin adjustments are one source wearing two hats, and the false comfort of agreement is worse than a visible disagreement would have been.

There is also a subtler problem with averaging probabilities near a line. If two models disagree because they hold different views of a discrete event, such as whether a player starts at all, the average probability can describe a world neither model believes in. A fifty fifty blend of a start scenario and a bench scenario is not a projection of a real night; it is a projection of an outcome that will never occur. In those cases the honest answer is to resolve the discrete question or pass, not to smooth it into a number.

The only tiebreaker that survives an argument

Two people can argue about methodology forever, because every choice has a story behind it. What ends the argument is a graded record: what did each system claim in advance, and what actually happened, across enough decided outcomes to mean something.

This is the point of backtesting, and its value is that it converts opinions about method into measurements of outcome. A model that says 60 percent should hit close to 60 percent of the time across the group of props where it said 60 percent. If it hits 48 percent, no amount of elegant reasoning about shrinkage saves it. If it hits 59 percent over a large decided sample, its shrinkage philosophy has earned the benefit of the doubt on the next disputed prop.

Two honest cautions. A record has to be per market and per sport before it settles a specific disagreement, because a system can be well calibrated on baseball strikeouts and poorly calibrated on rebounds. And a record has to be large enough that the comparison is not itself noise, which for probability calibration usually means hundreds of decided outcomes rather than dozens. A source that publishes a record without the sample size behind it has published a claim, not evidence.

Unexplained disagreement is itself a research output

Here is the rule that saves the most trouble. If two sources you respect disagree sharply and you cannot name the input responsible, that is a pass. Not a coin flip, not a compromise number, a pass.

The reasoning is straightforward. An unexplained gap means there is a material fact about tonight you have not identified, and whichever side you take, you are taking it without knowing what you are betting against. The gap has told you that your information is incomplete. Acting anyway converts a diagnosis into a guess, and the whole point of running a process is to avoid that conversion.

This connects to a broader habit covered in knowing when to pass: the number of props you decline is a feature of a research process, not a symptom of indecision. A slate offers thousands of offers and you are under no obligation to have an opinion about any particular one. Reserving your confidence for the props whose disagreements you can actually explain is what makes the confidence worth anything.

It also cuts the other way, usefully. When two independent systems agree closely and you can articulate why, that agreement is meaningful corroboration rather than a coincidence. Agreement between models built on different assumptions is a stronger signal than agreement between models that share a data pipeline, and knowing which kind you are looking at is part of reading the evidence honestly.

How Slateline treats the comparison

Slateline publishes one projection per offer, so you are usually comparing our number against something external rather than against a second internal model. That makes the labeling questions above more important, not less. Where a market reference exists we show the edge against a fair probability derived from sharp markets and say so; where none exists we say the comparison is model versus line, which is the weaker claim spelled out in model, line, and market.

We also try to make our disagreements diagnosable rather than mysterious. Each signal carries the factors that drove it, along with flags for the conditions that should reduce your confidence: unresolved availability, thin samples, unusual roles, and markets our own graded history has not earned confidence in yet. The record itself, including the markets where we have been wrong, lives in the Model Room, and the current slate with its labeled edge basis is on the signal board.

None of this makes disagreement go away, and it should not. Two careful processes looking at an uncertain night will keep landing in different places, and the researcher who treats that as information rather than as an argument to win is the one who ends up with a calibrated view. Keep any related activity recreational and bounded, and if it stops feeling that way, our responsible gaming page is the right next stop.

References

See the research in practice

Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.

Open the Model Room

Keep researching

Why Two Models Disagree, and What to Do About It · Slateline