How Player Prop Projections Work

Ask a projection what a player will do tonight and the honest answer is not a number. It is a shape: a range of outcomes with probabilities attached. This article walks through how that shape gets built, layer by layer, and why the shape matters more than the average sitting in the middle of it.

By SlatelinePublished
A bell shaped outcome distribution with a posted line marked as a vertical rule through it

Ask a projection system what a player will do tonight and the honest answer is not a number. It is a shape. A hitter projected for 1.1 total bases does not produce 1.1 total bases in any actual game; he produces zero in a lot of them, two in a fair number, and five or six in a few. The single number is a summary of that shape, and most of what makes a projection useful, or useless, happens in the parts of the shape the summary throws away.

This article walks through how a serious projection gets built, in the order the construction actually happens. The order is the point. Every layer depends on the one below it, and the layer almost everyone skips is the foundation.

Layer one: opportunity

Before any question about skill, a projection has to answer a question about volume. How many chances will this player get? In baseball the unit is the plate appearance. In basketball it is minutes, and beneath minutes, possessions. In soccer it is whether the player starts and how deep into the match they stay on the pitch. In a fight it is how long the fight lasts. Different sports, same principle: no rate means anything until you know what it gets multiplied by.

Opportunity is also where most of the volatility in props lives. A batter hitting sixth instead of second loses roughly half a plate appearance per game on average, which quietly moves every counting stat he can produce. A basketball rotation change moves scoring projections before anyone discusses shooting. This is why a projection pipeline resolves lineups, roles, probable starters, and availability before it touches a single rate. Get opportunity wrong and every downstream number is polished nonsense.

Layer two: rates, and why small samples get pulled toward typical

With opportunity established, the next layer is what the player tends to do with each unit of it. Per plate appearance outcome rates. Points per possession. Shots per ninety minutes. These come from the player's own history, and here the pipeline meets its oldest enemy: samples that are too small to trust at face value.

The standard remedy is called shrinkage, and the plain language version is simple. When you have very little evidence about a player, your best estimate is mostly what a typical player looks like. As real evidence accumulates, the estimate moves away from typical and toward what the evidence says. A rate built this way is a blend: some weight on the player's observed numbers, some weight on a league baseline, with the balance shifting toward the player as the sample grows.

Example: Shrinkage in one made up case

Suppose a fictional hitter has 3 home runs in his first 40 plate appearances of a season. Taken literally, that pace is superstar territory. A shrunk estimate does not take it literally. It starts from what an ordinary hitter does, moves a modest distance toward the hot start, and lands on a projection only slightly above average. Forty plate appearances is a whisper of evidence, and the estimate treats it like one. After 400 plate appearances at the same pace, the estimate would have moved much further, because now the evidence is loud.

Shrinkage is why good projections look boring in April and why they refuse to chase last week. It is not caution as a personality trait. It is the mathematically correct amount of trust to place in a given amount of evidence.

Layer three: context

A rate is an average across contexts, and tonight is one specific context. The third layer bends the baseline rates toward the actual matchup. The opponent matters: a strikeout rate faces a specific pitcher with his own tendencies, a scorer faces a specific defense. The venue matters: parks change how batted balls carry, altitude changes endurance, surfaces change tennis entirely. Role matters: the same player used differently produces different stat lines from identical talent.

The discipline here is that adjustments should be measured, not vibes with decimals. An adjustment worth applying is one you can estimate from data and one whose size you can defend. Stacking many small subjective bumps in the same direction is one of the most common ways an otherwise sound projection becomes an overconfident one.

Layer four: simulation, where the distribution comes from

Now the pipeline has opportunity, adjusted rates, and context. The final step is to put them together, and here there is a fork in the road. The simple path multiplies things out into a single expected value and stops. The serious path plays the game.

Simulation means running the game forward, event by event, thousands of times, letting randomness fall differently each run. In one simulated game the leadoff hitter walks twice and the lineup turns over five times. In another the starter exits in the third inning and the bullpen changes every matchup behind him. Each run produces a full stat line for every player, and across thousands of runs those stat lines pile up into a distribution: how often the player finishes with zero, one, two, three of the stat in question. This general approach, estimating a hard distribution by sampling random outcomes many times, is the Monte Carlo method, and it earns its keep wherever the interactions are too tangled for clean formulas.

The reason to pay the computational cost is that a distribution is the only object that can actually price a line. A posted line at 1.5 is not asking what the player averages. It is asking how much probability sits at 2 or more. A single point estimate cannot answer that question; two players with the same average can have wildly different amounts of probability above the same line, because one is steady and the other is streaky. The distribution carries that difference. The average erases it.

Simulation also gets the weird interactions right without special pleading. A pitcher who gets knocked out early takes his strikeout total with him. A blowout changes who plays the fourth quarter. A fight that ends in the first round settles every volume stat in it. When you simulate the game rather than the stat, these dependencies happen naturally inside each run, and the distribution inherits them for free.

Why the median and the mean disagree, and why 0.5 lines care

Counting stats are lopsided. A player cannot hit fewer than zero home runs, but he can hit three. Most games cluster at the low end and a thin tail stretches out to the right. On shapes like this the mean and the median split apart: the mean gets dragged upward by the tail, while the median, the outcome in the middle of the pile, stays low.

A fictional slugger might carry a mean of 0.35 home runs per game while his median game contains exactly zero. Both numbers are true. They answer different questions. The mean answers what he produces per game across a long season. The median tells you that in most individual games he does not homer at all.

At a 0.5 line this distinction is the entire question. Over 0.5 home runs settles on one thing only: the probability of at least one, which for most hitters on most nights is well under half even when the mean looks respectable. Research that reasons from the average toward a 0.5 line will systematically overrate the over on rare event stats, because the average is quietly borrowing credit from the tail. The distribution does not make this mistake. It states the probability at zero directly, and the line gets priced against that.

The layer after the last layer: grading

Everything above describes how a projection gets made. None of it tells you whether the projections are any good. That question can only be answered afterward, by comparing stated probabilities to what actually happened, across volume, at the lines that were actually posted.

This is how Slateline operates. Each sport on our board runs its own engine, simulating at the level of the events that sport is actually made of: plate appearances in baseball, possessions in basketball, points in tennis, fight time in MMA. And every projection those engines publish is graded after the fact, with the record and its calibration shown in the Model Room. When the honest answer is that a market has not earned trust yet, the grade says so.

If you build your own estimates, hold them to the same standard. Write the probability down before the game, grade it after, and let the accumulated record tell you whether your layers are sound. The construction order in this article, opportunity, rates, context, simulation, is a method anyone can follow. The grading is what makes it research instead of storytelling. The broader process around it is covered in the research guide, and you can see the full graded output of our own pipeline any day on the signal board.

References

See the research in practice

Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.

Open the Model Room

Keep researching

foundations

Player Prop Research: A Complete Guide to Better Decisions

Most people research player props backwards. They start from a name they like and look for evidence. This guide starts from the line and works outward: what the number claims, what a projection can add, and how to know whether your process is actually any good.

markets

Model, Line, and Market: Three Numbers That Mean Different Things

Open any prop research screen and you may see three numbers describing the same event. One came from a model. One is a product a platform wants to sell you. One is a distillation of what the sharpest markets believe. Confusing them is the most common structural error in prop research.

methodology

Why Sample Size Matters in Player Prop Analysis

Flip a fair coin ten times and seven heads is unremarkable. Watch a player for ten games and seven overs feels like destiny. The mathematics is identical; only the storytelling changes. This article covers how much evidence a sample really carries, why some stats settle down quickly while others take a season, and what to do when the data you have is all the data there is.

methodology

How Backtesting Builds Trust in Player Prop Models

Any model can look brilliant in a chart made after the games ended. The entire question is what the model said before. This article defines an honest backtest, catalogs the specific ways published records deceive, explains calibration and the Brier score in plain language, and argues that admitting a model is unproven is a feature, not a confession.

How Player Prop Projections Work · Slateline