Umpires and Strike Zones as a Strikeout Prop Factor
Umpire notes are the most confidently repeated small factor in baseball prop research. The effect is real, it is measurable, and it sits several rungs below the things that actually decide a strikeout line.

Search any pitcher's name on a slate morning and somebody will tell you who is behind the plate. The note usually arrives with more certainty than any other line of analysis in the thread, which is strange, because it describes the smallest of the inputs being discussed. Understanding why it is both real and small is a useful exercise in ranking evidence.
The claim under the note is defensible. The strike zone is defined in the rulebook as a region over the plate bounded by the batter's stance, but the called zone is the one a human applies in real time, and humans vary. Those variations propagate: a zone called slightly larger produces more called strikes, which produces more two strike counts, which produces more strikeouts and fewer walks. The mechanism is sound. The question is how much of a strikeout line it actually moves, and the answer is: less than the confidence in the note implies.
Where this factor sits in the order of operations
Strikeout props are opportunity problems first. A pitcher who records fifteen outs cannot strike out as many batters as one who records twenty one, and the difference between those two workloads swamps every rate adjustment you could make. Expected outs recorded, driven by the manager's tendencies, the bullpen situation, pitch count limits, and whether the start is likely to go badly early, is the dominant term. The strikeout research article works through that hierarchy in full.
After workload comes the pitcher's own strikeout rate, which is the single largest rate input and the one with the most reliable evidence behind it. After that comes the opposing lineup: contact profiles vary enormously across teams, and a lineup that rarely swings through pitches is a much larger obstacle than any umpire. Handedness splits and the specific batters likely to face him a third time matter here too.
Only after all of that does the umpire enter, alongside other environmental modifiers such as venue effects covered in the park factors article. This ordering is not a stylistic preference. It reflects the relative size of the terms, and putting a small term first is how research produces confident conclusions about nothing.
- Expected outs recorded, which sets the denominator for everything else.
- The pitcher's own strikeout rate, estimated over a meaningful sample.
- The opposing lineup's contact and swing tendencies, including handedness.
- Venue and conditions.
- The home plate umpire's called strike tendency, regressed heavily toward league average.
Estimating a tendency is a sample size problem
The reason umpire adjustments must be small is not that the effect is imaginary. It is that the effect is hard to estimate precisely, and an estimate you cannot trust deserves a small weight regardless of how large the underlying effect might be.
Consider what goes into an umpire tendency number. It is built from called pitches, which are only the pitches nobody swung at, across the games that umpire happened to work behind the plate. Those games came with particular pitchers, particular catchers, particular batters, and particular weather. Catcher receiving skill alone can shift called strike totals meaningfully, so an umpire who drew a run of games with skilled receivers will look different from one who did not, without either having a different zone.
That is confounding, and the honest response to a confounded estimate over a limited sample is heavy regression toward league average. The general principle is the one in the sample size article: the less certain the measurement, the closer your working number should sit to the population mean. Umpire tendencies fail this test more than most inputs, because the sample per umpire per season is limited by scheduling and the confounds are not small.
The effect is on rates, and a start is a small sample
Here is the dilution nobody accounts for. An umpire tendency is a rate effect on called pitches. To reach a strikeout prop, it has to survive three narrowings.
First, only some pitches are called. Swings, fouls, and balls in play never touch the umpire's judgment, so a tendency applies to a fraction of the pitches thrown. Second, only some called pitches are near enough to the edge for a tendency to change the outcome; a pitch down the middle and a pitch at the eyes are called the same by everyone. Third, a change in called strikes only sometimes changes a plate appearance's result, because many counts survive an extra strike or an extra ball without changing what eventually happens.
Then the whole thing gets applied across a single start, which is a small number of plate appearances. A rate nudge across roughly two dozen batters produces a fractional expected change in strikeouts, and fractional expected changes are invisible against the variance of a real start. Strikeout totals for the same pitcher against the same lineup swing by several in either direction for reasons no model captures.
Suppose a fictional umpire is estimated to expand the called zone enough to add half a percentage point to strikeout rate for the pitchers he works with, and suppose that estimate has already been regressed toward league average. Suppose a made up starter is projected to face 24 batters. Half a percentage point across 24 plate appearances is 0.12 expected strikeouts. Against a line of 6.5, that shifts the probability of clearing by a small amount, likely a point or two. Meanwhile suppose the same starter's expected batters faced could plausibly be 21 or 27 depending on how the game goes, a range worth well over a full strikeout. The workload uncertainty is roughly an order of magnitude larger than the umpire adjustment. Every number in this example is invented to illustrate proportions.
This is the honest shape of the factor. It is not zero, and a model that ignores it entirely is leaving a small amount of accuracy on the table. It is also not a reason to take a side, and it can never rescue a line you otherwise dislike.
Assignment timing limits what you can do with it anyway
There is a practical constraint on top of the statistical one. Crew assignments rotate through positions, and which member of a crew works behind the plate on a given day is generally not public information far in advance. In practice the identity of the home plate umpire becomes reliably known shortly before the game.
That timing matters for research workflow. A factor you learn about close to first pitch arrives when lines have already been posted, and by then whatever effect it has may already be in the number. Line movement in the hours before a game reflects information arriving, and there is no reason to assume this particular piece of information is the one the market ignored. Trading on a widely published, easily accessed factor at the moment everyone else receives it is not an edge, whatever the size of the underlying effect.
It also means umpire information cannot form the basis of a research process. You cannot build a slate plan around inputs you receive minutes before the market closes. The realistic use is as a final small adjustment on props you already researched for other reasons, and as an occasional reason to skip a marginal spot rather than to take one.
The real danger: stacking small adjustments in one direction
The failure mode this factor invites is not overrating the umpire. It is what happens when the umpire note joins four other small notes that all happen to point the same way, and the researcher adds them up.
The pattern is familiar. The pitcher has a good strikeout rate. The lineup strikes out often. The park plays favorably. The catcher receives well. The umpire has a large zone. Each of those is defensible in isolation and each is small, so the arithmetic feels conservative. But the notes were not selected independently: they were gathered after forming a view, which means the ones supporting the view got found and the ones opposing it got less attention. Five small adjustments in the same direction produce a large adjustment, and the confidence attached to it does not reflect the fact that the selection of adjustments was itself biased.
Some of these factors are also not independent of each other in the first place. A catcher's receiving and an umpire's called strike tendency are measured from the same pitches, so a favorable reading on both may be one effect counted twice. Double counting correlated inputs inflates confidence in exactly the same way as selective attention, and it is harder to notice.
The honest conclusion
Home plate umpires differ, the differences reach strikeout and walk rates, and a heavily regressed adjustment for them is a defensible part of a complete model. That is the entire claim, and it is smaller than the volume of discussion around it suggests.
Practically: research the workload, the pitcher, and the lineup first, and if those three do not produce a view, no umpire note is going to. When the assignment becomes known, apply a small adjustment if you apply one at all, and let it break ties rather than form conclusions. Be most suspicious of it when it agrees with everything else you found, for the reasons above. And keep the distinction clear between a projection, the line a platform sells, and a market reference derived from sharp prices, because a factor this small will rarely be the difference between them.
Slateline's MLB engine models called zone effects as one environmental input among many, regressed toward league average and reported in the per signal explanation rather than presented as a headline reason. Anything that small should be visible in the reasoning and invisible in the conclusion. Set limits before the slate, and if this stops being recreational, help is available at 1 800 GAMBLER and through the National Council on Problem Gambling. To see which factors actually carry weight on a given strikeout projection, the per signal explanations on the Signal Board show the ranking directly.
References
- Strike zone (Wikipedia)
- Umpire (baseball) (Wikipedia)
- National Problem Gambling Helpline (National Council on Problem Gambling)
See the research in practice
Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.
Open the Model RoomKeep researching
How to Research MLB Pitcher Strikeout Props
Strikeout rate is one of the most stable skills in sports, which makes strikeout props a rare market where process genuinely compounds. The catch is that the stable rate sits behind an unstable quantity: how long the pitcher stays in the game.
Park Factors in MLB Props: What They Move and What They Do Not
Every ballpark has a reputation, and most of those reputations are directionally true and quantitatively overstated. A park factor is a measured average across seasons and weather, applied to one game it never saw.
Why Sample Size Matters in Player Prop Analysis
Flip a fair coin ten times and seven heads is unremarkable. Watch a player for ten games and seven overs feels like destiny. The mathematics is identical; only the storytelling changes. This article covers how much evidence a sample really carries, why some stats settle down quickly while others take a season, and what to do when the data you have is all the data there is.