Why We Publish Our Misses

The industry standard for a track record is selective memory: screenshots of the winners, a counter that resets after a bad month, units with no definition. A record that omits its losses cannot produce a hit rate, cannot produce calibration, and cannot be checked by anyone. This is the argument for publishing the whole ledger, and how to hold us to it.

By SlatelinePublished
A ledger split into two columns, wins on one side and losses on the other, both fully filled in

Scroll any feed where projections are sold and you will find the same artifact repeated a thousand times: a cropped screenshot of a winning slip, timestamped after the fact, posted by an account whose losses have no address. Nobody is lying, exactly. The winner was real. It simply arrived alongside an unknown number of siblings that were never photographed, and the missing siblings are the entire dataset.

We publish our misses because a record with the misses removed is not a weaker version of a track record. It is a different object entirely, one that cannot answer any question you would want to ask it. This piece explains what a record has to satisfy before it means anything, why the losses carry more information than the wins, and what specifically we put in public so that a skeptical reader can check us rather than believe us.

What a record has to satisfy

There are four properties, and each one closes a specific escape route. Drop any one and the other three stop protecting anything, because whatever flattery the process is capable of will simply relocate to the gap.

  1. Recorded in advance. Every projection exists in writing, with its probability and its side, before the outcome is knowable. A prediction retrieved from memory after a result is not a prediction.
  2. Complete rather than curated. Every projection the system published gets graded, including the losses, the pushes, and the legs voided when a player never appeared. A record you assembled by choosing what to include measures your taste in highlights.
  3. Tied to lines that were actually posted. The result is graded against the number a named platform really sold, at the modifier it really sold it at, under that platform's settlement rules. A record graded against a consensus number nobody could buy is a record of a market that did not exist.
  4. Not quietly revisable. If yesterday's entry can be edited, reclassified as a no play, or retired along with the model version that produced it, then the record is a draft and its history is whatever the author most recently preferred.

None of those four is about being right. They are all about being checkable. A record that satisfies all four is allowed to look mediocre, and frequently will, because as the sample size piece works through, a genuinely sound process still loses in long ugly stretches on the way to its average. A record that never looks mediocre is not evidence of skill. It is evidence that something is being filtered.

Survivorship bias and the records that vanish

The statistical name for the missing siblings is survivorship bias: when you only get to see the cases that made it through some filter, the surviving cases look far better than the population they came from, and no arithmetic performed on the survivors can recover the truth. Applied to prediction records, the filter is usually one of three things.

  • Selection at the account level. Enough independent accounts start posting picks and, by chance alone, some finish a season well ahead. Those accounts grow; the rest go quiet. The visible population was produced by outcomes, not ability.
  • Selection at the market level. A tool covers thirty markets, publishes a record page for the eight that performed, and lets the silence around the other twenty two read as neutrality.
  • Selection in time. The counter resets after a bad run, the archive begins at a convenient date, the losing version gets deprecated and its ledger goes with it. Every individual number stays true while the record as a whole becomes fiction.
Example: How honest numbers produce a dishonest record

Suppose a made up service launches projections across ten sports and simply waits. After a few hundred settlements, chance alone will leave two or three sports comfortably above break even and two or three comfortably below. The service builds a landing page for its best two, prints a combined figure, and every digit on that page is accurate. The lie is not in any number. It is in the selection that chose which numbers got a page. This is why the only defensible unit of publication is the whole ledger, decided before anyone knew which sports would cooperate.

Notice that none of these filters requires bad intent. Selective memory is the default state of any process that publishes voluntarily, because posting a winner feels like evidence and posting a loser feels like an apology. Avoiding it takes a mechanism, not a temperament.

The misses carry the information

Start with the simplest measure. A hit rate is a fraction, and the losses are the denominator. Remove them and there is no fraction left, only a count of wins, which grows monotonically forever and therefore tells you nothing about quality. Any figure derived from a wins only ledger is arithmetic performed on a number that cannot go down.

Now go one level deeper, to the thing a projection model actually claims. A model does not claim it will win. Lines sit close to fair by design, so losing constantly is the normal condition of even a strong process. What a model claims is that its probabilities mean something: that the offers it called 60 percent really do land about 60 percent of the time. Testing that claim is calibration), and calibration is defined over both outcomes. Every bucket needs its failures inside it or the curve is undefined. A record without misses cannot be miscalibrated, which sounds flattering until you realize it also cannot be calibrated. It has simply left the domain where the question can be asked. Reading a calibration curve walks through what one looks like when the losses are present.

This is also why probability accuracy is much harder to fake than a win percentage. A good stretch can hand anyone a flattering hit rate. Matching your stated confidence to reality at every confidence level across hundreds of settlements is exactly as difficult as being right about uncertainty, which is the actual job.

There is a commercial cost here and it is worth naming plainly. A published loss is a screenshot someone else gets to use, and a visible losing week is a reason a subscriber leaves. The tools that hide their losses are not confused about statistics. They are responding rationally to an incentive. We would rather absorb that cost than sell a number that cannot be checked, because a research product whose claims are unfalsifiable is not a research product.

What Slateline publishes

Every sport with a live board carries a public graded record: MLB, WNBA, tennis, MMA, soccer, LoL, and CS2. The record for each sits in the Model Room, alongside the calibration view, the probability accuracy measures, and how model versions have performed against one another. The methodology for each engine is published in the same place, including the limitations we know about and have not solved.

The mechanics matter more than the summary. Projections are written with their probability and their side before the slate runs. Grading happens against the line a named platform actually posted, at the modifier it was sold under, using that platform's own settlement behavior, which includes recording a leg as void rather than as a win or a loss when the player never appeared. Voids are not silently dropped; the treatment of nonappearance is a documented rule rather than a convenience, and the void rules piece covers why that distinction changes what a probability even means.

Where an engine has not earned trust yet, the product says so structurally rather than in a disclaimer. Newer engines carry capped grades: the strongest recommendation tiers are simply unavailable until that specific market's own graded history reaches real volume and quality. Demonstration boards for sports without live lines are labeled as demonstrations and carry no edge claims at all. The point of the cap is that a top grade is a claim about proven edge, and an unproven engine making that claim is writing a check its ledger has not cashed.

How to hold us to it

Transparency that only the publisher can audit is decoration. So here is the useful version: treat the questions below as a standard you apply to every tool you consider, and apply them to us first.

  • Ask whether the record is browsable or only summarized. A single headline figure with no underlying rows is a claim, not evidence.
  • Ask what happened to the markets, sports, and model versions that are no longer mentioned. Absence is the most common form of curation.
  • Ask what line each result was graded against, and whether that line was really for sale on a named platform at that moment.
  • Ask whether probability accuracy is shown, not just a win percentage over a window the publisher selected.
  • Ask whether the tool has any way to say it does not know. A product with no unproven state, no skipped offers, and no losing stretches is describing a world that does not exist.

Then do the same thing for yourself. Log your own research decisions before results exist, keep the ones that went badly, and let your own ledger arbitrate between your instincts and your process. Hit rate versus projection explains why a recent streak is such a poor guide on its own, and the same reasoning applies to your read of your own judgment.

Our record will look unimpressive in places. Some markets will sit below break even for stretches long enough to be uncomfortable, and we will leave them visible, because the alternative is a number nobody can check. If you want to see the ledger rather than take a description of it on faith, it is in the Model Room, and today's projections that will end up inside it are on the signal board. If research ever stops being research for you, the resources on our responsible gaming page are there for a reason.

References

See the research in practice

Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.

Open the Model Room

Keep researching

Why We Publish Our Misses · Slateline