How to Keep a Research Log That Is Actually Useful

Most logs record what happened and nothing about why you decided it. That version teaches almost nothing, because outcomes in a high variance activity are mostly noise. The useful version captures the decision while you are still uncertain about it.

By SlatelinePublished
A ledger grid where each row carries a stated probability alongside a settled outcome

Ask ten people who research player props whether they keep records and most will say yes. Ask to see the records and you usually get one of two things: nothing at all, or a column of wins and losses with a running total at the bottom. The second version feels responsible. It is almost useless, and understanding why is the whole subject of this article.

A results only log records the one part of the process you did not control. Whether a shot fell, whether a manager made a substitution in the 61st minute, whether a fight ended in round one. Over any sample a person actually accumulates, that column is dominated by variance. It can tell you how a stretch went. It cannot tell you whether your reasoning was any good, which is the only thing a log could plausibly improve.

Log the decision, not the outcome

The shift is small to describe and hard to sustain: write down what you believed and why, at the moment you believed it, before anything settles. The outcome then becomes one field among several rather than the entire record. Everything useful you can later extract comes from the fields you wrote while you were still genuinely uncertain.

There is a minimum set. Fewer fields than this and the log cannot answer questions. More than this and you will abandon it within two weeks, which is worse.

  • Date and sport, so you can slice the record later by season, by month, and by category.
  • The offer exactly as sold: the stat, the line, the platform, and the side you took.
  • Your probability, written before settlement, as a number rather than a feeling.
  • One sentence of reasoning: the single input you believe the number is missing.
  • The closing number, if the offer was still available near the event start.
  • The settled result, including pushes and voids as their own outcomes rather than blanks.

Notice how much of that describes the offer rather than your opinion. Stat definitions, platform, and line all matter because the same nominal prop settles differently in different places, and a log that collapses them makes its own history unreadable a month later. If you have worked through the research checklist, these are the same four verification fields, simply written down instead of merely checked.

Why the probability field is the point of the whole exercise

One field carries most of the value: the probability you wrote down in advance. Without it, a log can only count wins. With it, a log can measure something far more interesting, which is whether your confidence means anything.

The mechanism is the same one used to grade a model. Group every decision you rated near 55 percent and ask what share of them actually happened. Do the same for the group near 65 percent, and the group near 75 percent. If the 55 percent group lands near 55 and the 75 percent group lands near 75, your confidence is tracking reality. If everything you rated 75 percent hits closer to 58, you are systematically overconfident, and that is a finding you can act on immediately: shade your estimates toward the middle, and treat your large stated edges with suspicion. The full method, including what the resulting curve looks like and how to read its shape, is in reading a calibration curve.

This is why a probability written after the fact is worth nothing. It is also why a log without pre committed probabilities can only ever measure luck. A record of results with no stated expectations has no baseline to compare against, so a good month and a lucky month look identical inside it. The number you write before settlement is what converts a diary into an instrument.

Memory is a hostile witness

The strongest argument for writing things down in advance is that you cannot reconstruct them afterward, no matter how vivid the night felt. Once an outcome is known, recollection of what you expected quietly reshapes itself around what happened. Psychologists call this hindsight bias, and the uncomfortable part is that it operates on careful people and careless people alike. You do not experience it as distortion. You experience it as remembering clearly.

In practice it produces two symmetric errors. Decisions that won get remembered as more confident than they were, so the reasoning behind them gets credit it never earned. Decisions that lost get remembered as doubts you supposedly had all along, so a sound process gets blamed for a bad roll. Both errors push in the same direction: they make your process look more informative than it is, and they make the record unusable for learning.

Example: The same night, two records

Suppose a made up hitter is listed at 1.5 total bases and you write down 54 percent to the over, with the note that the projected opposing starter is a fly ball pitcher in a small park. He hits two home runs. The results only log records a win and a plus sign. The decision log records that a 54 percent call landed, which is exactly what a 54 percent call is supposed to do about half the time, and that the reasoning was about park and pitcher type rather than about power outbursts. A month later, only the second record can tell you whether that reasoning is worth repeating.

How to review the log without lying to yourself

Reviewing is where most logs go wrong even when the recording is good. The natural instinct is to reopen the worst night and relitigate each decision in it. That review is guaranteed to mislead, because you are examining a sample selected on its outcome. Every decision in it looks worse than it was for the simple reason that it lost.

Review in aggregate instead, on a fixed schedule, and never in response to a result. Monthly is about right for most people. Look at calibration first, then at categories: by sport, by stat type, by whether the lineup was confirmed at decision time, by whether the offer carried a modified payout. Categories are where a log earns back the effort, because the finding you want is rarely global. It is usually narrow and specific, something like a market you consistently overrate or a sport where your unconfirmed lineup decisions perform noticeably worse than your confirmed ones.

  • Calibration first: does the stated probability track the realized frequency across buckets?
  • Then category: which sports, stat families, or offer types carry the misses?
  • Then process: how often were you deciding without confirmed opportunity information?
  • Then closing numbers, where you have them, as an independent read on your timing.
  • Never a single night, and never a review triggered by how the last one went.

The closing number deserves a note of its own. Where an offer was still on the board near the event start, recording what it closed at gives you a second, outcome independent measure: whether the market moved toward your side or away from it. That is far from proof of anything, and it is noisy on thin markets, but it accumulates faster than settled results do. The general logic of judging a process against a reference rather than against outcomes is the same logic in backtesting player prop models, and it is why our own engines are graded in public in the Model Room rather than by a highlight reel.

The log is also the honest ledger of what this costs

There is a second function that gets less attention and matters more. A log makes real frequency and real spending visible, and both are things memory systematically understates. People remember slates they researched carefully and forget the ones they clicked through. They remember the size of a good week and round the rest away.

Because your log already carries a date on every row, it answers questions no recollection can. How many days in the last month did this activity appear at all. Whether entries cluster after losing sessions, which is the clearest early signal of chasing. Whether the amounts drifted upward without any decision to raise them. Whether the total for the month matches what you told yourself you were spending. Those four questions cost nothing to ask of a log you were already keeping.

The version you will actually maintain

Ambitious logs die. A spreadsheet with twenty two columns, tagged matchup notes, and a dashboard survives about nine days, and then the gap in the record makes the whole thing feel ruined and you stop. The log that works is the one whose cost per row is under thirty seconds.

So start with six columns and a plain spreadsheet: date, offer, side, your probability, one sentence, result. Add the closing number when it is easy and leave it blank when it is not; a partly filled column still carries information. Add a category column only after you have two months of rows and a question you cannot answer without it. Resist every urge to add fields you have not yet needed, because the binding constraint here is never analytical sophistication. It is whether you fill in the row on a Tuesday when you are tired.

One more habit is worth the effort: log the passes. When you research an offer and decide against it, write the row anyway with your probability and your reason. Passes are where most of your good judgment lives, and a log that only contains actions cannot see it. It also makes the record honest about the denominator, which matters when you later ask how selective you actually are.

None of this predicts anything. A log is a mirror, not a model. What it gives you, after a few honest months, is the ability to say something most people in this activity cannot say at all: whether your confidence has historically meant what it claimed. If you want a graded example of that standard applied to a model rather than a person, the projections and their settled record sit in the Model Room, and the current slate is on the signal board.

References

See the research in practice

Slateline grades every projection it publishes and shows its record in the open. Browse the Model Room to see hit rates, calibration, and methodology for every sport we cover.

Open the Model Room

Keep researching

How to Keep a Research Log That Is Actually Useful · Slateline