Methods · October Is Coming
The Sports Page
Making the numbers mean something since the first pitch
Vol. I, No. 157September 1, 2026Distributed Free to Friends & Family

We Built 252 Ways to Predict the Playoffs. The Best One Lost to the Simplest One.

Our winner picked 59.2% of postseason series correctly. Shown a season it had never seen, it fell to 48.9% — worse than a coin. A plain run differential, with nothing tuned at all, beat it.

The Sports Page · The Professor · how a good number goes bad

59.2%
Our best metric, graded on its own homework
48.9%
The same metric on a season it had not seen
57.8%
Run differential, nothing tuned at all

The question: can you build a number that tells you who wins in October?

We tried 252 of them. The answer is no.

Here is the surprising part. The failure was not that our metrics were bad. The best one looked excellent. It only fell apart when we showed it a season it had not already studied.

What you get from it: the single question that separates a prediction from a story.

What we were trying to fix

A reader raised a fair objection to measuring hot teams by run differential. A club that outscores people by 42 runs might have earned it, or it might have played the Rockies six times.

There is a standard fix. Add the average quality of everyone a club faced to its own run differential, then solve until all the ratings agree. Then blend that with how the club played over its last N games, and you have a measure of current form that does not reward beating up the helpless.

Two dials: how many recent games to count, and how much weight to give them against the full season. We tried every setting. Windows from 5 games to 60, weights from none to all. That is 252 different metrics, each one tested against every postseason series from 1998 through 2025 — 223 series in all.

First, the part that worked

Before the tuning, one plain result is worth your time. We asked how often each simple measure picks the winner of a playoff series.

Won-lost record: 50.9%. The standings, the number every broadcast opens with, are a coin flip in October.

Run differential does better at 57.8%, and the schedule-adjusted last twenty games manages 57.4%. Real, but modest.

Then we optimised, and it felt wonderful

Out of the 252, the best came in at 59.2%: sixty-five percent weight on the season rating, thirty-five on the last twenty games.

That is better than anything simple. Written up carelessly it becomes a finding, with a name and a chart and a claim that we had beaten the standard measures.

Figure 1 · A winner, and what it was actually worth
252 CANDIDATE METRICS, AND WHAT HAPPENED WHEN WE TESTED THEMEach dot is one way of measuring a club: a blend of season rating and recent form.Position is how often it picked the winner of a postseason series, 1998–2025.a coin flipplain run differential, no tuningthe winner: 59.2%scored on the same data that chose it252 candidates, in-sampleNOW TEST THAT SAME WINNER HONESTLY:48.9% — held out one season at a time56.4% — tuned on past seasons only45%50%55%60%HOW OFTEN IT PICKED THE WINNER OF A PLAYOFF SERIES

Two honest tests. It failed both.

A metric picked for fitting the past has to be tried on something it has not seen. We did that twice, in two different ways.

Hide one season, tune on the other twenty-six, predict the hidden one, repeat: 48.9%. Worse than guessing.

Or be kinder and more realistic. Tune only on seasons that had already happened, then bet the next October, the way anyone actually would: 56.4%.

Plain run differential over those same seasons: 58.3%. Under neither test did our clever number beat the boring one.

When the best window is never the same twice, there is no best window.

The Professor

The tell we should have caught first

Look again at the crowd of dots. Seventeen of the 252 sit within a single point of the winner. The champion was not a champion; it was the tallest blade of grass.

Worse, the settings never held still. Tuned year by year, the best window jumped between 10 games, 20 games and 30 games, and the weight wandered from 0.35 to 0.90. If the ideal window is ten games in one season and thirty in the next, there is no ideal window. There is only noise, and we kept renaming it.

Why it had to happen

With 223 series, the uncertainty on any one of these rates is about six points either way. We were choosing among 252 numbers that all sat inside each other’s error bars, which is another way of saying we never had enough games to tell them aparti.

Take the maximum of 252 noisy numbers and you have mostly measured the noisei. Our winner beat the middle of the pack by four points; the error on any single one of them is larger than that.

And the ceiling was never high. Set the best measure in baseball, 57.8%, against the 75% we use as the point where a difference becomes noticeable, and it is under a third of the way from guessing to that line. To even prove 57.8% is not 50% would take about 332 series. Baseball produces eleven a year.

What to take home: ask what it had not seen

When somebody shows you a number, ask one question. Was it tested on data it had never seen?

If it was chosen for fitting the past, it is a description of the past. That can be worth having; it is simply not a prediction, and the two get sold at the same price.

You will meet this everywhere. A trading strategy that would have returned 40%. A diet study with the outcome picked after the data came in. A hiring model that is 94% accurate — on whom? Every one of them can be checked with the same question, and most of them have never been asked it.

We wanted the metric to work. That is exactly why we had to try to break it, and it is the part of this we would most like you to steal.

Notes & sources

All 223 postseason series from 1998 through 2025, built game by game from the MLB Stats API. 2020 is excluded for its 60-game season and 16-team field. Each series is scored by whether the club ahead on a given measure won it; series are labelled by team identifier rather than by seed, so no home-field or seeding information leaks into the comparison.

Schedule adjustment is the Simple Rating System: a club’s run differential per game plus the mean rating of its opponents, solved by iteration until consistent, centred on zero. The grid is every combination of window length in {5, 10, 15 … 60} and season weight in {0.00, 0.05 … 1.00}, giving 252 candidates.

Single-measure results: won-lost record 50.9%, last-20 raw 53.4%, last-30 adjusted 52.9%, season SRS 56.5%, last-20 adjusted 57.4%, season run differential 57.8%, Pythagorean 57.8%. The apparent gap between raw and schedule-adjusted recent form rests on only 21 series where the two disagree; a McNemar exact test puts it at p = 0.19. We report it as a direction, not a result.

Honest testing: leave-one-season-out re-selects the window and weight on the other 26 seasons and predicts the held-out one, scoring 48.9% over 223 series. The walk-forward version selects only on seasons already played and predicts the next, scoring 56.4% over the 181 series from 2004 onward; plain run differential scores 58.3% over that same subset. The two honest tests differ by more than we would like, which is itself a sign of how little any of these numbers are pinned down.

One limit, stated plainly. None of this proves recent form is worthless in October. It shows that with the data baseball actually produces, we cannot tell whether it helps — and that a search over 252 options will hand you a confident answer anyway.

What to Watch · Tuesday, September 1
Baseball
Blue Jays at Guardians6:40pm ET
27 points of playoff probability ride on it. Both clubs sit near a coin flip, which is where a single game is worth most. If the Guardians win they go to 55%; if they lose, 40%.

Baseball is ranked by championship leverage: we simulate the rest of the season 60,000 times, then split those seasons by who won each of today's games. College football cannot be simulated that way, so it is ranked by how close to a coin flip the ratings make it — a lopsided game teaches you nothing. Tomorrow this block reports what happened to today's pick.

Licensed under The Sports Page License · Borrow it, give it back better · Non-commercial, attribution required
Pass it on.
A few minutes to read. A few seconds to send.
Share on X Facebook LinkedIn Email
The Sports Page
Or scan, for sharing the old-fashioned way.
QR Code to thesportspage.net
thesportspage.net
© 2026 The Sports Page · A Statistical Dispatch for Friends & Family