You Can’t Fix a Broken Study With a Better Formula.
Every claim is secretly a comparison.
“Signing that quarterback made the team better.” Better than what? Than the team without him? Than an average replacement? Than the imaginary version of this season where they signed someone else? Until you can name the thing you are comparing to, you do not have a finding. You have a feeling with a number stapled to it.
This is the heart of research design, and it comes before any math. Take the classic: “They won eight straight after the trade.” Compared to what? Maybe the schedule softened. Maybe three injured starters came back. Maybe they were due. The eight-game streak is being silently measured against an alternate season that never happened — and whether that comparison is fair was decided by how the question was set up, not by how cleverly the streak is analyzed afterward.
So before you trust a sports statistic, ask the three design questions: what did they measure, on whom, and compared to what? If the comparison is unfair, a fancy model does not fix it. It just launders the unfairness into a result that looks confident.
The design sets a ceiling the analysis can’t break through.
Bad design shows up in a few familiar shapes, and none of them is a math problem. Selection: who got into the sample and who quietly didn’t — only the coaches who were hired, only the prospects who stuck, only the teams still standing (that one has its own primer, Survivorship Bias). The design chooses the sample, and the sample chooses the answer. No control: “the team improved after the change” with nothing to compare against — a before-and-after with no counterfactual, which is a story, not evidence. Baked-in confounds: comparing a starter’s numbers to a benchwarmer’s, or one era’s to another’s, when the roles or the conditions were never the same to begin with.
Here is the rule that makes design the most important word in this whole newsletter: the design of a study sets a ceiling on what can ever be learned from it, and no analysis can rise above that ceiling. A brilliant model applied to a broken comparison produces a brilliant-looking answer to the wrong question. This is why careful people obsess over the design before the data arrives — because a flaw in the formula can be fixed later, and a flaw in the design cannot be fixed at all.
It is also the humbler cousin of a bigger idea (see Do We Believe the Data, the Model, or the Theory?): even a perfect model of badly gathered data will faithfully, confidently answer a question you did not mean to ask.
The full treatment: designing before you compute.
What makes a comparison fair, how sampling and controls are built into a study, and why the design decides the ceiling — the throughline runs through The Sports Page’s companion statistics textbook, a free, open, graduate-level text, beginning in The Basics. Read it free here — the same “read, play, learn” idea, one rung deeper.
Where this concept shows up in The Sports Page
- Survivorship Bias and Base Rate (Concepts No. 4 and 1) — the two most common ways a sample is chosen unfairly before you start.
- Counterfactuals (Concept No. 8) — the “compared to what” that a fair design has to supply.
- Sunday Editions — every prediction is scored against a fair baseline, not against nothing.
- Any two-team comparison — we hold the roles and eras honest before we let the numbers race.