Two Teams a Rank Apart? The Higher One Wins 48.7% of the Time.
When #6 plays #7, the poll says one is better. The field says it is a coin flip, and slightly the wrong way. Two teams have to sit about ten places apart before a ranking tells you anything you could not have guessed.
The question: how far apart do two teams have to be in the rankings before the ranking actually tells you who will win?
The answer: about ten places. Below that, you are reading noise. Teams one or two ranks apart split their games 48.7% to 51.3% — and the 48.7% belongs to the team the poll says is better.
Here is the surprising part. We checked it two completely separate ways, and they agree. One asks the field; the other asks the raters. Neither knows about the other, and both land on ten.
What you get from it: a way to look at any ranked list this autumn and know which gaps are real.
Borrowing a trick from psychophysics
Put two weights in someone’s hands and ask which is heavier. If they are identical, the answer is a guess and you get it right half the time. Make one heavier and heavier, and at some point people start getting it right reliably.
That crossing point has a name. It is the just noticeable difference, and by long convention we set it at 75% — halfway between guessing and certainty.
So ask the same question about rankings. How much heavier does one team have to be before we can tell?
Test one: ask the field
Take every game from 2015 through 2025 where both teams were ranked in that week’s poll. That is 493 games. For each one, note the gap between them and whether the better-ranked team won.
Across all 493, the better-ranked team wins 62.3% of the time. That sounds respectable until you cut it by gap.
One or two apart: 48.7%. Three or four: 53.8%. Five to seven: 58.8%. Not until the gap reaches the low teens does the curve cross 75%.
Test two: ask the raters
Now forget who won. Massey Ratings collects 25 separate rating systems — polls, computers, prediction models — each ranking all 138 teams.
Treat them as 25 judges holding the same two weights. For every pair of teams, ask what share of the systems agree on which is better. That is 9,453 pairs.
Teams one rank apart in the composite: the systems agree only 54.5% of the time. They do not cross 75% until the gap reaches nine to twelve.
Watch the two curves climb together. One is made of football; the other is made of arithmetic. They cross the line in the same place.
A ranking is more precise than it is accuratei. It is printed to the nearest single place and it is good to about ten.
The ProfessorHow wrong can one number be?
Look at what the 25 systems do to a single team. The middle team on the list has a range of 41 ranks between its highest and lowest placement.
North Dakota State sits 83rd in the composite. The systems place it anywhere from 23rd to 138th.
The top is much firmer: Ohio State is first, and no system has it worse than fourth. Notre Dame is fourth, with a range of first to ninth. So the poll is a real measurement at the very top and close to noise through the middle.
And one thing beats the ranking entirely
Split those 493 games three ways — home, away, and neutral site — and something jumps out.
When the better-ranked team plays at home, it wins 77% of the time. On the road, 48%. At a neutral site, in between: 67%.
A ranked team away from home is worse than a coin flip, whatever the number beside its name. Where a game is played is worth about thirteen ranks of polling, and no poll shows you that. We take that number apart in a companion piece later this week.
What to take home: ask for the error bars
None of this says rankings are worthless. The top ten is real, big gaps are real, and the systems agree strongly once teams are far apart.
It says the number is reported far more finely than it can support. We print a ranking as a single integer, and it deserves an interval around iti about ten places wide.
So this autumn, when you see #6 play #7, do not ask which is better. Ask whether anything separates them at all — and then check who is at home.
The habit travels. Any time you meet a ranked list — universities, hospitals, countries, restaurants — the ranking is precise and the measurement underneath is not. Ask how far apart two things must be before the order survives a second opinion. Usually nobody has checked.
Notes & sources
Game data from the CollegeFootballData API: every regular-season game 2015–2025 in which both teams appeared in that week’s AP poll, 493 in total. Ties excluded. Of the 493, 190 were played at the better-ranked team’s home field, 219 at its opponent’s, and 84 at neutral sites. The gap is the absolute difference in AP rank; the outcome is whether the better-ranked side won. Confidence intervals are Wilson 95%: 1–2 apart is 48.7% [0.38, 0.60], 12–16 is 77.8% [0.67, 0.86]. Note that no single band clears 0.75 with its whole interval — the point estimate crosses in the low teens, and the honest reading is “around ten” rather than an exact threshold.
Rating systems from a Massey Ratings composite export, 2026 preseason: 138 teams, 25 systems with sufficient coverage. Pairs required at least 15 systems rating both teams, giving 9,453 pairs. Agreement is the share of those systems ordering the pair the same way as the composite. Median across-system range per team, 41 ranks; median standard deviation, 10.2 ranks.
The 75% criterion is Thurstone’s, from the law of comparative judgment (1927): two stimuli have a 50% chance of being ordered either way, so the halfway point between chance and certainty is 75%. The application of that idea to measures that are not physical — and the argument that a measure needs calibrating against behaviour before its numbers mean anything — is from Sechrest, McKnight & McKnight, Calibration of measures for psychotherapy outcome studies, American Psychologist 51(10), 1996.
One limit worth stating: preseason ratings are the noisiest of the year, so the 41-rank median spread should tighten as real games arrive. We will re-run this in November and report whether it did.