xG & Advanced Stats

Regression to the Mean, Read Properly: the numbers

Advanced / contestedEvergreen metric reference
An aerial view of a pitch divided into zones

Regression to the mean is the most misunderstood idea in football analytics and the most useful. It does not say a hot player will become bad; it says the part of his run that was luck will not repeat.

The idea in one paragraph

Whenever a measurement combines a stable component with a random one, the extreme readings are extreme partly because of the stable part and partly because of luck. On a repeat measurement the stable part persists and the lucky part does not, so extremes move towards the average. Nothing has improved or declined; the measurement has simply shed the component that was never going to repeat.

This applies to football totals with unusual force because a goal is a rare event. A player who scores in six consecutive matches is being described partly by his ability and largely by a sequence of coin flips that happened to land well, and the sequence is what regresses.

What actually regresses

Goals regress towards expected goals, because a conversion rate above the population average is the least stable component of a scoring record. Save percentage regresses towards a value near the league mean, because shot-stopping above expectation is mostly variance over a season. League position itself regresses, because a table is built from a handful of one-goal margins.

What does not regress is volume. A team that takes a great many shots and creates a great many high-value chances is not benefiting from luck; its underlying numbers are the stable signal. The distinction is always between a rate, which regresses, and a volume, which does not.

The trap of regression as a prediction

It is tempting to convert the idea into a forecast: this team has overperformed, so it will now decline. That is a misuse. Regression to the mean describes what happens to an average case; it does not tell you which case you have, and it says nothing about the direction of the next result for a specific team.

The correct use is comparative and probabilistic. Between two sides with identical points totals, the one whose underlying numbers are stronger is likelier to sustain its position. That is a statement about likelihood across many teams, and it is silently converted into a claim about one team whenever it is stated as a prediction.

Why the effect is strongest at the extremes

The distance a reading travels back towards the mean is proportional to how far out it sat, because a larger deviation implies a larger share of luck in the observation. A club that finished far above its underlying numbers has more to give back than one that finished slightly above, and the apparent harshness of the effect at the top of a table is a direct consequence of that proportionality.

This also explains why league tables are more informative in the middle than at the extremes over a single season. A mid-table side's points total is mostly its true level; a title winner's or a relegated side's total is a mixture of level and sequence, which is why the following season so often looks like a correction.

Applying it to a single player

For an individual, the honest procedure is to identify which part of a record is a rate and which is a volume. A striker with thirty goals from thirty high-value chances has a repeatable profile. A striker with thirty goals from twelve expected has a profile that will most likely settle lower, not because he is a worse player than he looked but because the goals that exceeded his chance quality were never going to be reproducible at that rate.

None of this licenses saying that a player will fall away. It licenses saying that the part of his record that was a rate is the part to discount, and that the volume of chances is the part to keep. Stated that way, regression becomes a description of which numbers to trust rather than a prophecy about a person.

Why it feels counter-intuitive

The intuition that resists regression is narrative. A hot run feels like form, and form feels like a state that a player carries forward. The statistical view is colder: a run is an observed sequence of outcomes, and sequences contain luck by construction. Both views can be reconciled by saying that form exists as a small effect, while the run in which it is observed is inflated by a larger random one.

The practical tell is the shape of the underlying numbers during the run. If the volume of chances rose alongside the goals, something real is happening. If the goals rose while the chances stayed flat, the run is a conversion spike, and conversion spikes are the least durable statistic in the sport.

Key reference points

  • Regression describes what happens to an average case, not to a specific one.
  • Rates regress; volumes generally do not.
  • The further a reading sits from the mean, the further it tends to travel back.
  • League position regresses because tables are built from one-goal margins.
  • Use it to choose which numbers to trust, never as a forecast about a club.
  • If the chance volume rises with the goals, the improvement is real.
What regresses and what does not
QuantityTendency to regress
Conversion rateStrong
Save percentageStrong
Goals versus expected goalsStrong over one season
League points totalModerate to strong
Volume of shots createdWeak, it is the stable signal
Volume of high-value chancesWeak, treat as the true level

Regression to the mean is not a claim that a run was fake; it is a claim that the unrepeatable part of a run will not be repeated, which is a weaker and more useful idea.