xG & Advanced Stats

Sample Size and When xG Becomes Reliable — explained

Advanced / contestedEvergreen metric reference
A football analytics scene

Every number on an analytics site is a sample estimate, and the size of the sample decides how much of the figure is signal. Knowing the thresholds at which football's metrics stabilise is what separates a reading from a guess.

Stabilisation, not truth

A statistic stabilises when adding more observations stops moving it in a systematic direction. Stabilisation is not the same as correctness: a metric can be stable and still biased if its definition is wrong or its sample is drawn from one competition. What stability buys is the guarantee that the number is describing a level rather than a streak.

Football's event rates are low enough that this matters for every metric in common use. The usual way to measure it is to split a dataset and compare the halves; the point at which the halves agree is the point at which the metric has stopped being a description of luck.

The ordering of stability

Metrics stabilise at very different speeds. Volume measures such as the number of passes attempted or the number of defensive actions settle quickly, because each event is common. Value measures such as expected goals take longer, because shots are rarer than passes. Rate measures such as conversion percentage and save percentage take longest of all, because they are ratios of rare events.

This ordering is the practical core of the subject. A manager can be described by his passing volume within a couple of months, by his expected goals over most of a season, and by his finishing rate almost never. Quoting them with the same confidence is the most common analytical error in public football writing.

Rough thresholds in practice

As a working guide, team-level expected goals becomes reasonably stable across a season, team-level volume statistics across ten to fifteen matches, and player-level conversion only across multiple seasons. Individual defensive metrics are noisier still, because a defender may face a handful of contested actions per match and the sample grows slowly.

The thresholds are approximations and they differ between competitions with different tempo. They are useful not as a rule to recite but as a prompt: whenever a figure is quoted, ask which column of this ordering it belongs to before treating the difference between two values as meaningful.

Why the average can hide the case

A stable team average can conceal two very different processes. A side with a solid season-long expected-goals figure may have earned it evenly or may have produced two spectacular matches inside an otherwise poor run. The average treats those identically, and the distribution behind it is where the information about reliability lives.

The remedy is to look at the shape, not only the mean: the spread of match values, the number of matches in which the side created almost nothing, and whether the strong performances cluster around a particular opponent type. A mean with a heavy tail is a mean that should be quoted with a warning attached.

Practical consequences for reading analytics

Any claim of the form this team is better than its results needs a sample big enough for the underlying metric to have stabilised, and the smaller the league gap the bigger that sample must be. Over ten matches almost every side is within noise of almost every other on the metrics that matter, which is why early-season underlying-numbers tables are a genre of confident nonsense.

The same discipline applies to players, where the samples are smaller still. A midfielder who has played eight matches has taken perhaps a dozen shots: not enough to say anything about his shooting, though possibly enough, if the volumes are extreme, to say something about his role. Distinguishing role from level is exactly what the sample size allows and what the headline figure conceals.

Reporting a figure honestly

The honest sentence names the sample. Nine goals from a non-penalty expected goals of six across a full season is a claim; he is overperforming is not, because the second version has quietly deleted the size of the evidence. The same applies to a metric quoted from a single tournament, where the sample is so small that the figure is closer to an anecdote than to a measurement.

Where the sample is too small to support a reading, saying so is not a failure of analysis. It is the analysis. A reference site that marks unstable numbers as unstable is more useful than one that prints them beside stable numbers with no distinction, and the marking is what turns a table of figures into a table of evidence.

Key reference points

  • Stabilisation means a metric describes a level rather than a streak, not that it is true.
  • Volumes stabilise first, value measures next, rates last.
  • Team expected goals needs roughly a season; conversion needs multiple seasons.
  • A stable mean can conceal a heavy tail of very poor matches.
  • Name the sample in the sentence, or the claim is unfalsifiable.
  • Marking a figure as unstable is itself a valid analytical conclusion.
Working stabilisation order
Metric classApproximate horizon
Pass volume, defensive-action volumeA handful of matches
Possession and territory sharesTen to fifteen matches
Team expected goalsMost of a season
Player shot volumeAround twenty to thirty matches
Player conversion rateMultiple seasons
Individual defensive duelsVery slow, few events per match

The sample size is not a caveat appended to a metric; it is part of the metric, and a number quoted without it has been quietly modified.