Why Two Credible Totals Disagree: a data reference

When two reputable providers publish different figures for the same thing, the instinct is to ask which one is wrong. The correct question is almost always which definition each one used, because both are usually arithmetically correct.
Mechanism one: the counting rule
The most common cause is a different counting rule. One source counts league matches only; another counts all competitions. One counts appearances as a start; another counts a substitute appearance as an appearance. The two figures are then different because they are answers to different questions.
Identifying this mechanism is usually a matter of reading the caption. Where a source does not state its rule, the disagreement cannot be resolved at all, and the reasonable response is to treat both figures as applicable under unknown conditions.
Mechanism two: the event taxonomy
The second mechanism is the definition of the underlying events, described elsewhere on this site. A provider that records blocked shots as attempts will report a higher shot count, and every rate derived from shots will be correspondingly displaced. The difference is systematic, so it appears across every row of the table.
This mechanism is harder to spot because it does not usually appear in a caption. The tell is that the disagreement is consistent in direction across many rows, whereas a random error would scatter. Uniform disagreement between two tables is almost always a definitional difference.
Mechanism three: the sample window
A third cause is that the two figures cover different periods. One was computed after the final match of a campaign; the other was computed a week earlier. For any quantity that accumulates, the difference is the events of the missing week, and there is no error at all.
This mechanism is the reason dates matter even for figures that look permanent. A career total, a head-to-head series and a season aggregate are all moving quantities, and two figures that disagree may simply have been taken at different times.
Mechanism four: the model vintage
When a provider revises its shot-value model, every figure computed under the old model differs from the corresponding figure computed under the new one. Historical totals are frequently refreshed under the new model, but not always, and a table assembled from re-computed and un-recomputed values will be internally inconsistent.
This is the most insidious mechanism because it is invisible. A table can look entirely coherent while mixing two model versions, and the resulting distortions are small enough to pass unnoticed. The only defence is to source every figure in one table from one extraction.
Mechanism five: the aggregation choice
The last mechanism is how the numbers were combined. A team's season expected goals can be the sum of its match values or the mean multiplied by matches, and where matches are missing the two differ. A rate can be computed per ninety minutes or per match. A career rate can be weighted by seasons or by appearances.
Each of those choices is defensible and each produces a slightly different figure. Where two totals disagree by a small amount, this is the likeliest explanation, and it is usually detectable by recomputing the figure from its components.
How to resolve a disagreement
The procedure is short and mechanical. Check the counting rule, check the taxonomy, check the date, check the model vintage, check the aggregation. In the great majority of cases one of the five explains the whole difference, and in the remainder the disagreement is small enough to be within rounding.
Where none of the five applies, the possibility of a genuine error arises and the figures should be treated with caution. But that case is much rarer than the frequency of public disagreement suggests, because almost every disagreement is definitional and almost nobody checks the definitions before arguing.
Key reference points
- Counting rules, taxonomies, dates, model vintages and aggregation choices explain most gaps.
- Systematic differences across many rows indicate a definitional cause.
- Moving quantities disagree when they were taken at different times.
- A table can silently mix two model vintages and look coherent.
- Recomputing from components usually reveals the aggregation choice.
- Genuine error is the least likely explanation for a disagreement.
| Mechanism | Detection |
|---|---|
| Counting rule | Read the caption |
| Event taxonomy | Difference is uniform in direction |
| Sample window | Compare the dates |
| Model vintage | Compare extraction dates |
| Aggregation choice | Recompute from components |
Two credible totals for the same thing are almost always two correct answers to two different questions, and the work of reading analytics is largely the work of finding out which question was asked.