Data Hygiene and Verification: definition and limits

A dataset is a constructed object, and most analytical errors are introduced before any analysis starts. The checks that catch them are unglamorous and they are the difference between a table that can be trusted and one that merely looks tidy.
The class of errors
Data hygiene addresses errors introduced by the process rather than by the football: duplicated rows, matches recorded twice under different names, coordinates from the wrong end of the pitch, players listed under inconsistent spellings, and fixtures joined to the wrong season. None of these is a modelling problem and all of them corrupt a model's output.
The distinguishing feature is that they are invisible in the result. A duplicated match inflates every total derived from it by a fixed amount, and the inflated total looks entirely plausible. Hygiene checks exist to find errors that the analysis cannot detect because the analysis trusts its input.
The essential checks
Five checks catch most problems. Uniqueness, to confirm that each match and each player appears once. Range, to confirm that coordinates and counts fall inside physically possible bounds. Referential integrity, to confirm that every event belongs to a match that exists. Internal consistency, to confirm that team totals equal the sum of their players. And temporal consistency, to confirm that dates and sequences are ordered correctly.
Each check is cheap and each has caught a known class of error. The five together form a short routine that should run on every dataset before any metric is computed, and running it takes a small fraction of the time that correcting its absence afterwards requires.
Name and identity resolution
Player and club names are a persistent source of error. The same player may appear under several transliterations across competitions, and the same club may appear under different names across seasons for commercial reasons. Joining two datasets on a name field without resolution will silently fail for the affected rows.
The standard remedy is a stable identifier for each entity, maintained independently of the display name, and a mapping table where names change. Where an identifier scheme is not available, the resolution has to be manual and it should be documented, because the affected rows are invisible in the output.
Provenance and extraction discipline
Every figure in a table should be traceable to a single extraction from a single provider under a single definition. Tables assembled from several extractions taken at different times will mix model vintages and produce internal inconsistencies that no amount of analysis will reveal.
The practical discipline is to extract once, store the extraction, and compute every figure in the table from that store. Re-deriving numbers from a fresh extraction mid-project is the standard way a coherent table becomes quietly incoherent, and the provenance record is what prevents it.
The verification of records
For the record material on this site, hygiene takes a different form: verification against archives. A tally is only admitted once the matches it counts can be identified in a source that lists them, and disputed classes are reported as ranges rather than resolved by preference.
Verification also means checking the negative. Before publishing a record, the verification asks whether a better mark exists that the compiler has missed, and whether the definition admits a class of fixtures that would change the ordering. Both checks are cheap and both have caught errors in published lists.
Why hygiene is the whole game
A model can be elegant and a dataset can be wrong, and the wrong dataset wins. Almost every well-publicised analytical failure traces back to an input problem rather than to a modelling one, because the inputs are assembled quickly and trusted absolutely.
This site's method is therefore to describe the rule, verify the input, mark the uncertainty and date the reading. Those four steps are slower than printing a number, and they are the only reason the number can be cited by anyone else.
Key reference points
- Most analytical errors are introduced before the analysis begins.
- Uniqueness, range, integrity, consistency and temporal checks catch most problems.
- Name and identity resolution needs stable identifiers, not display names.
- Every figure in a table should come from one extraction and one definition.
- Verification of records includes checking whether a better mark was missed.
- A wrong dataset defeats a correct model every time.
| Check | Error it catches |
|---|---|
| Uniqueness | Duplicated matches or players |
| Range | Impossible coordinates or counts |
| Referential integrity | Events joined to a missing match |
| Internal consistency | Team totals not matching player sums |
| Temporal ordering | Fixtures assigned to the wrong season |
| Provenance | Mixed model vintages inside one table |
Hygiene is the least interesting part of analytics and the part that decides whether any of the rest is true, which is why it belongs in a reference rather than in a footnote.