How Expected Goals Models Actually Work — explained

Expected goals is not one formula but a family of models trained on historical shots. Knowing what the model saw before it printed a number is the difference between using the metric and being used by it.
What the model is trained on
Every expected-goals model begins with a historical archive of shots. Each attempt is labelled with whether it became a goal and with a set of descriptive features: where it was struck from, the angle to the goal, the part of the body used and the manner of the pass that preceded it. The model learns the relationship between those features and the probability of scoring, and applies it to a new shot.
The output is therefore a probability, not a tally. A shot valued at 0.10 is one that a comparable population converted one time in ten. Because probabilities can be summed, a team's expected goals for a match is the total of the values of its attempts, which is what makes the metric portable across sides and seasons.
Why the model is deliberately average
The training population is the whole league or the whole archive, not the individual. A model built this way answers the question how often is a shot like this converted by an average player, and the deliberate consequence is that exceptional finishing is invisible to it. A striker who repeatedly beats the model is not proving the model wrong; he is demonstrating exactly the skill the model was designed to leave out.
That design choice is the source of most public confusion about the metric. Expected goals describes chance quality, not player quality, and the two are different objects. Comparing a player's goals with his expected goals measures the residual, and the residual entangles finishing skill, shot selection and luck.
The features that carry the most weight
Distance and angle dominate every published model. A shot from the centre of the penalty area is worth several times one from the edge of the box, and the effect of angle is nearly as strong: the same distance at a tight angle to the goal line is worth a fraction of the central opportunity. Everything else in the feature set is a refinement on top of those two geometric facts.
Beyond geometry the model looks at how the chance was created. A shot following a cut-back across the face of goal, or one struck after the goalkeeper has already moved, is valued above a shot of the same geometry taken from a standing start. Providers differ most from one another in how finely they describe that preceding action, because it needs tracking data rather than a list of coordinates.
Where the model leaks
The most persistent leak is the identity of the shooter and the state of the defence. Two shots from identical coordinates are not identical events if one is taken by a specialist finisher and the other by a centre-back, and no average-based model can express that without becoming a player model. The second leak is defensive pressure, which is visible only in tracking data and is the reason a video-only provider produces systematically flatter shot values.
A third leak is scoreline and time. A team chasing a match takes shots it would not otherwise take, which pushes its expected-goals total up without improving its attack. Reading a single match without the game state is the most common way the metric is misapplied, and it is a reading error rather than a model error.
Summing probabilities across a season
Over a large number of shots the sum of probabilities is a well-behaved statistic, because the errors of individual estimates cancel more than they compound. That is why expected goals is far more informative over a season than over a match, and why a club's underlying numbers stabilise long before its league position does.
The practical rule is that match-level expected goals describes what happened in that match, while season-level expected goals describes the quality of the team. Treating the first as evidence about the second is the same error as judging a batsman's technique from a single innings, and it produces the same kind of confident nonsense.
How providers diverge on the same match
Two providers watching the same fixture will usually produce totals within a few tenths of each other, but not identical ones. The differences accumulate from small decisions: whether a blocked shot is logged at all, how a deflected effort is attributed, whether a header from a set piece is treated with a separate feature set. None of those choices is a mistake, and each of them shifts the total.
The useful habit is to fix one provider for a comparison and to say which one. Mixing totals from two providers inside one table creates apparent changes in performance that are artefacts of the definition rather than facts about the football, which is the single most common error in published analytics tables.
Key reference points
- Expected goals is a learned probability over a historical shot archive, not a formula.
- Distance and angle carry most of the weight; every other feature is a refinement.
- The model is trained on an average player, so finishing skill is deliberately excluded.
- Summed probabilities behave well over a season and poorly over a single match.
- Provider differences come from event definitions, not from carelessness.
- Fix one provider and state it before comparing two totals.
| Feature | What the model adds |
|---|---|
| Distance to goal | The dominant term in every published model |
| Angle to goal | Nearly as strong, especially at tight angles |
| Body part | Head and weaker foot carry a discount |
| Pass type | Cut-backs and low crosses raise the value |
| Phase of play | Open play, set piece and counter differ |
| Defensive pressure | Requires tracking data; absent from video feeds |
| Game state | Scoreline bias, a reading trap rather than a feature |
Expected goals describes the chance, not the finisher.