Touchline Intelligence
Loading live model data…
The page reads the model API on every request, so the first paint waits for the real numbers rather than showing placeholders.
Touchline Intelligence
The page reads the model API on every request, so the first paint waits for the real numbers rather than showing placeholders.
Evaluation design, leakage controls, and release engineering behind the served probabilities. The model card carries the complete record.
Everything builds on StatsBomb Open Data pinned to one commit. The fixed cohort keeps 5,606 eligible non-penalty shots with 507 goals: regulation and shootout penalties, own goals, and rows missing required fields are outside the modeled population. Post-shot information never enters the feature space.
| Tournament | Matches | Events | Shots | Role |
|---|---|---|---|---|
| FIFA World Cup 2018 | 64 | 227,825 | 1,706 | Development |
| UEFA Euro 2020 | 51 | 192,664 | 1,289 | Development |
| FIFA World Cup 2022 | 64 | 234,637 | 1,494 | Calibration |
| UEFA Euro 2024 | 51 | 187,924 | 1,340 | Holdout |
| Total | 230 | 843,050 | 5,829 | — |
Coverage terms were reviewed against the source repository and are documented with the pinned revision in DATA_SOURCE.md.
Each tournament has exactly one role, decided in advance. Development rows may shape the model; calibration rows may fit one transform; the holdout may answer one question once. What each split is forbidden from doing matters as much as what it does.
| Split | Tournament | Matches | Shots | May | May not |
|---|---|---|---|---|---|
| Development | WC 2018 + Euro 2020 | 115 | 2,872 | Features, preprocessing, model selection, grouped cross-validation | Calibration or holdout claims |
| Calibration | WC 2022 | 64 | 1,430 | Fit the Platt transform and apply a frozen adoption rule | Base refit, feature selection, candidate selection |
| Tournament holdout | Euro 2024 | 51 | 1,304 | One predeclared raw-versus-calibrated evaluation | Any retrospective decision at all |
Cross-validation inside development uses five deterministic, match-grouped folds: every shot from a match stays on one side of a fold boundary, so no model ever validates on shots from a match it trained on. The Euro 2024 holdout is a tournament holdout, not a pure time-series claim, because date and composition change together.
The model sees 16 columns: two continuous geometry features from the shot location, plus categorical indicators for body part, technique, and play pattern. Distance uses the recorded StatsBomb coordinate system; the goal angle uses a numerically stable two-post form chosen because the common single-arctangent expression is measurably wrong for 38 shots near goal on this cohort.
Categorical vocabulary is fitted once on development rows without labels; levels under 25 shots merge into a rare bucket; one reference level per field stays implicit. An unseen future level maps to the all-zero reference encoding rather than crashing a serving request or silently expanding the contract. Scaling during cross-validation is fitted on each fold's training rows only.
Excluded on purpose: provider xG, post-shot fields, outcomes, future events, and two true-only presence annotations that failed a pre-registered consistency gate despite slightly better aggregate metrics.
Complexity had to earn its place. Every candidate trained on identical development rows and folds under one locked protocol; a challenger could replace the incumbent only by beating it on mean log loss beyond the incumbent's fold variance without worsening Brier, calibration, or stability. Ties stayed with the simpler model.
| Candidate | Log loss | Brier | Decision |
|---|---|---|---|
| Constant (training-fold goal rate) | 0.3019 | 0.0815 | Evaluation reference |
| Geometry logistic (distance + angle) | 0.2700 | 0.0747 | Improved on constant; not the final feature set |
| Full logistic (context + two presence flags) | 0.2620 | 0.0725 | Presence flags failed a pre-registered consistency gate |
| Selected: full logistic minus presence flags | 0.2634 | 0.0730 | Selected base estimator |
| Gradient boosting (registered 12-point grid) | 0.2680 | 0.0745 | Did not meet the replacement conditions |
| PyTorch MLP (16 → 8 → 1, 145 parameters) | 0.2667 | 0.0739 | Did not meet the replacement conditions |
Boosting and the neural challenger were real experiments with real engineering, and both lost. Recording losing challengers is part of the point: the shipped model is the one the protocol chose, not the most impressive one available.
After the base model was frozen, World Cup 2022 fitted a single Platt transform under a rule fixed in advance: adopt only if calibration deviation improved on supported bins without worsening either proper score. It passed, and the calibrated variant was locked in before Euro 2024 was opened.
The holdout then disagreed: calibration slightly worsened log loss and Brier there. The shipped release keeps calibration anyway, because redeciding after seeing the holdout would convert the final evaluation into a second selection set. Both variants are reported on the model page exactly as measured.
Uncertainty comes from a match-clustered paired bootstrap: 2,000 resamples that draw whole matches with replacement, since shots within a match share context.
The release is a content-hashed packet: model artifact, metrics, calibration decision, and split assignment each carry a SHA-256 digest, and the serving bundle refuses to start from material that does not match. A reproduction run at the registered commit, under the recorded lockfile, produced a byte-identical artifact and equal metrics on the development rows.
The model API publishes this identity on every response, and this site verifies it across endpoints on every request. If the hashes disagree, the page shows you the disagreement instead of a blended view.
What is not claimed: equivalence to any commercial xG model, live drift monitoring, or license to redistribute the source data. Historical row-level publication stays closed pending written provider direction; the API fails closed on that boundary by design.
Four international tournaments from one pinned Open Data revision. Nothing here establishes performance in domestic leagues, women's or youth football, other eras, or another provider's event definitions.
Euro 2024 differs from earlier tournaments in both date and composition, so the result cannot separate temporal drift from distribution shift.
The World Cup 2022 transform met its adoption rule but worsened both proper scores on Euro 2024. One source and one destination tournament cannot characterize calibration transport generally.
A support rule (at least 50 shots, 5 goals, 5 misses, 10 matches) blocks interpretation of the thinnest slices, and even supported slices carry sampling uncertainty.
Location, body part, technique, and play pattern. No tracking data, no StatsBomb 360, no goalkeeper or defender positions, no player ability, no game state. Provider xG is excluded by constraint.
Coefficients describe associations in recorded event data. The model cannot say how conversion would change under an intervention.
This page is the summary. The canonical records are the model card, the data source review, the technical write-up, the immutable experiment artifacts in /experiments, and the model API reference.