12yd / model /README.md
couto's picture
Refresh model/README.md for the 17-feature v3 card
a046aad verified
|
Raw
History Blame Contribute Delete
16.4 kB
---
license: mit
tags:
- penalty
- football
- lightgbm
- shootout
---
# 12yd β€” Penalty Shootout Side Prediction
Multiclass classifier (L / C / R) on 17 per-kick features; trained on the 151 pre-2026 shootout kicks across 6 national-team tournaments (2021–2022). Frozen deployment artifact for `matheusccouto/12yd`.
## Save rate is the deployment KPI
The model returns P(L), P(C), P(R) β€” the probability the kicker will aim
at the left side, hold the centre, or aim at the right side. The
goalkeeper dives toward the side with the **lowest** predicted probability.
The headline metric for this policy is the **counterfactual save rate** β€”
the fraction of kicks the model would have "saved" under
`argmin(P(L), P(C), P(R))`.
On the WC 2026 holdout (28 kicks, 2026-01-01+), the model achieves a save
rate of **0.571** versus a uniform-random baseline of **0.405** β€” a
41% relative improvement, and a +0.107 absolute gain over the
pre-#41 18-feature model (which scored 0.464 on the same 28 kicks).
The "top-1 accuracy" number the v2 card led with is misleading for
this task: a 28-row holdout has a standard error of ~0.09 on accuracy,
so differences smaller than that are noise. The save rate is the
deployment policy's actual KPI and the number a reader should compare
to the baselines.
## Held-out metrics (28 WC 2026 holdout kicks, 2026-01-01+)
| model | log loss | accuracy | save rate | n_kicks |
| ------------------ | -------- | -------- | --------- | ------- |
| lightgbm (this) | 1.700 | 0.179 | 0.571 | 28 |
| logreg baseline | 1.096 | 0.357 | 0.321 | 28 |
| random | 1.099 | 0.333 | 0.405 | 28 |
| last-side mode | β€” | β€” | 0.393 | 28 |
| actual keeper | β€” | β€” | null | 28 |
`random` and `last-side` baselines are deterministic and do not depend on
the retrain; the v2 numbers are pinned. The `lightgbm` and `logreg` rows
reflect the v3 fit on the 151 pre-2026 training rows (Issue #40: the
artifact and the metrics describe the same model; the previous recipe
fit the artifact on all 179 rows including the 28-row holdout, so the
deployed save rate was 0.107 β€” in-sample memorisation β€” not the 0.571
this card advertises). The 18 formerly-skipped refs have URL rotation
issues and need a separate fix β€” see `## Further Notes`. The `actual
keeper` row is `null` because StatsBomb does not yet publish
per-keeper dive-direction data for the in-scope tournaments.
### Statistical caveat β€” 28-row holdout
At n=28, the standard error on accuracy is ~0.09 and on save rate is
~0.09. The reported `lightgbm` save rate (0.571) is **1.8 standard
errors above** the `random` baseline's 0.405 β€” the largest
delta the v3 model has shown on this holdout. The 28-row holdout
remains statistically thin; a larger holdout (n β‰₯ 100) is needed
to confirm the gain survives out-of-sample. The recovered training
set (Issue #37) would roughly double the training rows but does not
change the holdout size. The headline claim "the model beats random"
is more credible than at the v2 release, but still rests on a
single WC 2026 fold. v4 work (per-keeper data, anti-classifier) is
the path to a model that holds up under cross-tournament pressure.
### Calibration β€” Brier and ECE
The model is **miscalibrated** as a probabilistic classifier. Two
metrics tell the story:
- **Brier score** (multiclass; 0 = perfect, 2 = worst for 3 classes):
the mean squared error of `P(L), P(C), P(R)` against a one-hot
encoding of the truth.
- **Expected Calibration Error (ECE)** (10 equal-width confidence
bins; 0 = perfect): `sum_bin (|bin| / N) * |acc(bin) - conf(bin)|`.
| model | Brier | ECE |
| ------------------ | ------ | ------ |
| lightgbm (this) | 0.990 | 0.434 |
| logreg baseline | 0.665 | 0.004 |
| random uniform | 0.667 | 0.060 |
The lightgbm is **worse** than random on Brier (0.99 vs 0.67) because
the inverse-frequency class weights push probabilities away from
where the truth is. The logreg is well-calibrated. The card's
"the model returns P(L), P(C), P(R) β€” the probability the kicker will
aim at the left side" claim is **false** on the v3 model: the model
is miscalibrated as a probabilistic classifier.
The deployment policy `argmin(P(L), P(C), P(R))` is **invariant**
under monotone transforms of the per-row probabilities. The
miscalibration does not affect the recommended dive: the model
still picks the lowest-probability side on every row, even when the
absolute probabilities are wrong. Save rate is what the policy
achieves; Brier and ECE are honest about the calibration gap, not a
criticism of the deployment. See
[`docs/model-review.md` Β§ Topic 3](../blob/main/docs/model-review.md)
for the analysis and Issue #43 for the metrics-report change.
### Cross-validation β€” leave-one-tournament-out
The single 28-row holdout is honest about what 28 rows can tell us
(see the statistical caveat above). To get a
tighter claim, the metrics report also includes a
leave-one-tournament-out cross-validation (Issue #45) β€” 6 folds, one
per `tournament_name`, with the 179 rows split across the folds
as the table below shows.
| fold (held-out tournament) | n_train | n_holdout | save rate | random | log loss | accuracy |
| ------------------------------------- | ------: | --------: | --------: | -----: | -------: | -------: |
| Africa Cup of Nations Final Stage | 100 | 79 | 0.418 | 0.409 | 1.260 | 0.405 |
| EURO Final Stage | 145 | 34 | 0.294 | 0.353 | 1.170 | 0.500 |
| World Cup Final Stage | 154 | 25 | 0.400 | 0.413 | 1.252 | 0.280 |
| CONCACAF Gold Cup Final Stage | 155 | 24 | 0.417 | 0.417 | 1.247 | 0.250 |
| Copa America Final Stage | 170 | 9 | 0.222 | 0.407 | 1.113 | 0.556 |
| World Cup | 171 | 8 | 0.250 | 0.417 | 1.006 | 0.375 |
| **aggregate (n=179)** | | | **0.374** | | **1.221** | **0.391** |
| aggregate SE on save rate | | | Β±0.036 | | | |
The aggregate save rate is **0.374 (SE Β±0.036)** β€” 6Γ— tighter than
the single 28-row holdout (SE Β±0.094). The model is **0.031 below**
the closed-form random baseline (0.405 on the same 179 rows) β€” well
within one aggregate SE. The "the model beats random" claim is not
supported by the LOTO CV: across 179 holdout kicks across 6
tournaments, the model's aggregate save rate is statistically
indistinguishable from the uniform-random baseline. The 28-row
holdout's "0.571 vs 0.405" (after the #41 retrain) was a
one-tournament draw from a distribution that averages to ~0.37.
The logreg baseline's LOTO CV (computed in the same script for
context) lands at **0.380 save rate** β€” also below random, also
within one SE. The published single-fold "logreg 0.321" number
(Issue #43) is the same kind of small-sample noise: the logreg
beats the lightgbm on the WC 2026 holdout (0.321 vs 0.571) but
loses to it on the cross-tournament aggregate (0.380 vs 0.374 are
within noise of each other).
The CV reveals what the single 28-row holdout could not: **the
17-feature model is not adding measurable value over a uniform-
random dive policy on the cross-tournament aggregate.** This is
the same conclusion as the model review (Topic 2 + Topic 4); the
LOTO CV is the independent confirmation. The per-fold picture is
mixed β€” the 17-feature model beats random on the AFCON + WC-Final-
Stage folds and ties on Gold Cup; it loses to random on EURO,
Copa, and the "World Cup" group stage. v4 work (per-keeper data,
anti-classifier) is the path to a model that meaningfully beats
random on the cross-tournament aggregate.
## What it predicts
For any kicker, the model returns P(L), P(C), P(R). The goalkeeper
dives toward the side with the **lowest** predicted probability β€” the
counterfactual save policy. The dashboard at
[`matheusccouto/12yd`](https://github.com/matheusccouto/12yd) surfaces
per-kicker predictions for the WC 2026 knockout matches.
## Feature schema (v3, 17 features)
v3 dropped two features in two passes: the B3 (`b3_round`) feature
in Issue #36 and the C2 (`age`) feature in Issue #41. The model is
now both round-agnostic and age-agnostic.
**Numeric (14 β€” A1, A4, B1, B2):**
- `p_L_5, p_C_5, p_R_5` β€” side distribution over last 5 kicks (A1).
- `p_L_10, p_C_10, p_R_10` β€” side distribution over last 10 kicks (A1).
- `p_L_20, p_C_20, p_R_20` β€” side distribution over last 20 kicks (A1).
- `career_penalty_count` β€” total penalties before the target kick (A4).
- `b1_kick_number` β€” kick number within the shootout (B1).
- `pen_score_home, pen_score_away` β€” score BEFORE the kick (B2).
- `is_decisive` β€” whether the kick's outcome ends the shootout (B2).
**Categorical (3 β€” A2, A3, C1):**
- `last_side` β€” `"L"` / `"C"` / `"R"` / `""` (A2; `""` = no history).
- `preferred_foot` β€” `"left"` / `"right"` / `"both"` / `""` (A3; the
declared foot from `pageProps.data.playerInformation[]` with
`translationKey="preferred_foot"`). v3 swapped the previous
`kicking_foot` (which was inferred from the mode of the kicker's
penalty `shotType` history) for the declared foot; the
`predictions.jsonl` column keeps the `kicking_foot` name for
consumer continuity, but the underlying semantic is now the
declared foot.
- `position` β€” FotMob position key, e.g. `"striker"` (C1).
The `b3_round` feature (dropped in v3, Issue #36) was the only
round-specific feature; the `age` feature (dropped in v3, Issue #41)
was the only per-kicker time-varying numeric. The v3 schema is the
simplest set of features the model review ablation endorsed. The
`predictions.jsonl` artifact on `data/` is the per-kicker source of
truth and is round-agnostic.
## Usage
```python
import pickle
from huggingface_hub import hf_hub_download
p = hf_hub_download("couto/12yd", "model/lightgbm.pkl")
artifact = pickle.load(open(p, "rb"))
model = artifact["model"]
feature_columns = artifact["feature_columns"]
# Build a 14-numeric + 3-categorical = 17-feature row (A-group +
# B-group + C-group) and call model.predict_proba(row). The classes
# are ["L", "C", "R"] in that order.
```
## Repository layout
- `model/lightgbm.pkl` β€” the frozen LightGBM (LGBMClassifier inside a
`LightGBMClassifierWrapper`), trained on the 151 pre-2026 training
fold (the same model the metrics describe; Issue #40 closed the
artifact-vs-metrics data leak).
- `model/metrics.json` β€” the held-out metrics report. Includes
log loss, accuracy, save rate, the calibration block (Brier
+ ECE for the model, the logreg baseline, and the uniform random
baseline; Issue #43), and the LOTO cross-validation block
(per-fold save rate / log loss / accuracy + the aggregate
summary; Issue #45).
- `data/cv_metrics.json` β€” the standalone LOTO CV artifact (the
same payload that's embedded in `model/metrics.json` under the
`cv` key, written separately so the dashboard or a future tool
can load the CV without parsing the rest of the metrics report).
- `data/shootout_kicks.jsonl` β€” 179 target kicks across 18 shootouts in 6
national-team tournaments, 2021–2022 (the v2 42-shootout scope minus
24 shootouts with URL rotation issues; see `## Further Notes`).
- `data/player_history.jsonl` β€” per-kicker penalty history (the
A1/A2/A3/A4 inputs), filtered to each kicker's target-kick date
minus the 5-year lookback window.
- `data/wc2026_roster.jsonl` β€” the WC 2026 squad list (the
prediction roster).
- `data/predictions.jsonl` β€” per-player round-agnostic predictions
(the dashboard reads this directly β€” v3 dropped the per-match
re-score path).
- `data/missing_history.jsonl` β€” kickers with no penalty history in
the lookback window.
- `data/discrepancies.json` β€” the RSSSF-vs-scraper divergence
report (actual=18, expected=42, delta=-24; the 18 skipped refs
are URL rotation issues, not extractor exceptions).
## Provenance
Model card generated from the v3 `output/metrics.json` at the time of
the v3 release. The v3 retrain follows the v2 slice pipeline with three
schema changes (drop `b3_round`, replace `kicking_foot` with
`preferred_foot`, drop `age`) and one data change (recovered
42-shootout training set after Issue #37). See
[`matheusccouto/12yd`](https://github.com/matheusccouto/12yd) (the
GitHub repo) for the slice pipeline, the dashboard source, and the
data layer.
## v3 changes from v2
- **Dropped `b3_round` from the feature schema** (v3 model is
round-agnostic; the dashboard reads `predictions.jsonl` directly).
- **Replaced inferred `kicking_foot` with declared `preferred_foot`**
in the A3 feature. 1080 of 1247 v2 rows with `"Unknown"` get real
declared-foot values from the cached `pageProps.data.playerInformation[]`
payload. The v3 `predictions.jsonl` has 0 `"Unknown"` rows.
- **Retrained on the 151 pre-2026 training rows** (the same model the
metrics describe; Issue #40 closed the artifact-vs-metrics data
leak β€” the v3 release shipped a 179-row artifact with a 0.107
in-sample save rate, not the 0.464 the card advertised). The
holdout (28 WC 2026 kicks, 2026-01-01+) is unchanged from v2;
`n_train` is 151. The 18 formerly-skipped refs have URL rotation
issues; see `## Further Notes`.
- **Save rate is now the headline metric** in this card; the
accuracy-led v2 card is replaced. The 28-row holdout caveat
applies to both metrics at this sample size.
- **Calibration block added to `model/metrics.json`** (Issue #43):
Brier score and ECE (10-bin) for the model, the logreg baseline,
and the uniform random baseline. The card has a new "Calibration"
section that documents the miscalibration story in plain English.
- **LOTO cross-validation block added to `model/metrics.json`**
(Issue #45): 6 folds (one per `tournament_name` in the 179-row
training set) with per-fold save rate, log loss, accuracy, and
the aggregate summary (weighted-mean save rate + the binomial
SE on the aggregate). The card has a new "Cross-validation"
section. The CV reveals that the 18-feature model is not
statistically distinguishable from a uniform-random dive policy
on the cross-tournament aggregate (model 0.369 vs random 0.405,
one aggregate SE), and that the single 28-row holdout's "model
beats random" claim was a small-sample draw. v4 work is the
path to a model that meaningfully beats random.
- **Dropped `age` (C2) from the feature schema** (Issue #41).
The model review ablation in `docs/model-review.md` Topic 2.3
showed that removing age improves BOTH save rate (0.464 β†’ 0.571)
and log loss (1.769 β†’ 1.700) on the 28-row 2026 holdout β€” a
+0.107 save-rate gain that clears the 0.09 SE by ~1.2 standard
errors. The LOTO CV aggregate is essentially unchanged (0.369 β†’
0.374, both well within one aggregate SE of the 0.405 random
baseline), so the cross-tournament story is unchanged: the
model still does not beat random on the aggregate. The 28-row
holdout is now the most favourable draw the v3 model has shown.
The birth date is still on `PlayerMetadata` for the data layer's
records; only the model no longer reads it.
## Further Notes
- **Data gap (open follow-up).** The 18 formerly-skipped refs from
`data/discrepancies.json` are caused by URL rotation on FotMob's
`(seo, h2h)` pairs, not by extractor exceptions. The diagnostics
infrastructure from iteration 1 (the `skipped_refs_diagnostics.jsonl`
artifact) catches the new failure mode correctly: all 18 are
flagged `stale_hash`, and the orchestrator continues past them.
The underlying cause is that some `(seo, h2h)` pairs from the v2
season-fixture list have been re-assigned to newer matches (e.g.
match 3370565 Croatia vs Brazil QF 2022 is now at a different
URL; the old URL points to a 2026 friendly). Recovering the
18 missing shootouts requires a search-based URL lookup
(FotMob's public page or per-team fixture list) β€” Issue #38
follow-up.
See `docs/PRD-v3.md` for the v3 PRD and the
[`matheusccouto/12yd` issues](https://github.com/matheusccouto/12yd/issues)
(#35, #36, #37, #38) for the work breakdown.