Spaces:
Running on Zero
03 - Modelling and the Severson reproduction
Back to 02: Feature engineering | README
Covers Phase 6, including Gate 2 - the credibility anchor. Every table on
this page is generated from a file in outputs/reports/.
1. GATE 2 - reproducing Severson et al. (2019)
Verdict: PASSED. Full model, primary test: 12.42% mean percent error against a published 9.1% - a ratio of 1.36x, inside the 2.0x tolerance.
Everything downstream - conformal risk control, the cost frontier, chamber allocation - assumes this pipeline extracts the same signal the source literature extracted. If that assumption were wrong, the later results would be internally consistent and externally meaningless. Reproducing a known number is the only way to test it.
| Model | Features | Train | Primary test | Secondary test |
|---|---|---|---|---|
| variance | 1 | 14.18% (pub 9.8%) | 14.74% (pub 14.1%) | 11.39% (pub 14.1%) |
| discharge | 6 | 7.97% (pub 5.6%) | 14.94% (pub 7.5%) | 9.73% (pub 10.7%) |
| full | 9 | 6.77% (pub 4.9%) | 12.42% (pub 9.1%) | 13.99% (pub 15.6%) |
| dummy_mean (floor) | 0 | 29.38% | 34.65% | 35.90% |
| dummy_median (floor) | 0 | 23.98% | 29.27% | 44.94% |
The evaluation design is the published one
Severson et al. used a fixed three-way split: train and "primary test" interleaved over batches 1 and 2, with batch 3 held out entirely as a "secondary test". Our cohort reproduces the published set sizes exactly: 41 / 43 / 40.
Using this project's own repeated grouped CV instead would give a number that is arguably better methodology but not comparable to the published one, and comparability is the entire point of a reproduction. Both are reported: the published design here, the project's CV in section 3.
The gate did its job - it caught a real error
The first attempt read the design as train = batch1, test = batch2, because batch 1 contains exactly 41 cells and the published training set also contains 41. That coincidence is a trap. Gate 2 failed at 25.84%, 2.84x the benchmark.
The failure was diagnostic rather than merely bad:
- Train and secondary-test errors reproduced well; only the primary test was wrong. A broken feature pipeline would have degraded all three.
- The error ordering was inverted. Published primary (9.1%) is lower than secondary (15.6%), which is only possible if the primary test is in-distribution. Ours had primary above secondary.
- The magnitude had a mechanism: batch 2 contains cells down to 150 cycles while batch 1 bottoms out at 533, so training on batch 1 alone forces extrapolation far below the training range - and mean percent error punishes short-lived cells hardest, since predicting 300 cycles for a cell that lasts 150 is a 100% error on its own.
This is precisely what a gate is for: it caught a methodological error before anything was built on top of it. The corrected split is pinned by a regression test.
Honest accounting of the residual gap
The reproduction is 1.36x the published value, not 1.0x. There are two candidate explanations, and they are not equally favourable, so they were distinguished by measurement rather than assumed.
| Explanation | What it would imply | Prediction it makes |
|---|---|---|
| (a) Feature-list reconstruction - the multi-feature sets were rebuilt from the paper's prose, not copied from the authors' code | The pipeline is sound; the inferred feature lists differ in detail | The gap should be largest where reconstruction was hardest and near zero where there was nothing to infer |
| (b) Feature-pipeline error - something upstream computes the ΔQ(V) quantities differently from the source | The pipeline is wrong, and every downstream result is suspect | The gap should appear everywhere, including in the simplest model |
The variance model settles it. It uses exactly ONE feature,
log10|var(ΔQ(V))|, so there is no feature-list ambiguity to reconstruct at
all - either the pipeline computes that quantity as the source did, or it does
not. It lands at 14.74% against a published 14.1%, a ratio of 1.05x.
The multi-feature sets, whose composition had to be inferred, sit further out (discharge 1.99x, full 1.36x). That is the pattern explanation (a) predicts and the opposite of what (b) predicts: a broken feature pipeline would fail the unambiguous single-feature model first, and it does not.
Gate 1 supports the same conclusion independently: the log10 var(ΔQ(V)) vs log10 cycle-life relationship reproduced at R² = 0.8588 against a published ρ = -0.93 (implied R² ≈ 0.86), which is the same ΔQ(V) machinery measured a different way.
This is stated at length because the favourable explanation is the one a reviewer should be most sceptical of, and asserting it without the discriminating test would be exactly the move that deserves scepticism.
Two further differences that cannot be ruled out and are not claimed to be excluded: the cohort here is independently re-derived (including a cycle-life definition recomputed from the capacity series rather than inherited), and three excluded cells could not be independently justified (see docs/01).
2. Does linear fragility limit the reproduction? No.
Phase 5 raised a concern that heavy-tailed features would make linear models fragile, and suggested robust or tree estimators might be needed. Tested directly on the identical split:
| Estimator | Train | Primary test | Secondary test |
|---|---|---|---|
| huber (robust linear) | 6.34% | 13.35% | 16.70% |
| random_forest | 5.43% | 14.70% | 14.62% |
| lightgbm | 2.11% | 17.66% | 15.42% |
The concern does not hold in this setting. Elastic net is the best of the four on the primary test. LightGBM achieves by far the lowest training error and the worst test error - textbook overfitting with only 41 training cells and 9 features.
The earlier concern was not unfounded, it was scoped differently: it arose with the full 57-feature matrix under cross-validation, where a single extrapolating cell (b1c41) dominated MAPE. On the curated 9-feature set the linear model is well-conditioned. Both statements are true; the recommendation that followed from the first was wrong for this case, and the measurement settles it.
3. Benchmark: model x budget
10 models x 5 budgets, evaluated over 50 outer folds
of the repeated grouped scheme (5 folds x 10 repeats). Every value is mean ± std
across folds; the full table with bootstrap 95% CIs is in
benchmarks.csv.
RMSE on log10 cycle life
| Model | N=5 | N=10 | N=20 | N=50 | N=100 |
|---|---|---|---|---|---|
extra_trees |
0.0798 ± 0.0271 | 0.0752 ± 0.0247 | 0.0776 ± 0.0257 | 0.0665 ± 0.0232 | 0.0580 ± 0.0240 |
xgboost |
0.0879 ± 0.0258 | 0.0861 ± 0.0242 | 0.0831 ± 0.0245 | 0.0684 ± 0.0208 | 0.0603 ± 0.0215 |
random_forest |
0.0860 ± 0.0253 | 0.0845 ± 0.0228 | 0.0871 ± 0.0244 | 0.0704 ± 0.0213 | 0.0627 ± 0.0214 |
lightgbm |
0.0878 ± 0.0264 | 0.0836 ± 0.0231 | 0.0894 ± 0.0254 | 0.0686 ± 0.0227 | 0.0665 ± 0.0207 |
ridge |
0.0853 ± 0.0170 | 0.0808 ± 0.0179 | 0.0923 ± 0.0247 | 0.1610 ± 0.1766 | 0.1016 ± 0.0316 |
lasso |
0.0861 ± 0.0170 | 0.0823 ± 0.0177 | 0.1002 ± 0.0301 | 0.1754 ± 0.1423 | 0.1164 ± 0.0468 |
elastic_net |
0.0850 ± 0.0169 | 0.0821 ± 0.0170 | 0.0973 ± 0.0270 | 0.1666 ± 0.1555 | 0.1190 ± 0.0452 |
huber |
0.0870 ± 0.0188 | 0.0768 ± 0.0190 | 0.0991 ± 0.0369 | 0.1786 ± 0.2212 | 0.1201 ± 0.0654 |
dummy_mean |
0.1868 ± 0.0278 | 0.1868 ± 0.0278 | 0.1868 ± 0.0278 | 0.1868 ± 0.0278 | 0.1868 ± 0.0278 |
dummy_median |
0.1879 ± 0.0271 | 0.1879 ± 0.0271 | 0.1879 ± 0.0271 | 0.1879 ± 0.0271 | 0.1879 ± 0.0271 |
Mean absolute percentage error, raw scale
| Model | N=5 | N=10 | N=20 | N=50 | N=100 |
|---|---|---|---|---|---|
extra_trees |
12.83 ± 4.26 | 12.30 ± 3.87 | 12.67 ± 4.08 | 10.81 ± 3.41 | 9.57 ± 3.35 |
xgboost |
14.94 ± 4.67 | 14.54 ± 4.55 | 14.06 ± 4.23 | 11.37 ± 3.20 | 10.18 ± 3.22 |
random_forest |
14.68 ± 4.37 | 14.63 ± 4.08 | 14.99 ± 4.48 | 12.01 ± 3.30 | 10.39 ± 2.94 |
lightgbm |
14.89 ± 4.54 | 13.89 ± 3.94 | 15.24 ± 4.33 | 11.53 ± 3.50 | 10.98 ± 3.12 |
ridge |
14.13 ± 2.77 | 13.45 ± 2.77 | 15.69 ± 4.65 | 462.65 ± 1384.36 | 18.35 ± 9.26 |
lasso |
13.91 ± 2.68 | 13.20 ± 2.51 | 17.04 ± 5.89 | 322.25 ± 1494.57 | 22.22 ± 17.92 |
elastic_net |
13.73 ± 2.64 | 13.22 ± 2.55 | 16.45 ± 5.18 | 335.73 ± 1218.39 | 23.58 ± 25.54 |
huber |
13.69 ± 2.96 | 11.84 ± 2.44 | 16.43 ± 7.55 | 1601.84 ± 4476.87 | 26.10 ± 40.86 |
dummy_mean |
36.95 ± 7.89 | 36.95 ± 7.89 | 36.95 ± 7.89 | 36.95 ± 7.89 | 36.95 ± 7.89 |
dummy_median |
37.94 ± 8.01 | 37.94 ± 8.01 | 37.94 ± 8.01 | 37.94 ± 8.01 | 37.94 ± 8.01 |
The linear-model tail failure, and why it disqualifies them here
Read the RMSE and MAPE tables together at N=50. Ridge reports a MAPE of 462.65% and Huber 1601.84% - many times worse than the dummy floor of 36.95%.
Per-fold inspection shows this is not a poor average but a catastrophic tail: ridge's median fold at N=50 is a perfectly reasonable 13.95%, yet 9 of 50 folds exceed 100%, the worst reaching 7136%. Extra trees over the same folds never exceeds 19.93% and has no fold above 100%.
The mechanism is extrapolation. A linear model extended beyond its training
range produces an unbounded prediction, and Phase 4 documented exactly the kind
of feature that triggers it - b1c41's dq_kurtosis sits ~140 robust-z from the
cohort median because its DeltaQ(V) contains one localized spike. Tree
ensembles cannot extrapolate past the training range by construction, so they
are structurally immune.
For this project the tail is what matters, not the median. A QC system is accountable for its escape rate, and a model that is excellent 82% of the time and unbounded the rest cannot carry a risk guarantee. This is the concrete reason Phases 7-8 build on tree ensembles.
It also settles the Phase 5 question precisely. The concern raised there was right in this setting - the full 57-feature matrix under cross-validation - and wrong for the curated 9-feature Severson reproduction, where elastic net was the best of four estimators. Both measurements stand; the scope of the claim is what needed correcting.
Spearman rank correlation
Reported because grading depends on order, not absolute accuracy: a model that ranks perfectly but is biased in level still grades perfectly once thresholds are calibrated.
| Model | N=5 | N=10 | N=20 | N=50 | N=100 |
|---|---|---|---|---|---|
extra_trees |
0.8957 ± 0.0400 | 0.8986 ± 0.0415 | 0.8967 ± 0.0385 | 0.9226 ± 0.0380 | 0.9436 ± 0.0251 |
random_forest |
0.8791 ± 0.0475 | 0.8749 ± 0.0443 | 0.8640 ± 0.0582 | 0.9121 ± 0.0399 | 0.9377 ± 0.0275 |
xgboost |
0.8815 ± 0.0427 | 0.8866 ± 0.0500 | 0.8716 ± 0.0647 | 0.9133 ± 0.0344 | 0.9342 ± 0.0361 |
lightgbm |
0.8778 ± 0.0390 | 0.8852 ± 0.0443 | 0.8678 ± 0.0644 | 0.9176 ± 0.0387 | 0.9271 ± 0.0313 |
elastic_net |
0.8863 ± 0.0435 | 0.8918 ± 0.0427 | 0.8796 ± 0.0455 | 0.8878 ± 0.0588 | 0.8824 ± 0.0626 |
huber |
0.8808 ± 0.0491 | 0.9087 ± 0.0335 | 0.8687 ± 0.0509 | 0.9096 ± 0.0358 | 0.8813 ± 0.0514 |
ridge |
0.8855 ± 0.0459 | 0.8940 ± 0.0446 | 0.8795 ± 0.0453 | 0.9042 ± 0.0385 | 0.8760 ± 0.0603 |
lasso |
0.8825 ± 0.0454 | 0.8911 ± 0.0455 | 0.8765 ± 0.0458 | 0.8731 ± 0.0669 | 0.8572 ± 0.0841 |
dummy_mean |
n/a ± n/a | n/a ± n/a | n/a ± n/a | n/a ± n/a | n/a ± n/a |
dummy_median |
n/a ± n/a | n/a ± n/a | n/a ± n/a | n/a ± n/a | n/a ± n/a |
Escape rate
The manufacturing metric, at a fixed reference operating point. This is what a plant is accountable for, and reporting only the statistical family would contradict the thesis of the project.
| Model | N=5 | N=10 | N=20 | N=50 | N=100 |
|---|---|---|---|---|---|
extra_trees |
0.0525 ± 0.0525 | 0.0486 ± 0.0481 | 0.0428 ± 0.0531 | 0.0480 ± 0.0530 | 0.0454 ± 0.0460 |
xgboost |
0.1316 ± 0.0987 | 0.1174 ± 0.0798 | 0.0965 ± 0.0762 | 0.0562 ± 0.0629 | 0.0514 ± 0.0563 |
random_forest |
0.1060 ± 0.0709 | 0.1016 ± 0.0708 | 0.0971 ± 0.0742 | 0.0849 ± 0.0774 | 0.0567 ± 0.0554 |
lightgbm |
0.1367 ± 0.0885 | 0.1046 ± 0.0676 | 0.1153 ± 0.0830 | 0.0658 ± 0.0682 | 0.0646 ± 0.0656 |
huber |
0.0781 ± 0.0743 | 0.0875 ± 0.0651 | 0.0913 ± 0.0774 | 0.0850 ± 0.0664 | 0.1101 ± 0.0791 |
lasso |
0.1066 ± 0.0798 | 0.1306 ± 0.0831 | 0.1681 ± 0.1200 | 0.2698 ± 0.1249 | 0.2170 ± 0.1102 |
ridge |
0.1161 ± 0.0750 | 0.1419 ± 0.0832 | 0.1465 ± 0.0846 | 0.1800 ± 0.0853 | 0.2288 ± 0.0963 |
elastic_net |
0.1057 ± 0.0764 | 0.1308 ± 0.0829 | 0.1668 ± 0.1179 | 0.2407 ± 0.1105 | 0.2498 ± 0.1164 |
dummy_median |
0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 |
dummy_mean |
0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 | 0.3466 ± 0.0888 |
4. Grading route: ordinal regression vs direct classification
The reviewed decision after Phase 3: grades come from thresholding a predicted cycle life, not from a three-class classifier. Phase 3 measured the realised balance as A = 11 (8.9%), B = 70 (56.5%), C = 43 (34.7%) - roughly two grade-A cells per fold. A direct classifier would estimate a boundary for a class it sees twice per fold.
Ordinal treatment avoids that structurally: every cell informs one continuous target regardless of which side of a boundary it falls on. It also preserves what Phase 8 needs - the CONTINUE action requires a continuous predictive distribution to compute value of information, which a three-class posterior cannot supply.
| Route | Metric | mean ± std | 95% CI | Caveat |
|---|---|---|---|---|
PRIMARY ordinal_from_regression |
escape_rate | 0.0789 ± 0.0596 | [0.0451, 0.1142] | |
PRIMARY ordinal_from_regression |
overkill_rate | 0.0420 ± 0.0381 | [0.0192, 0.0638] | |
PRIMARY ordinal_from_regression |
yield | 0.6737 ± 0.0884 | [0.6270, 0.7297] | |
PRIMARY ordinal_from_regression |
recall_A | 0.9167 ± 0.1800 | [0.8000, 1.0000] | small-class: grade A has only 11 cells |
PRIMARY ordinal_from_regression |
recall_B | 0.9462 ± 0.0418 | [0.9223, 0.9705] | |
PRIMARY ordinal_from_regression |
recall_C | 0.9033 ± 0.1102 | [0.8350, 0.9650] | |
secondary direct_classification |
escape_rate | 0.0431 ± 0.0512 | [0.0139, 0.0745] | |
secondary direct_classification |
overkill_rate | 0.0675 ± 0.0700 | [0.0303, 0.1127] | |
secondary direct_classification |
yield | 0.6455 ± 0.1071 | [0.5840, 0.7090] | |
secondary direct_classification |
recall_A | 0.8333 ± 0.3600 | [0.6000, 1.0000] | small-class: grade A has only 11 cells |
secondary direct_classification |
recall_B | 0.9337 ± 0.0285 | [0.9185, 0.9524] | |
secondary direct_classification |
recall_C | 0.9481 ± 0.0730 | [0.9025, 0.9889] |
4a. Stacked ensemble
Out-of-fold predictions only, ridge meta-learner over elastic net / random forest / LightGBM (10 outer folds, N=100):
| Metric | mean ± std |
|---|---|
| RMSE (log10) | 0.0623 ± 0.0217 |
| MAPE (raw) | 10.18% ± 3.04 |
| Spearman | 0.9408 ± 0.0283 |
| Escape rate | 0.0536 ± 0.0441 |
Mean meta-learner weights: random forest +0.564, LightGBM +0.455, elastic net +0.146 (std 0.443 - unstable, consistent with the linear tail failure above).
Honest result: the ensemble does not beat the best single model. Extra trees alone reaches RMSE 0.0580 ± 0.0240 at the same budget over 50 folds. With n = 124 and three base models whose errors are correlated, stacking has little disagreement to exploit and the meta-learner has very few rows to fit on. The ensemble is retained for Phase 8 comparison, not promoted as the headline.
5. Limitations
- Compact hyperparameter grids. The full protocol is 50 outer fits per (model, budget), each wrapping an inner grid search. Grids are deliberately small, so a wider search might find better settings for any individual model. The comparison is between models under equal search budgets, which is what Phase 6 needs in order to choose a family to carry forward.
- The reproduction is 1.36x the published value, not an exact match, for the reasons in section 1.
- Published comparison values for the discharge and variance models are recorded from the source report; a reader verifying this work should check them against the paper directly.
← Feature engineering · README · Uncertainty and risk control →