CellTriage / docs /03_modeling.md
Sarvarbek13's picture
CellTriage QC operator console - inference only, CPU-bound classical ML
749bffa verified
|
Raw
History Blame Contribute Delete
17.3 kB

03 - Modelling and the Severson reproduction

Back to 02: Feature engineering | README

Covers Phase 6, including Gate 2 - the credibility anchor. Every table on this page is generated from a file in outputs/reports/.


1. GATE 2 - reproducing Severson et al. (2019)

Verdict: PASSED. Full model, primary test: 12.42% mean percent error against a published 9.1% - a ratio of 1.36x, inside the 2.0x tolerance.

Everything downstream - conformal risk control, the cost frontier, chamber allocation - assumes this pipeline extracts the same signal the source literature extracted. If that assumption were wrong, the later results would be internally consistent and externally meaningless. Reproducing a known number is the only way to test it.

Model Features Train Primary test Secondary test
variance 1 14.18% (pub 9.8%) 14.74% (pub 14.1%) 11.39% (pub 14.1%)
discharge 6 7.97% (pub 5.6%) 14.94% (pub 7.5%) 9.73% (pub 10.7%)
full 9 6.77% (pub 4.9%) 12.42% (pub 9.1%) 13.99% (pub 15.6%)
dummy_mean (floor) 0 29.38% 34.65% 35.90%
dummy_median (floor) 0 23.98% 29.27% 44.94%

The evaluation design is the published one

Severson et al. used a fixed three-way split: train and "primary test" interleaved over batches 1 and 2, with batch 3 held out entirely as a "secondary test". Our cohort reproduces the published set sizes exactly: 41 / 43 / 40.

Using this project's own repeated grouped CV instead would give a number that is arguably better methodology but not comparable to the published one, and comparability is the entire point of a reproduction. Both are reported: the published design here, the project's CV in section 3.

The gate did its job - it caught a real error

The first attempt read the design as train = batch1, test = batch2, because batch 1 contains exactly 41 cells and the published training set also contains 41. That coincidence is a trap. Gate 2 failed at 25.84%, 2.84x the benchmark.

The failure was diagnostic rather than merely bad:

  • Train and secondary-test errors reproduced well; only the primary test was wrong. A broken feature pipeline would have degraded all three.
  • The error ordering was inverted. Published primary (9.1%) is lower than secondary (15.6%), which is only possible if the primary test is in-distribution. Ours had primary above secondary.
  • The magnitude had a mechanism: batch 2 contains cells down to 150 cycles while batch 1 bottoms out at 533, so training on batch 1 alone forces extrapolation far below the training range - and mean percent error punishes short-lived cells hardest, since predicting 300 cycles for a cell that lasts 150 is a 100% error on its own.

This is precisely what a gate is for: it caught a methodological error before anything was built on top of it. The corrected split is pinned by a regression test.

Honest accounting of the residual gap

The reproduction is 1.36x the published value, not 1.0x. There are two candidate explanations, and they are not equally favourable, so they were distinguished by measurement rather than assumed.

Explanation What it would imply Prediction it makes
(a) Feature-list reconstruction - the multi-feature sets were rebuilt from the paper's prose, not copied from the authors' code The pipeline is sound; the inferred feature lists differ in detail The gap should be largest where reconstruction was hardest and near zero where there was nothing to infer
(b) Feature-pipeline error - something upstream computes the ΔQ(V) quantities differently from the source The pipeline is wrong, and every downstream result is suspect The gap should appear everywhere, including in the simplest model

The variance model settles it. It uses exactly ONE feature, log10|var(ΔQ(V))|, so there is no feature-list ambiguity to reconstruct at all - either the pipeline computes that quantity as the source did, or it does not. It lands at 14.74% against a published 14.1%, a ratio of 1.05x.

The multi-feature sets, whose composition had to be inferred, sit further out (discharge 1.99x, full 1.36x). That is the pattern explanation (a) predicts and the opposite of what (b) predicts: a broken feature pipeline would fail the unambiguous single-feature model first, and it does not.

Gate 1 supports the same conclusion independently: the log10 var(ΔQ(V)) vs log10 cycle-life relationship reproduced at R² = 0.8588 against a published ρ = -0.93 (implied R² ≈ 0.86), which is the same ΔQ(V) machinery measured a different way.

This is stated at length because the favourable explanation is the one a reviewer should be most sceptical of, and asserting it without the discriminating test would be exactly the move that deserves scepticism.

Two further differences that cannot be ruled out and are not claimed to be excluded: the cohort here is independently re-derived (including a cycle-life definition recomputed from the capacity series rather than inherited), and three excluded cells could not be independently justified (see docs/01).

2. Does linear fragility limit the reproduction? No.

Phase 5 raised a concern that heavy-tailed features would make linear models fragile, and suggested robust or tree estimators might be needed. Tested directly on the identical split:

Estimator Train Primary test Secondary test
huber (robust linear) 6.34% 13.35% 16.70%
random_forest 5.43% 14.70% 14.62%
lightgbm 2.11% 17.66% 15.42%

The concern does not hold in this setting. Elastic net is the best of the four on the primary test. LightGBM achieves by far the lowest training error and the worst test error - textbook overfitting with only 41 training cells and 9 features.

The earlier concern was not unfounded, it was scoped differently: it arose with the full 57-feature matrix under cross-validation, where a single extrapolating cell (b1c41) dominated MAPE. On the curated 9-feature set the linear model is well-conditioned. Both statements are true; the recommendation that followed from the first was wrong for this case, and the measurement settles it.

3. Benchmark: model x budget

10 models x 5 budgets, evaluated over 50 outer folds of the repeated grouped scheme (5 folds x 10 repeats). Every value is mean ± std across folds; the full table with bootstrap 95% CIs is in benchmarks.csv.

RMSE on log10 cycle life

Model N=5 N=10 N=20 N=50 N=100
extra_trees 0.0798 ± 0.0271 0.0752 ± 0.0247 0.0776 ± 0.0257 0.0665 ± 0.0232 0.0580 ± 0.0240
xgboost 0.0879 ± 0.0258 0.0861 ± 0.0242 0.0831 ± 0.0245 0.0684 ± 0.0208 0.0603 ± 0.0215
random_forest 0.0860 ± 0.0253 0.0845 ± 0.0228 0.0871 ± 0.0244 0.0704 ± 0.0213 0.0627 ± 0.0214
lightgbm 0.0878 ± 0.0264 0.0836 ± 0.0231 0.0894 ± 0.0254 0.0686 ± 0.0227 0.0665 ± 0.0207
ridge 0.0853 ± 0.0170 0.0808 ± 0.0179 0.0923 ± 0.0247 0.1610 ± 0.1766 0.1016 ± 0.0316
lasso 0.0861 ± 0.0170 0.0823 ± 0.0177 0.1002 ± 0.0301 0.1754 ± 0.1423 0.1164 ± 0.0468
elastic_net 0.0850 ± 0.0169 0.0821 ± 0.0170 0.0973 ± 0.0270 0.1666 ± 0.1555 0.1190 ± 0.0452
huber 0.0870 ± 0.0188 0.0768 ± 0.0190 0.0991 ± 0.0369 0.1786 ± 0.2212 0.1201 ± 0.0654
dummy_mean 0.1868 ± 0.0278 0.1868 ± 0.0278 0.1868 ± 0.0278 0.1868 ± 0.0278 0.1868 ± 0.0278
dummy_median 0.1879 ± 0.0271 0.1879 ± 0.0271 0.1879 ± 0.0271 0.1879 ± 0.0271 0.1879 ± 0.0271

Mean absolute percentage error, raw scale

Model N=5 N=10 N=20 N=50 N=100
extra_trees 12.83 ± 4.26 12.30 ± 3.87 12.67 ± 4.08 10.81 ± 3.41 9.57 ± 3.35
xgboost 14.94 ± 4.67 14.54 ± 4.55 14.06 ± 4.23 11.37 ± 3.20 10.18 ± 3.22
random_forest 14.68 ± 4.37 14.63 ± 4.08 14.99 ± 4.48 12.01 ± 3.30 10.39 ± 2.94
lightgbm 14.89 ± 4.54 13.89 ± 3.94 15.24 ± 4.33 11.53 ± 3.50 10.98 ± 3.12
ridge 14.13 ± 2.77 13.45 ± 2.77 15.69 ± 4.65 462.65 ± 1384.36 18.35 ± 9.26
lasso 13.91 ± 2.68 13.20 ± 2.51 17.04 ± 5.89 322.25 ± 1494.57 22.22 ± 17.92
elastic_net 13.73 ± 2.64 13.22 ± 2.55 16.45 ± 5.18 335.73 ± 1218.39 23.58 ± 25.54
huber 13.69 ± 2.96 11.84 ± 2.44 16.43 ± 7.55 1601.84 ± 4476.87 26.10 ± 40.86
dummy_mean 36.95 ± 7.89 36.95 ± 7.89 36.95 ± 7.89 36.95 ± 7.89 36.95 ± 7.89
dummy_median 37.94 ± 8.01 37.94 ± 8.01 37.94 ± 8.01 37.94 ± 8.01 37.94 ± 8.01

The linear-model tail failure, and why it disqualifies them here

Read the RMSE and MAPE tables together at N=50. Ridge reports a MAPE of 462.65% and Huber 1601.84% - many times worse than the dummy floor of 36.95%.

Per-fold inspection shows this is not a poor average but a catastrophic tail: ridge's median fold at N=50 is a perfectly reasonable 13.95%, yet 9 of 50 folds exceed 100%, the worst reaching 7136%. Extra trees over the same folds never exceeds 19.93% and has no fold above 100%.

The mechanism is extrapolation. A linear model extended beyond its training range produces an unbounded prediction, and Phase 4 documented exactly the kind of feature that triggers it - b1c41's dq_kurtosis sits ~140 robust-z from the cohort median because its DeltaQ(V) contains one localized spike. Tree ensembles cannot extrapolate past the training range by construction, so they are structurally immune.

For this project the tail is what matters, not the median. A QC system is accountable for its escape rate, and a model that is excellent 82% of the time and unbounded the rest cannot carry a risk guarantee. This is the concrete reason Phases 7-8 build on tree ensembles.

It also settles the Phase 5 question precisely. The concern raised there was right in this setting - the full 57-feature matrix under cross-validation - and wrong for the curated 9-feature Severson reproduction, where elastic net was the best of four estimators. Both measurements stand; the scope of the claim is what needed correcting.

Spearman rank correlation

Reported because grading depends on order, not absolute accuracy: a model that ranks perfectly but is biased in level still grades perfectly once thresholds are calibrated.

Model N=5 N=10 N=20 N=50 N=100
extra_trees 0.8957 ± 0.0400 0.8986 ± 0.0415 0.8967 ± 0.0385 0.9226 ± 0.0380 0.9436 ± 0.0251
random_forest 0.8791 ± 0.0475 0.8749 ± 0.0443 0.8640 ± 0.0582 0.9121 ± 0.0399 0.9377 ± 0.0275
xgboost 0.8815 ± 0.0427 0.8866 ± 0.0500 0.8716 ± 0.0647 0.9133 ± 0.0344 0.9342 ± 0.0361
lightgbm 0.8778 ± 0.0390 0.8852 ± 0.0443 0.8678 ± 0.0644 0.9176 ± 0.0387 0.9271 ± 0.0313
elastic_net 0.8863 ± 0.0435 0.8918 ± 0.0427 0.8796 ± 0.0455 0.8878 ± 0.0588 0.8824 ± 0.0626
huber 0.8808 ± 0.0491 0.9087 ± 0.0335 0.8687 ± 0.0509 0.9096 ± 0.0358 0.8813 ± 0.0514
ridge 0.8855 ± 0.0459 0.8940 ± 0.0446 0.8795 ± 0.0453 0.9042 ± 0.0385 0.8760 ± 0.0603
lasso 0.8825 ± 0.0454 0.8911 ± 0.0455 0.8765 ± 0.0458 0.8731 ± 0.0669 0.8572 ± 0.0841
dummy_mean n/a ± n/a n/a ± n/a n/a ± n/a n/a ± n/a n/a ± n/a
dummy_median n/a ± n/a n/a ± n/a n/a ± n/a n/a ± n/a n/a ± n/a

Escape rate

The manufacturing metric, at a fixed reference operating point. This is what a plant is accountable for, and reporting only the statistical family would contradict the thesis of the project.

Model N=5 N=10 N=20 N=50 N=100
extra_trees 0.0525 ± 0.0525 0.0486 ± 0.0481 0.0428 ± 0.0531 0.0480 ± 0.0530 0.0454 ± 0.0460
xgboost 0.1316 ± 0.0987 0.1174 ± 0.0798 0.0965 ± 0.0762 0.0562 ± 0.0629 0.0514 ± 0.0563
random_forest 0.1060 ± 0.0709 0.1016 ± 0.0708 0.0971 ± 0.0742 0.0849 ± 0.0774 0.0567 ± 0.0554
lightgbm 0.1367 ± 0.0885 0.1046 ± 0.0676 0.1153 ± 0.0830 0.0658 ± 0.0682 0.0646 ± 0.0656
huber 0.0781 ± 0.0743 0.0875 ± 0.0651 0.0913 ± 0.0774 0.0850 ± 0.0664 0.1101 ± 0.0791
lasso 0.1066 ± 0.0798 0.1306 ± 0.0831 0.1681 ± 0.1200 0.2698 ± 0.1249 0.2170 ± 0.1102
ridge 0.1161 ± 0.0750 0.1419 ± 0.0832 0.1465 ± 0.0846 0.1800 ± 0.0853 0.2288 ± 0.0963
elastic_net 0.1057 ± 0.0764 0.1308 ± 0.0829 0.1668 ± 0.1179 0.2407 ± 0.1105 0.2498 ± 0.1164
dummy_median 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888
dummy_mean 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888 0.3466 ± 0.0888

4. Grading route: ordinal regression vs direct classification

The reviewed decision after Phase 3: grades come from thresholding a predicted cycle life, not from a three-class classifier. Phase 3 measured the realised balance as A = 11 (8.9%), B = 70 (56.5%), C = 43 (34.7%) - roughly two grade-A cells per fold. A direct classifier would estimate a boundary for a class it sees twice per fold.

Ordinal treatment avoids that structurally: every cell informs one continuous target regardless of which side of a boundary it falls on. It also preserves what Phase 8 needs - the CONTINUE action requires a continuous predictive distribution to compute value of information, which a three-class posterior cannot supply.

Route Metric mean ± std 95% CI Caveat
PRIMARY ordinal_from_regression escape_rate 0.0789 ± 0.0596 [0.0451, 0.1142]
PRIMARY ordinal_from_regression overkill_rate 0.0420 ± 0.0381 [0.0192, 0.0638]
PRIMARY ordinal_from_regression yield 0.6737 ± 0.0884 [0.6270, 0.7297]
PRIMARY ordinal_from_regression recall_A 0.9167 ± 0.1800 [0.8000, 1.0000] small-class: grade A has only 11 cells
PRIMARY ordinal_from_regression recall_B 0.9462 ± 0.0418 [0.9223, 0.9705]
PRIMARY ordinal_from_regression recall_C 0.9033 ± 0.1102 [0.8350, 0.9650]
secondary direct_classification escape_rate 0.0431 ± 0.0512 [0.0139, 0.0745]
secondary direct_classification overkill_rate 0.0675 ± 0.0700 [0.0303, 0.1127]
secondary direct_classification yield 0.6455 ± 0.1071 [0.5840, 0.7090]
secondary direct_classification recall_A 0.8333 ± 0.3600 [0.6000, 1.0000] small-class: grade A has only 11 cells
secondary direct_classification recall_B 0.9337 ± 0.0285 [0.9185, 0.9524]
secondary direct_classification recall_C 0.9481 ± 0.0730 [0.9025, 0.9889]

4a. Stacked ensemble

Out-of-fold predictions only, ridge meta-learner over elastic net / random forest / LightGBM (10 outer folds, N=100):

Metric mean ± std
RMSE (log10) 0.0623 ± 0.0217
MAPE (raw) 10.18% ± 3.04
Spearman 0.9408 ± 0.0283
Escape rate 0.0536 ± 0.0441

Mean meta-learner weights: random forest +0.564, LightGBM +0.455, elastic net +0.146 (std 0.443 - unstable, consistent with the linear tail failure above).

Honest result: the ensemble does not beat the best single model. Extra trees alone reaches RMSE 0.0580 ± 0.0240 at the same budget over 50 folds. With n = 124 and three base models whose errors are correlated, stacking has little disagreement to exploit and the meta-learner has very few rows to fit on. The ensemble is retained for Phase 8 comparison, not promoted as the headline.

5. Limitations

  • Compact hyperparameter grids. The full protocol is 50 outer fits per (model, budget), each wrapping an inner grid search. Grids are deliberately small, so a wider search might find better settings for any individual model. The comparison is between models under equal search budgets, which is what Phase 6 needs in order to choose a family to carry forward.
  • The reproduction is 1.36x the published value, not an exact match, for the reasons in section 1.
  • Published comparison values for the discharge and variance models are recorded from the source report; a reader verifying this work should check them against the paper directly.

← Feature engineering · README · Uncertainty and risk control →