Spaces:
Running on Zero
Running on Zero
File size: 37,913 Bytes
749bffa 3f044cb d92710f 749bffa 3f044cb 749bffa 3f044cb 749bffa 3f044cb 749bffa 3f044cb 749bffa 3f044cb 749bffa d92710f 749bffa fa92df8 749bffa 410d8c4 749bffa 410d8c4 749bffa 410d8c4 749bffa fa92df8 749bffa | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 409 410 411 412 413 414 415 416 417 418 419 420 421 422 423 424 425 426 427 428 429 430 431 432 433 434 435 436 437 438 439 440 441 442 443 444 445 446 447 448 449 450 451 452 453 454 455 456 457 458 459 460 461 462 463 464 465 466 467 468 469 470 471 472 473 474 475 476 477 478 479 480 481 482 483 484 485 486 487 488 489 490 491 492 493 494 495 496 497 498 499 500 501 502 503 504 505 506 507 508 509 510 511 512 513 514 515 516 517 518 519 520 521 522 523 524 525 526 527 528 529 530 531 532 533 534 535 536 537 538 539 540 541 542 543 544 545 546 547 548 549 550 551 552 553 554 555 556 557 558 559 560 561 562 563 564 565 566 567 568 569 570 571 572 573 574 575 576 577 578 579 580 581 582 583 584 585 586 587 588 589 590 591 592 593 594 595 596 597 598 599 600 601 602 603 604 605 606 607 608 609 610 611 612 613 614 615 616 617 618 619 620 621 622 623 624 625 626 627 628 629 630 631 632 633 634 635 636 637 638 639 640 641 642 643 644 645 646 647 648 649 650 651 652 653 654 655 656 657 658 659 660 661 662 663 664 665 666 667 668 669 670 671 672 673 674 675 676 677 678 679 680 681 682 683 684 685 686 687 688 689 690 691 692 693 694 695 696 697 698 699 700 701 702 703 704 705 706 707 708 709 710 711 712 713 714 715 716 717 718 719 720 721 722 723 724 725 726 727 728 729 730 731 732 733 734 735 736 737 738 739 740 741 | ---
title: CellTriage QC Operator Console
emoji: π
colorFrom: gray
colorTo: blue
sdk: gradio
sdk_version: 6.22.0
python_version: 3.12.12
app_file: app/app.py
pinned: false
license: mit
short_description: Cost-optimal, risk-controlled QC for Li-ion cells
---
> **This Space is the operator console for a research project.** The full
> write-up, results and limitations are below; the deployment configuration
> is at the end. Every number here is generated from `outputs/reports/`,
> which ships with this Space so each claim can be traced to its file.
# CellTriage
**Cost-optimal, risk-controlled end-of-line QC for lithium-ion cells** β predict
cycle life from early cycles, attach a distribution-free guarantee, and convert
it into a grade decision under an explicit cost matrix.
[](https://www.python.org/)
[](https://scikit-learn.org/)
[](tests/)
[](configs/models.yaml)
[](LICENSE)
[](https://huggingface.co/spaces/Sarvarbek13/CellTriage)
**βΆ [Try the operator console](https://huggingface.co/spaces/Sarvarbek13/CellTriage)** β screen a cell, price the
diagnostic budget, and watch the conformal guarantee be withdrawn when a cell
from a shifted production campaign is selected.
A cell's true lifetime is knowable only after months of cycling to end-of-life,
but a manufacturer must grade it within days. **A point prediction is not a
decision**: the cost of shipping a weak cell into an automotive pack and the
cost of scrapping a good one differ by more than an order of magnitude, so the
useful output is a graded decision with a bounded escape rate.
---
## Headline results
All figures are **mean Β± std across outer cross-validation folds with a
bootstrap 95% CI**. n = 124 cells. Every number traces to a
file in [`outputs/reports/`](outputs/reports/); none is hand-typed.
| Result | Value | Source |
|---|---|---|
| **Gate 2 β Severson reproduction** | 12.42% vs published 9.1% (**1.36Γ**) | `severson_reproduction.json` |
| Variance-only model, same split | **1.05Γ** published | `severson_reproduction.csv` |
| **Extra trees, budget 100** | MAPE **9.57% Β± 3.35** [8.71, 10.53] | `benchmarks.csv` |
| Rank correlation, budget 100 | Ο = 0.944 Β± 0.025 [0.936, 0.950] | `benchmarks.csv` |
| **Triage vs AQL acceptance sampling** | **14.2Γ cheaper**, 50/50 folds, p = 1.8e-15 | `baseline_fold_scorecards.csv` |
| **Chamber time released** | **99.17%** vs cycling to end-of-life | `budget_frontier.json` |
| Economic knee | **N = 5** cycles (5 budgets indistinguishable) | `budget_frontier.json` |
| **Conformal coverage (CQR, Ξ±=0.1)** | **90.3% Β± 4.7 in-distribution** β **60.0% under campaign shift** | `conformal_summary.csv`, `conformal_ood.csv` |
| Conformal coverage (split, Ξ±=0.1) | 90.8% Β± 8.0 β **57.5%** | same |
**The coverage collapse is in the headline table on purpose.** It is the most
transferable result here, and a QC system whose guarantee fails under process
drift should say so before it says anything else.
---
## The five findings
### 1. The model becomes *confidently wrong* under campaign shift

*Accuracy degrades and the interval narrows at the same time β the two lines a
plant would monitor move in opposite directions.*
| Axis | In-distribution | New campaign | Change |
|---|---:|---:|---:|
| RMSE (logββ) | 0.0640 | 0.1233 | **+92.5%** |
| **Interval width** | **0.1975** | **0.1352** | **-31.5%** |
| **Coverage** | **90.7%** | **42.5%** | **-53.2%** |
Error nearly doubles while the prediction interval shrinks by a third.
**A monitor watching interval width alone would read a failing model as a more
confident one.** For any system selling a bounded escape rate, that is the
failure mode to design against β the confidence signal moves in the
*reassuring* direction exactly as it stops being trustworthy.
The mechanism is the conformal calibration set: residuals come from the
qualification campaigns, so the inherited width reflects an easier distribution
than the one it is applied to. Exchangeability fails, and every conformal
guarantee rests on it.
**Phase 3 predicted this seven phases before it was modelled**, which makes it
evidence rather than anecdote: the exploratory analysis had already measured
batch 3 as distribution-shifted (two-sample **KS D = 0.651,
p = 1.7e-11**) and the Gate-1 relationship weakening across
campaigns. β [docs/07](docs/07_protocol_shift.md)
### 2. Cost-blind uncertainty targeting is worse than random

*Holding the cells the model is least sure about is the worst policy tested.*
| Policy | Cost per cell | vs greedy |
|---|---:|---:|
| **greedy VOI per cycle** | **5.973 Β± 1.965 [5.428, 6.526]** | β |
| random | 6.422 Β± 1.722 [5.964, 6.909] | +0.449, p = 0.32 |
| **confidence-only** | **7.467 Β± 1.490 [7.072, 7.896]** | **+1.495, p = 6.4e-06** |
Every policy holds the same fraction (20.2%) and
incurs identical chamber cost, so this compares only *which* cells were chosen.
**Greedy and confidence-only consume identical uncertainty estimates and differ
only in whether the cost matrix enters the ranking** β which isolates the cost
matrix as the cause rather than leaving it as one candidate among several.
Greedy does **not** significantly beat random at this fold size
(25/50 folds,
p = 0.32), and that is reported as a null rather
than dropped. β [docs/05](docs/05_decision_framework.md)
### 3. A hypothesis passed while its physical interpretation failed

*The attribution is real; the mechanism first attached to it was not.*
H3 predicted that thermal-exposure features would be predictive, and they are.
**The proposed mechanism was wrong.** In a 30 Β°C-controlled chamber, the
variance of a temperature integral comes overwhelmingly from **cycle duration**,
not from temperature β so the feature works as a *time-under-load* proxy, not as
the Arrhenius effect the hypothesis named.
**SHAP reports what the model used, not whether the mechanism is real.** No
attribution method could have caught this; it took an external consistency check
against a measurement made six phases earlier. The audit sheets now carry the
caveat in the text a process engineer actually reads.
β [docs/06](docs/06_explainability.md)
### 4. Recipe descriptors: redundant in-distribution, harmful under shift
| Scheme | RMSE with recipe features | without | Change |
|---|---:|---:|---:|
| in-distribution CV | 0.0640 | 0.0649 | **+1.3%** |
| **campaign holdout** | **0.1233** | **0.0681** | **-44.7%** |
| **campaign holdout β coverage** | **42.5%** | **77.5%** | β |
Phase 9 found cross-validated RMSE essentially unchanged when the recipe
descriptors were removed β the model does not need them. Yet removing them under
campaign shift nearly halves RMSE and recovers coverage substantially.
> **Redundancy means the information exists elsewhere, not that the feature is
> inert.** A redundant feature still carries a learned mapping, and when that
> mapping is campaign-specific it becomes a liability the moment the campaign
> changes.
**The deployed console ships without them.** β [docs/07](docs/07_protocol_shift.md)
### 5. The decision is insensitive to budget where the prediction is not

*Cost is flat across the whole budget range; the knee sits at the cheapest end.*
All 5 budgets tested
(5, 10, 20, 50, 100 cycles) are
**statistically indistinguishable on expected cost** β the rule being the
smallest budget whose bootstrap CI overlaps that of the cheapest. The knee is at
**N = 5**, releasing **99.17%** of chamber
time against cycling to end-of-life
(6.7 vs 801.6
chamber-cycles per cell).
**The caveat that must travel with it:** accuracy *does* improve with more
cycles β RMSE falls steadily. The decision is insensitive because the cost
matrix is dominated by a few expensive misgrades, not because the extra cycles
are uninformative. β [docs/05](docs/05_decision_framework.md)
---
## The physical basis

*The variance of ΞQ(V) between cycles 10 and 100 predicts cycle life before any
meaningful capacity fade is visible β this is why early prediction works at all.*
**RΒ² = 0.859** over n = 124 cells
(Pearson r = -0.927;
Severson et al. report Ο = β0.93). Reproducing this published relationship was
**Gate 1**: if it had not appeared, the parser was wrong and every downstream
number would have been fiction. β [docs/01](docs/01_dataset_and_eda.md)
---
## Method
**What is not novel, stated plainly.** Early cycle-life prediction from
ΞQ(V)-derived features is established β Severson et al. (2019) introduced it,
Attia et al. (2020) built closed-loop protocol optimisation on top, and Weng et
al. (2021) pushed the horizon back to formation. Conformal prediction for
battery prognostics already exists. **This project does not claim to have
discovered early prediction**, and the literature itself identifies
post-manufacture screening as the natural application.
Each contribution below is paired with the prior work it sits on and what
actually differs. **Two are downgraded** after reading that literature.
#### 1. Three-way triage under an asymmetric cost matrix β *downgraded*
*Prior work.* The accept / reject / defer structure is not new. Chow (1970)
established the optimal reject option, and Yao (2010) formalised **three-way
decisions** β positive, negative and boundary regions derived from a loss
function and a threshold pair under Bayesian decision theory. That is precisely
the structure used here.
*What differs.* Not the structure, which is textbook. In the standard
formulation the boundary region is a free deferral; here **deferring has a
physical price and a hard capacity limit** β chamber-cycles, a bottleneck
resource β so "keep testing" competes for slots against every other cell, and
the thresholds come from a finite-sample conformal interval rather than
estimated class probabilities.
#### 2. A conformal escape-rate guarantee, priced in yield
*Prior work.* Split conformal and CQR (Vovk et al. 2005; Romano et al. 2019),
conformal risk control (Angelopoulos et al. 2022). In batteries specifically,
Zhang et al. (2022) produce **calibrated cycle-life prediction intervals** with
quantile regression forests, and conformal inference has been applied to SoH and
RUL estimation. Javanmardi & HΓΌllermeier (2022) conformalize RUL, though on
turbofan rather than battery data.
*What differs.* The guarantee is stated as an **escape rate on a QC decision**
rather than as interval coverage, and is **never reported without the yield it
costs** β a bound quoted without its price is not decision-grade. The
infeasibility is also reported: Ξ± = 0.01 needs 99 calibration cells and this
cohort cannot supply them, so it is stated as unreachable rather than omitted.
#### 3. A diagnostic-budget cost frontier
*Prior work.* Value-of-information analysis (Howard 1966) and VOI-based
inspection planning are long established in reliability engineering, as is
optimal stopping for sequential testing.
*What differs.* The instantiation β **diagnostic cycles as the budget variable**
for end-of-line cell screening β and the empirical result, which is that the
frontier is *flat*: the decision is insensitive to budget precisely where the
prediction is not.
#### 4. Capacity-constrained allocation by VOI per cycle β *downgraded*
*Prior work.* Greedy knapsack allocation against a value-of-information score is
a standard construction, and sequential budget allocation under a testing
constraint is well covered in the reliability literature.
*What differs.* Little, mechanically β and the contribution is not the mechanism
but the **measured negative result**: cost-blind uncertainty targeting performs
*worse than random*, with the cost matrix isolated as the cause by an ablation
that holds the uncertainty estimates fixed.
#### 5. Evaluation under production-campaign shift
*Prior work.* Conformal validity under covariate shift has a strong theoretical
literature β weighted conformal prediction (Tibshirani et al. 2019), adaptive
conformal inference (Gibbs & Candès 2021), and conformal prediction beyond
exchangeability (Barber et al. 2023). For batteries, da Silva et al. (2026)
combine transfer learning with conformal prediction under manufacturing
variability, using an LSTM on a *virtual* dataset.
*What differs.* That work proposes **corrections** for shift. This measures the
**failure mode itself on real production campaigns** and reports its diagnostic
signature β error rising while interval width *falls* β as an operational
monitoring hazard, then wires it into the console as an alarm. **We did not find
prior work reporting the coverage/width divergence as a monitoring failure mode
for battery QC**; that is a statement about our search, not a claim that none
exists.
#### 6. Mechanistic validation of learned attributions
*Prior work.* TreeSHAP gives exact attributions for tree ensembles (Lundberg et
al. 2020), and interpretability for battery cycle-life models is established β
Zhang et al. (2022) use permutation importance and partial dependence.
*What differs.* Not the attribution method. The **falsification protocol around
it**: hypotheses stated in advance with the measurement that would refute each,
including an external consistency check. That is what caught a hypothesis which
passed on the evidence while its physical interpretation failed (finding 3) β
something no attribution method can detect on its own.
#### Where this sits
Contributions 1 and 4 are **integrative rather than methodologically novel**,
and are labelled as such above. The defensible claim is the assembly: a screening
system where the guarantee, the diagnostic budget, the capacity constraint, the
shift behaviour and the attribution audit are evaluated *together* under one
cost model β and where the negative results are reported rather than tuned away.
#### Adjacent work in battery production QC
Machine-learned screening in cell manufacturing is an active area: Stock et al.
(2022) classify cycle life early in production, and Zou et al. (2024) classify
quality and predict lifetime from formation data alone. **Neither attaches a
distribution-free guarantee to the decision, prices the diagnostic budget, or
evaluates under campaign shift** β which is the gap this project targets.
**Pipeline.** Features use only cycles 1..N, enforced structurally and verified
by a mutation test that deliberately shuffles cycles beyond the budget. Extra
trees over **56 usable features** of
57 defined; the deployed console model uses
**48**, dropping the 8
recipe descriptors for the reason in finding 4. Selection and preprocessing are
fitted **inside** CV folds only. Split-conformal and CQR intervals. Repeated
grouped nested cross-validation, grouped by physical cell so no cell appears on
both sides of a split.
---
## The falsification design
**Every phase carries a condition under which the method must fire.** This is
the part that makes the nulls interpretable.
| Check | What it would have falsified | Outcome |
|---|---|---|
| **Gate 1** β ΞQ(V) canary | The parser, and with it every downstream number | Passed at RΒ² = 0.859; **fired once** during development on a duplicate-key join |
| **Gate 2** β Severson reproduction | The whole modelling pipeline | Passed at 1.36Γ; **fired once** on a wrong split composition |
| **b1c41 with/without** | Whether one anomalous cell drove the attributions | Passed |
| **H3 external cross-check** | The *interpretation* of a hypothesis that had already passed | **Fired** β see finding 3 |
| **OOD arm** | The conformal guarantee | **Fired** β see finding 1 |
| **Leakage mutation test** | The budget slice, by shuffling cycles > N | Passed |
*(The two "fired once" entries are development incidents, written up in
[docs/01](docs/01_dataset_and_eda.md) and [docs/03](docs/03_modeling.md); the
current artifacts carry the passing values shown.)*
**Three results in this project are nulls**, and each is interpretable only
because the instrument was shown to work first. That is the difference between
*"we found nothing"* and *"there is nothing to find"* β and it transfers well
beyond batteries.
---
## Engineering findings
**One named pattern, caught five times: the acceptance check must test the
property being claimed, not a proxy compatible with the failure mode.**
| # | The check that passed | The property that was not tested |
|---|---|---|
| 1 | "Figures regenerate" β files appeared on disk | No GUI window opens; the interactive backend was still live |
| 2 | Gate 2 split *sizes* matched the paper | Split *composition* did not β batch 1 has exactly 41 cells, which made a wrong design look right |
| 3 | Conformal coverage met its target | Coverage cannot detect a **conservative** off-by-one; over-wide intervals pass by construction |
| 4 | `python -m app.app` served the console | Spaces runs `python app/app.py` β the two differ by one `sys.path` entry, and the latter crashed |
| 5 | This README's generator ran clean | Three `.get(key, <literal>)` lookups fell through to hard-coded defaults when the report layout differed. **A defensive default in a document generator is how hand-typed numbers sneak back into a generated artifact** β two matched the real value by luck, one did not |
Two more from the console build, both consistency failures rather than crashes:
- **The console committed the exact failure it exists to warn about** β it
showed an accept-green escape risk of 0.0% directly beneath a red
*guarantee-void* warning. A monitoring tool that contradicts its own alarm is
the finding above reproduced inside the interface. The number is now withdrawn
rather than hidden, and a test asserts no accept-colour survives an alarm.
- **`StabilitySelector` index misalignment** demoted the project's strongest
predictor (|Ο| = 0.908) to apparent noise (0.023) β caught only because it
contradicted a measurement made in an earlier phase.
Full write-up: [docs/ENGINEERING_PATTERNS.md](docs/ENGINEERING_PATTERNS.md).
---
## Architecture
**The data flow, which is the part worth knowing:**
```
raw .mat (7.7 GiB, HDF5)
ββ parse + validate + reconcile βββββββββββ per-cell parquet (124 cells)
ββ slice to cycles 1..N, build features β per-budget matrices (100 β 56 usable)
ββ nested grouped CV, in-fold selection β fitted models + conformal quantiles
ββ cost matrix + grade boundaries β triage decisions, budget frontier, allocation
ββ freeze bundle (0.2 MB assets + 6.2 MB models) β operator console
```
Each arrow is a module with a `run_*()` entry point; the whole chain is
`python -m src.pipelines.run_all`. **Nothing downstream of a failed gate runs** β
Gate 1 (the ΞQ(V) canary) and Gate 2 (the Severson reproduction) halt the chain
rather than letting later phases build on a broken parser or an unreproduced
benchmark.
```
CellTriage/
βββ configs/ All assumptions as YAML; no magic numbers in code
βββ src/
β βββ data/ MATLAB v7.3 parsing, validation, cohort assembly
β βββ features/ Budget-sliced feature builder + in-fold selectors
β βββ models/ Baselines, registry, Severson reproduction
β βββ uncertainty/ Split conformal, CQR, conformal risk control
β βββ decision/ Cost model, triage policy, budget + capacity optimisation
β βββ explain/ TreeSHAP, physics validation, per-cell audit sheets
β βββ evaluation/ Nested CV protocol, protocol-shift evaluation
β βββ pipelines/ Orchestration, app bundle, this README
βββ app/ Gradio operator console (inference only)
βββ tests/ 382 tests, including the leakage mutation test
βββ docs/ 01β07 deep dives + engineering patterns
βββ outputs/ reports/ Β· figures/ Β· models/ Β· logs/
```
---
## Reproducibility
```bash
python -m venv .venv && .venv/Scripts/pip install -r requirements.lock.txt
python -m src.pipelines.run_all # regenerates every number and figure
python -m pytest tests/ -q # 382 tests
python app/app.py # operator console at :7860
```
Everything is seeded; each artifact carries a provenance stamp with the config
hash and package versions. `requirements.lock.txt` is the exact resolved
environment β pins are load-bearing, since a scikit-learn minor bump can change
estimator internals and silently move results.
---
## Limitations
**n = 124 cells**, from 140 raw
records (11 excluded with logged reasons, 5
consumed by continuation joins). This is a small dataset for the number of
questions asked of it, which is why every number carries fold variance and why a
single test-set figure appears nowhere.
- **The conformal guarantee does not survive a new production campaign** β
coverage 90.3% β 60.0% (CQR). Recalibration on the new
campaign is required before the escape bound means anything. This is the
headline finding, not a caveat.
- **The recipe/campaign confound cannot be resolved by this dataset.**
12 of 18 recipe families occur in exactly
one batch, covering 48.0% of the cohort, so
recipe holdout is partly a campaign holdout. The ~20Γ difference in coverage
degradation between the two axes is offered as an inference, not a claim the
data can support cleanly.
- **Ξ± = 0.01 is unreachable at this sample size** β the finite-sample conformal
correction needs 99 calibration cells and the cohort cannot supply them
alongside a training set.
- **VOI second-order uncertainty is not propagated.** The allocation ranks by
expected value of information using point estimates of the posterior.
- **Grade A has 11 cells.** Grading uses ordinal regression from predicted cycle
life rather than three-class classification for this reason.
- **The advantage over a well-tuned static threshold is conditional**, appearing
above roughly 6:1 escape:overkill and not below it. Against AQL it is
unconditional.
- **This is a research cycling dataset, not a factory dataset.** Cells were
cycled in a temperature-controlled laboratory at 30 Β°C. The QC framing β each
cell a unit at end-of-line, each charging protocol a process recipe, each
batch a production campaign β is a faithful analogue, not a literal production
log. Costs are relative units; every conclusion is tested across an
escape:overkill sweep from 2:1 to 500:1.
---
## Documentation
| Doc | Contents |
|---|---|
| [01 β Dataset and EDA](docs/01_dataset_and_eda.md) | Parsing, exclusions, Gate 1 |
| [02 β Feature engineering](docs/02_feature_engineering.md) | Budget slicing, the leakage rule |
| [03 β Modeling](docs/03_modeling.md) | Gate 2, estimator comparison |
| [04 β Uncertainty](docs/04_uncertainty.md) | Conformal methods, coverage, the off-by-one |
| [05 β Decision framework](docs/05_decision_framework.md) | Cost matrix, triage, budget frontier, allocation |
| [06 β Explainability](docs/06_explainability.md) | TreeSHAP, H1βH4, audit sheets |
| [07 β Protocol shift](docs/07_protocol_shift.md) | OOD evaluation, the recipe finding |
| [Engineering patterns](docs/ENGINEERING_PATTERNS.md) | The acceptance-check pattern, four cases |
---
## References
**Battery cycle life and production screening**
1. Severson, K. A., Attia, P. M., Jin, N., Perkins, N., Yang, B., Chen, M. H.,
Braatz, R. D. et al. Data-driven prediction of battery cycle life before
capacity degradation. *Nature Energy* **4**, 383β391 (2019).
<https://doi.org/10.1038/s41560-019-0356-8>
2. Attia, P. M., Grover, A., Jin, N., Severson, K. A., Markov, T. M., Liao, Y.-H.
et al. Closed-loop optimization of fast-charging protocols for batteries with
machine learning. *Nature* **578**, 397β402 (2020).
<https://doi.org/10.1038/s41586-020-1994-5>
3. Weng, A., Mohtat, P., Attia, P. M., Sulzer, V., Lee, S., Less, G. &
Stefanopoulou, A. Predicting the impact of formation protocols on battery
lifetime immediately after manufacturing. *Joule* **5**, 2971β2992 (2021).
<https://doi.org/10.1016/j.joule.2021.09.015>
4. Stock, S., Pohlmann, S., GΓΌnter, F. J., Hille, L., Hagemeister, J. &
Reinhart, G. Early quality classification and prediction of battery cycle
life in production using machine learning. *Journal of Energy Storage* **50**,
104144 (2022). <https://doi.org/10.1016/j.est.2022.104144>
5. Zou, J., Gao, Y., Frieges, M. H., BΓΆrner, M. F., Kampker, A. & Li, W. Machine
learning for battery quality classification and lifetime prediction using
formation data. *Energy and AI* **18**, 100451 (2024).
<https://doi.org/10.1016/j.egyai.2024.100451>
6. Zhang, H., Su, Y., Altaf, F., Wik, T. & Gros, S. Interpretable battery cycle
life range prediction using early degradation data at cell level. *IEEE
Transactions on Transportation Electrification* (2022).
<https://doi.org/10.1109/TTE.2022.3226683>
**Conformal prediction, distribution shift and prognostics**
7. Vovk, V., Gammerman, A. & Shafer, G. *Algorithmic Learning in a Random World.*
Springer (2005).
8. Romano, Y., Patterson, E. & Candès, E. J. Conformalized quantile regression.
*Advances in Neural Information Processing Systems* **32** (2019).
arXiv:1905.03222
9. Tibshirani, R. J., Barber, R. F., Candès, E. J. & Ramdas, A. Conformal
prediction under covariate shift. *Advances in Neural Information Processing
Systems* **32** (2019). arXiv:1904.06019
10. Gibbs, I. & Candès, E. J. Adaptive conformal inference under distribution
shift. *Advances in Neural Information Processing Systems* **34** (2021).
11. Barber, R. F., Candès, E. J., Ramdas, A. & Tibshirani, R. J. Conformal
prediction beyond exchangeability. *Annals of Statistics* **51**(2) (2023).
12. Angelopoulos, A. N., Bates, S., Fisch, A., Lei, L. & Schuster, T. Conformal
risk control. *The Twelfth International Conference on Learning
Representations (ICLR)* (2024). Spotlight.
<https://openreview.net/forum?id=33XGfHLtZg> (arXiv:2208.02814)
13. Javanmardi, A. & HΓΌllermeier, E. Conformal prediction intervals for
remaining useful lifetime estimation. arXiv:2212.14612 (2022).
14. da Silva, S. F., Ozkan, M. F., El Idrissi, F. & Canova, M. Conformalized
transfer learning for Li-ion battery state of health forecasting under
manufacturing and usage variability. arXiv:2603.24475 (2026).
**Decision theory, value of information and inspection**
15. Chow, C. K. On optimum recognition error and reject tradeoff. *IEEE
Transactions on Information Theory* **16**, 41β46 (1970).
<https://doi.org/10.1109/TIT.1970.1054406>
16. Yao, Y. Three-way decisions with probabilistic rough sets. *Information
Sciences* **180**, 341β353 (2010).
<https://doi.org/10.1016/j.ins.2009.09.021>
17. Howard, R. A. Information value theory. *IEEE Transactions on Systems
Science and Cybernetics* **2**, 22β26 (1966).
<https://doi.org/10.1109/TSSC.1966.300074>
18. ISO 2859-1:1999. *Sampling procedures for inspection by attributes.*
**Explainability**
19. Lundberg, S. M., Erion, G., Chen, H. et al. From local explanations to global
understanding with explainable AI for trees. *Nature Machine Intelligence*
**2**, 56β67 (2020).
---
## License
MIT β see [LICENSE](LICENSE). Covers the source code only; the cycling data is
distributed by its original authors under their own terms and is not
redistributed here.
<sub>README generated by `src/pipelines/build_readme.py` β every number is read
from `outputs/reports/` at build time, never typed.</sub>
---
# Deployment notes
# CellTriage β operator console
Inference-only front end for the CellTriage QC decision system. It screens a
cell against a cost matrix, states a conformal bound on the escape rate, prices
the diagnostic budget, allocates scarce aging-chamber slots, and explains any
decision in physical units.
**The console is a viewer, not a trainer.** It loads five fitted scikit-learn
pipelines and a 34-cell demo subset. No training code runs here, and no raw
cycling data is shipped β the source corpus is 7.7 GiB of `.mat` files that are
redistributable only from their original source.
---
## Deployment configuration, and why
### The hardware tier is ZeroGPU, and the workload does not need a GPU
This is the part worth stating plainly rather than burying.
**Every model in this project is a CPU-bound classical estimator by design.**
The deployed predictor is an `ExtraTreesRegressor` inside a nested sklearn
`Pipeline`. "No neural networks" is a hard project constraint β
a methodological choice, not a resource one β because feature engineering is the
bottleneck in this problem and TreeSHAP gives a process engineer exact,
auditable attributions. A single prediction takes **~50 ms on CPU** once the
model is cached.
**ZeroGPU is selected because it is the free tier that exists, not because the
workload calls for one.** HuggingFace's `cpu-basic` free tier is no longer
offered for new Spaces; ZeroGPU is the available free option. Nothing in this
repository was changed to justify it: no estimator was swapped for a
GPU-requiring implementation, no dependency on `torch` or any accelerator
library was introduced, and the `spaces` import is a shim, not a compute path.
Choosing a deployment tier is not a reason to weaken a stated methodological
constraint, and the constraint won.
If a CPU tier becomes available again, this Space should move to it. Nothing
would need to change but the setting.
**One thing a careful reader will find in the container, so it is stated here
rather than left to look concealed: `torch` IS installed in this image.** The
ZeroGPU builder appends it to the install command itself β the observed build
step is `pip install -r requirements.txt "torch<=2.11.0" gradio[oauth,mcp]
"uvicorn" "websockets" spaces==0.51.1`. It is not in `requirements.txt`, no
module in this repository imports it, and no model uses it. The claim being
made is that *this project* introduces no GPU dependency, not that the platform
image contains none.
### How ZeroGPU is actually selected
**In the Space's Settings UI, not in this YAML header.** The
`suggested_hardware` key accepts `cpu-basic` through `a100-large`; there is no
ZeroGPU value in that enumeration, so a header cannot request it. The steps are:
1. Create the Space with `sdk: gradio` β **ZeroGPU supports the Gradio SDK
only**, which this console already uses.
2. In *Settings β Hardware*, choose **ZeroGPU (Nvidia H200)**.
3. Leave `python_version: 3.12.12` in the header. **ZeroGPU accepts only Python
3.12.12 or 3.10.13**; any other value fails to schedule. The research
environment was frozen on 3.12.10, so this is a deliberate patch-level
divergence β see *Known divergences* below.
### Why nothing is decorated with `@spaces.GPU`
**Nothing in this console requests a GPU slot, and that is deliberate β after
getting it wrong once on a live deployment.**
The first deployed version decorated the inference path:
```python
@spaces.GPU(duration=15) # β the bug
def _infer(model, X): ...
```
The reasoning was that a ZeroGPU Space should mark its heavy path so it
"provisions correctly". **That reasoning was wrong, and the deployed Space
proved it.** `@spaces.GPU` *reserves* a GPU slot on every call and bills the
declared duration against a daily quota β **5 minutes on a free account**. The
campaign-shift check builds its reference distribution by predicting one
qualification cell at a time: **22 calls Γ 15 s = 330 s**, so a *single page
load* exhausted the entire day's quota. The console rendered
`You have exceeded your ZeroGPU runs limit` and every decision panel failed.
Every local test passed throughout, because the decorator is documented as
effect-free off ZeroGPU. **The failure was only observable on the deployed
Space** β the same class of defect as the `python app/app.py` launch-mode crash.
The fix is to decorate nothing. **Un-decorated code runs normally on a ZeroGPU
Space, on CPU, consuming no quota**, which is exactly right for CPU-bound
scikit-learn inference at ~45 ms per prediction. `@spaces.GPU` is for functions
that genuinely need the accelerator; applying it to CPU work buys nothing and
spends a scarce shared resource to do it.
`spaces` is therefore **not imported** by the console. It remains in
`requirements.txt` only because the ZeroGPU builder installs it regardless. A
test parses the module AST β not its text β and fails if any GPU-reserving
decorator or `spaces` import reappears.
### Files
| File | Role |
|---|---|
| `app/app.py` | Console; `app_file` target |
| `app/panels.py` | Matplotlib decision panels |
| `app/theme.py` | Industrial console theme; decision-colour rule |
| `app/assets/` | Manifest, demo cells, headline results (0.2 MB) |
| `outputs/models/` | Five fitted pipelines (6.2 MB) |
| `app/requirements-space.txt` | Slim runtime set β see note below |
| `src/features/selector.py` | **Required for unpickling**, not optional |
Spaces installs from a root `requirements.txt`. `app/requirements-space.txt`
documents the minimal runtime set β it drops LightGBM, XGBoost, CatBoost,
Optuna, SHAP and MAPIE, none of which the console imports. To deploy the slim
set, copy it over the root file in the Space repo.
---
## Cold start
A spun-down ZeroGPU Space pays cold start on every wake, so it was measured
rather than assumed:
| Stage | Time (3 runs) |
|---|---:|
| Import + bundle load (manifest, demo cells, results) | 4.18 β 4.38 s |
| First prediction (loads + caches one model) | 0.70 β 0.85 s |
| **Total to first rendered decision** | **5.01 β 5.16 s** |
| Subsequent predictions (cached) | 0.052 β 0.068 s |
Measured on a local CPU with `spaces==0.51.1` installed, so the import cost of
the shim is included. **A ZeroGPU container will be slower** β cold storage,
container scheduling and a first-touch filesystem are not reproducible on a
laptop β so treat this as a floor, not a prediction.
Budget was 30 s. Models load **lazily and are cached per budget** β the
constructor touches only the 0.2 MB metadata bundle, so a session pays for the
budgets it actually uses instead of deserialising all five (6.2 MB) up front.
---
## Known divergences from the research environment
Stated rather than smoothed over:
- **Python 3.12.12 here vs 3.12.10 in `requirements.lock.txt`.** Forced by
ZeroGPU's accepted-version list. A patch release, and the console performs no
numerical work beyond `.predict()` β but it means the Space is not
byte-identical to the environment every reported number was produced in. **All
headline results come from the research environment, never from this Space.**
- **pydantic 2.12.5 here vs 2.13.4 in the lock file.** Forced: the builder
installs `gradio[oauth,mcp]`, whose `mcp` extra caps pydantic at 2.12.5, and
the first deploy failed with `ResolutionImpossible` on exactly this. Verified
safe before changing β no pydantic object is pickled inside any shipped model,
so the package is needed only to make `src.utils.config` importable during
unpickling, and that import was confirmed under 2.12.5 in a clean venv.
- **The console is not the evidence.** It renders artifacts built by
`src/pipelines/build_app_artifacts.py`. Reproducing the numbers means running
the pipeline, not reading the demo.
- **34 demo cells, not 124.** Enough to exercise every grade, the flagged-data
case, and both campaigns; not a redistribution of the corpus.
## What this console will not let you do
- **Turn off the campaign-shift check.** The incoming lot is derived from the
cell, never chosen. Under a fired alarm the escape figure is still shown but
rendered as withdrawn β never in the accept colour β because the measured
failure mode is a model that looks *more* confident as it stops being
trustworthy, and an interface that showed green reassurance under that alarm
would commit the exact error this project exists to expose.
- **Read a number without its caveat.** *Method & limitations* carries n = 124,
the coverage collapse from 90.7% to 42.5%, the relative-cost caveat, and the
conditions under which the advantage over a static threshold disappears.
|