Spaces:
Running on Zero
Running on Zero
| # 04 - Uncertainty and risk control (RQ2) | |
| [Back to 03: Modelling](03_modeling.md) | [README](../README.md) | |
| Covers **Phase 7**. Every table is generated from | |
| [`outputs/reports/`](../outputs/reports/). | |
| **RQ2: can conformal risk control bound the escape rate at a chosen alpha, and | |
| what does that guarantee cost in yield?** Short answer: yes in-distribution, at | |
| a cost of about 5 percentage points of yield -- and **no under campaign shift**, | |
| which is the more important finding. | |
| --- | |
| ## 1. The baseline being corrected is measured, not hypothetical | |
| Phase 6 recorded the uncalibrated LightGBM quantile models achieving | |
| **66.5% empirical coverage on a nominal 90% interval** -- a | |
| **23.5-point shortfall**. Conformalization is fixing that specific gap. | |
| Stating the starting point matters: "our intervals achieve 93% coverage" means | |
| little without knowing they achieved 66.5% before calibration. | |
| ## 2. Coverage, in-distribution and out-of-distribution | |
| Budget 100, 10 outer folds, mean ± std with bootstrap 95% CI. | |
| | Method | Nominal | In-distribution coverage | Batch-3 OOD coverage | OOD gap | | |
| |---|---:|---|---:|---:| | |
| | `cqr` | 99% | 94.8% ± 5.0 [91.6, 97.6] | 75.0% | **-24.0 pts** | | |
| | `cqr` | 95% | 95.2% ± 4.2 [92.7, 97.6] | 47.5% | **-47.5 pts** | | |
| | `cqr` | 90% | 90.3% ± 4.7 [87.6, 93.2] | 60.0% | **-30.0 pts** | | |
| | `split_conformal` | 99% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-29.0 pts** | | |
| | `split_conformal` | 95% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-25.0 pts** | | |
| | `split_conformal` | 90% | 90.8% ± 8.0 [86.0, 95.2] | 57.5% | **-32.5 pts** | | |
| **In-distribution the guarantee holds**, and slightly over-covers -- expected | |
| behaviour for a finite-sample conformal method, which is deliberately | |
| conservative. | |
| ### The OOD result, reported first because it is unfavourable | |
| **Coverage collapses on batch 3**, by 24 to 47 percentage points. At nominal | |
| 90%, split conformal achieves 57.5% and CQR 62.5% -- both *below* the | |
| 66.5% uncalibrated baseline the method exists to improve on. | |
| This is not a bug, and it is the honest answer to "does your guarantee survive | |
| contact with a new manufacturing campaign?" **It does not.** | |
| Every conformal guarantee assumes calibration and test data are **exchangeable**. | |
| Batch 3 was produced later under different conditions, and Phase 3 measured the | |
| shift directly: **KS D = 0.651, p = 1.7e-11**, with per-batch Gate-1 R² falling | |
| 0.901 / 0.665 / 0.647 across campaigns. The exchangeability assumption is | |
| violated by construction, so the guarantee has no force there. | |
| **The operational consequence** is concrete: a conformal escape-rate guarantee | |
| calibrated on one production campaign **cannot be trusted on the next one | |
| without recalibration**. Phase 10 develops the monitoring trigger this implies. | |
| ### An off-by-one that a coverage check alone would not have caught | |
| The conformal quantile was initially computed as | |
| `np.quantile(scores, k/n, method="higher")`. NumPy maps a quantile onto index | |
| positions `[0, n-1]`, so that expression returns the **(k+1)-th** smallest score | |
| where the conformal definition wants the **k-th**: for n=10, alpha=0.2 it | |
| returned 10.0 where the definition gives 9.0. | |
| The error was **conservative** -- intervals came out wider than necessary, so | |
| empirical coverage still met the nominal level and every coverage assertion | |
| passed. That is precisely why it survived: *checking that coverage holds cannot | |
| detect an interval that is too wide.* It was found by testing the quantile | |
| definition directly against a hand-computed order statistic. | |
| Fixing it took nominal-90% coverage from **92.8% to 90.3%** (CQR) and **95.2% to | |
| 90.8%** (split conformal) -- from over-covering to essentially exact -- and | |
| narrowed the intervals by **4.8%** and **24.9%** respectively. Both the | |
| superseded artifacts (`conformal_*.superseded_offbyone.csv`) and the corrected | |
| ones are kept. | |
| This is the third instance in this project of an acceptance check passing while | |
| the underlying property was wrong (after the headless-figure criterion and | |
| Gate 2's first failure), and the same lesson: **check the property, not a proxy | |
| compatible with the failure mode.** | |
| ## 3. Interval width vs budget | |
| A guarantee that costs an uninformatively wide interval is useless to a QC | |
| engineer. Mean width in log10 cycle life at nominal 90%; the parenthesised | |
| figure is the multiplicative band in raw cycle life. | |
| | Method | N=5 | N=10 | N=20 | N=50 | N=100 | | |
| |---|---|---|---|---|---| | |
| | `cqr` | 0.449 (×2.81) | 0.449 (×2.81) | 0.487 (×3.07) | 0.406 (×2.55) | 0.401 (×2.52) | | |
| | `split_conformal` | 0.349 (×2.23) | 0.328 (×2.13) | 0.302 (×2.01) | 0.244 (×1.75) | 0.205 (×1.60) | | |
| Width narrows monotonically with budget, which is the shape RQ1 needs: more | |
| diagnostic cycles buy a tighter statement. | |
| **Split conformal is narrower than CQR here**, which inverts the usual | |
| expectation that adaptive intervals beat constant-width ones. The reason is | |
| sample size: the quantile models are fitted on roughly 64 proper-training cells, | |
| so they are weak, and CQR must then apply a large conformal correction on top. | |
| Adaptivity needs enough data to estimate the quantiles well, and at n = 124 | |
| there is not. | |
| ## 4. Conformal risk control -- the escape-rate guarantee | |
| The commercially meaningful statement, bounding | |
| **P(ship a cell whose true life < warranty target) ≤ alpha**. | |
| | alpha | Attainable? | Escape rate (population) | Yield | Oracle ceiling | Yield cost | | |
| |---:|---|---|---|---|---| | |
| | 0.01 | **NO** (needs n≥99, have ~35) | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 | | |
| | 0.05 | yes | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 | | |
| | 0.10 | yes | 0.0000 ± 0.0000 | 0.601 ± 0.117 | 0.653 | 0.052 ± 0.033 | | |
| ### alpha = 0.01 is NOT ATTAINABLE, and saying so matters | |
| The finite-sample conformal correction needs `ceil((n+1)(1-alpha)) ≤ n`. Below | |
| that, the level clips to 1.0 and the method simply returns the largest | |
| calibration score -- so **every alpha under the threshold produces an identical | |
| interval**. | |
| With ~35 calibration cells per fold, alpha = 0.01 requires at least | |
| 99 and is unreachable. It produces results **numerically identical** to | |
| alpha = 0.05, which the table shows. Reporting them as two distinct operating | |
| points would be misleading, so the attainability flag is carried in the artifact | |
| itself. | |
| This is a hard consequence of n = 124, not a tuning choice. | |
| ### The guarantee is doing real work, not shipping nothing | |
| A zero escape rate is trivially achievable by rejecting everything, so the | |
| number is only meaningful against the **zero-escape yield ceiling**: since every | |
| cell below the warranty target must be held by any zero-escape policy, the | |
| ceiling is the fraction of genuinely fit cells. | |
| At alpha = 0.10 the policy ships **60.1%** against a ceiling of **65.4%** -- a | |
| cost of **5.3 ± 3.3 percentage points of yield**. That is the answer to the | |
| second half of RQ2. | |
| The achieved escape rate is 0.0000 at every alpha: the bound is not merely met | |
| but met with large margin, which is the conservatism of finite-sample conformal | |
| methods showing through. | |
| ### A caveat on the OOD escape number | |
| Batch 3 contains **1 cell below the warranty target out of 40**. A zero escape | |
| rate there is therefore largely a property of the test set rather than evidence | |
| the guarantee transfers -- unlike the coverage collapse above, which is | |
| informative. Batch 2, by contrast, is 41/43 below target. The escape guarantee | |
| is only meaningfully testable where unfit cells exist. | |
| ## 5. Probability calibration (secondary route) | |
| Reliability of grade probabilities from the **secondary** direct-classification | |
| route. The primary grading route is ordinal thresholding, whose uncertainty is | |
| handled by the conformal methods above. | |
| | Method | ECE | MCE | Accuracy | | |
| |---|---|---|---| | |
| | `isotonic` | 0.0762 ± 0.0162 | 0.4810 ± 0.1100 | 0.9277 ± 0.0524 | | |
| | `sigmoid` | 0.1590 ± 0.0429 | 0.4798 ± 0.0992 | 0.9237 ± 0.0607 | | |
| | `uncalibrated` | 0.0692 ± 0.0374 | 0.2915 ± 0.2151 | 0.9313 ± 0.0430 | | |
| **Reported first because it is unfavourable: post-hoc calibration makes things | |
| worse.** The uncalibrated model has the lowest ECE, and both isotonic and Platt | |
| scaling roughly *double* the maximum calibration error. | |
| The cause is sample size again. Each calibrator is fitted by inner | |
| cross-validation over ~99 training cells, so it sees ~33 cells across three | |
| classes -- with grade A contributing about two. Isotonic regression on that is | |
| badly overfit, and Platt scaling assumes a sigmoid link the data cannot support. | |
| **Conclusion: use the uncalibrated probabilities.** At this sample size, | |
| recalibration costs more than it buys. | |
| ## 6. Does CQR inherit the linear extrapolation tail? | |
| Checked directly, because Phase 6 documented linear models producing unbounded | |
| predictions on 9 of 50 folds. **It does not.** The CQR interval width has a | |
| max/median ratio across folds of **1.24×** (split conformal 1.66×), and there | |
| were **zero quantile crossings** in 248 test predictions. | |
| The reason is structural: the quantile models are LightGBM, and a tree cannot | |
| extrapolate beyond its training range. The tail behaviour that disqualified | |
| linear models does not transfer to the tree-based quantile machinery. | |
| ## 7. Limitations | |
| - **The guarantee is in-distribution only.** Demonstrated to fail under campaign | |
| shift; see section 2. | |
| - **alpha = 0.01 is unreachable** at n = 124; see section 4. | |
| - Calibration and risk control are evaluated over 10 outer folds rather than the | |
| full 50, a compute decision recorded in the artifacts. | |
| - The OOD escape rate is uninformative because batch 3 has almost no unfit cells. | |
| --- | |
| [← Modelling](03_modeling.md) · [README](../README.md) · [Decision framework →](05_decision_framework.md) | |