# 04 - Uncertainty and risk control (RQ2) [Back to 03: Modelling](03_modeling.md) | [README](../README.md) Covers **Phase 7**. Every table is generated from [`outputs/reports/`](../outputs/reports/). **RQ2: can conformal risk control bound the escape rate at a chosen alpha, and what does that guarantee cost in yield?** Short answer: yes in-distribution, at a cost of about 5 percentage points of yield -- and **no under campaign shift**, which is the more important finding. --- ## 1. The baseline being corrected is measured, not hypothetical Phase 6 recorded the uncalibrated LightGBM quantile models achieving **66.5% empirical coverage on a nominal 90% interval** -- a **23.5-point shortfall**. Conformalization is fixing that specific gap. Stating the starting point matters: "our intervals achieve 93% coverage" means little without knowing they achieved 66.5% before calibration. ## 2. Coverage, in-distribution and out-of-distribution Budget 100, 10 outer folds, mean ± std with bootstrap 95% CI. | Method | Nominal | In-distribution coverage | Batch-3 OOD coverage | OOD gap | |---|---:|---|---:|---:| | `cqr` | 99% | 94.8% ± 5.0 [91.6, 97.6] | 75.0% | **-24.0 pts** | | `cqr` | 95% | 95.2% ± 4.2 [92.7, 97.6] | 47.5% | **-47.5 pts** | | `cqr` | 90% | 90.3% ± 4.7 [87.6, 93.2] | 60.0% | **-30.0 pts** | | `split_conformal` | 99% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-29.0 pts** | | `split_conformal` | 95% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-25.0 pts** | | `split_conformal` | 90% | 90.8% ± 8.0 [86.0, 95.2] | 57.5% | **-32.5 pts** | **In-distribution the guarantee holds**, and slightly over-covers -- expected behaviour for a finite-sample conformal method, which is deliberately conservative. ### The OOD result, reported first because it is unfavourable **Coverage collapses on batch 3**, by 24 to 47 percentage points. At nominal 90%, split conformal achieves 57.5% and CQR 62.5% -- both *below* the 66.5% uncalibrated baseline the method exists to improve on. This is not a bug, and it is the honest answer to "does your guarantee survive contact with a new manufacturing campaign?" **It does not.** Every conformal guarantee assumes calibration and test data are **exchangeable**. Batch 3 was produced later under different conditions, and Phase 3 measured the shift directly: **KS D = 0.651, p = 1.7e-11**, with per-batch Gate-1 R² falling 0.901 / 0.665 / 0.647 across campaigns. The exchangeability assumption is violated by construction, so the guarantee has no force there. **The operational consequence** is concrete: a conformal escape-rate guarantee calibrated on one production campaign **cannot be trusted on the next one without recalibration**. Phase 10 develops the monitoring trigger this implies. ### An off-by-one that a coverage check alone would not have caught The conformal quantile was initially computed as `np.quantile(scores, k/n, method="higher")`. NumPy maps a quantile onto index positions `[0, n-1]`, so that expression returns the **(k+1)-th** smallest score where the conformal definition wants the **k-th**: for n=10, alpha=0.2 it returned 10.0 where the definition gives 9.0. The error was **conservative** -- intervals came out wider than necessary, so empirical coverage still met the nominal level and every coverage assertion passed. That is precisely why it survived: *checking that coverage holds cannot detect an interval that is too wide.* It was found by testing the quantile definition directly against a hand-computed order statistic. Fixing it took nominal-90% coverage from **92.8% to 90.3%** (CQR) and **95.2% to 90.8%** (split conformal) -- from over-covering to essentially exact -- and narrowed the intervals by **4.8%** and **24.9%** respectively. Both the superseded artifacts (`conformal_*.superseded_offbyone.csv`) and the corrected ones are kept. This is the third instance in this project of an acceptance check passing while the underlying property was wrong (after the headless-figure criterion and Gate 2's first failure), and the same lesson: **check the property, not a proxy compatible with the failure mode.** ## 3. Interval width vs budget A guarantee that costs an uninformatively wide interval is useless to a QC engineer. Mean width in log10 cycle life at nominal 90%; the parenthesised figure is the multiplicative band in raw cycle life. | Method | N=5 | N=10 | N=20 | N=50 | N=100 | |---|---|---|---|---|---| | `cqr` | 0.449 (×2.81) | 0.449 (×2.81) | 0.487 (×3.07) | 0.406 (×2.55) | 0.401 (×2.52) | | `split_conformal` | 0.349 (×2.23) | 0.328 (×2.13) | 0.302 (×2.01) | 0.244 (×1.75) | 0.205 (×1.60) | Width narrows monotonically with budget, which is the shape RQ1 needs: more diagnostic cycles buy a tighter statement. **Split conformal is narrower than CQR here**, which inverts the usual expectation that adaptive intervals beat constant-width ones. The reason is sample size: the quantile models are fitted on roughly 64 proper-training cells, so they are weak, and CQR must then apply a large conformal correction on top. Adaptivity needs enough data to estimate the quantiles well, and at n = 124 there is not. ## 4. Conformal risk control -- the escape-rate guarantee The commercially meaningful statement, bounding **P(ship a cell whose true life < warranty target) ≤ alpha**. | alpha | Attainable? | Escape rate (population) | Yield | Oracle ceiling | Yield cost | |---:|---|---|---|---|---| | 0.01 | **NO** (needs n≥99, have ~35) | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 | | 0.05 | yes | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 | | 0.10 | yes | 0.0000 ± 0.0000 | 0.601 ± 0.117 | 0.653 | 0.052 ± 0.033 | ### alpha = 0.01 is NOT ATTAINABLE, and saying so matters The finite-sample conformal correction needs `ceil((n+1)(1-alpha)) ≤ n`. Below that, the level clips to 1.0 and the method simply returns the largest calibration score -- so **every alpha under the threshold produces an identical interval**. With ~35 calibration cells per fold, alpha = 0.01 requires at least 99 and is unreachable. It produces results **numerically identical** to alpha = 0.05, which the table shows. Reporting them as two distinct operating points would be misleading, so the attainability flag is carried in the artifact itself. This is a hard consequence of n = 124, not a tuning choice. ### The guarantee is doing real work, not shipping nothing A zero escape rate is trivially achievable by rejecting everything, so the number is only meaningful against the **zero-escape yield ceiling**: since every cell below the warranty target must be held by any zero-escape policy, the ceiling is the fraction of genuinely fit cells. At alpha = 0.10 the policy ships **60.1%** against a ceiling of **65.4%** -- a cost of **5.3 ± 3.3 percentage points of yield**. That is the answer to the second half of RQ2. The achieved escape rate is 0.0000 at every alpha: the bound is not merely met but met with large margin, which is the conservatism of finite-sample conformal methods showing through. ### A caveat on the OOD escape number Batch 3 contains **1 cell below the warranty target out of 40**. A zero escape rate there is therefore largely a property of the test set rather than evidence the guarantee transfers -- unlike the coverage collapse above, which is informative. Batch 2, by contrast, is 41/43 below target. The escape guarantee is only meaningfully testable where unfit cells exist. ## 5. Probability calibration (secondary route) Reliability of grade probabilities from the **secondary** direct-classification route. The primary grading route is ordinal thresholding, whose uncertainty is handled by the conformal methods above. | Method | ECE | MCE | Accuracy | |---|---|---|---| | `isotonic` | 0.0762 ± 0.0162 | 0.4810 ± 0.1100 | 0.9277 ± 0.0524 | | `sigmoid` | 0.1590 ± 0.0429 | 0.4798 ± 0.0992 | 0.9237 ± 0.0607 | | `uncalibrated` | 0.0692 ± 0.0374 | 0.2915 ± 0.2151 | 0.9313 ± 0.0430 | **Reported first because it is unfavourable: post-hoc calibration makes things worse.** The uncalibrated model has the lowest ECE, and both isotonic and Platt scaling roughly *double* the maximum calibration error. The cause is sample size again. Each calibrator is fitted by inner cross-validation over ~99 training cells, so it sees ~33 cells across three classes -- with grade A contributing about two. Isotonic regression on that is badly overfit, and Platt scaling assumes a sigmoid link the data cannot support. **Conclusion: use the uncalibrated probabilities.** At this sample size, recalibration costs more than it buys. ## 6. Does CQR inherit the linear extrapolation tail? Checked directly, because Phase 6 documented linear models producing unbounded predictions on 9 of 50 folds. **It does not.** The CQR interval width has a max/median ratio across folds of **1.24×** (split conformal 1.66×), and there were **zero quantile crossings** in 248 test predictions. The reason is structural: the quantile models are LightGBM, and a tree cannot extrapolate beyond its training range. The tail behaviour that disqualified linear models does not transfer to the tree-based quantile machinery. ## 7. Limitations - **The guarantee is in-distribution only.** Demonstrated to fail under campaign shift; see section 2. - **alpha = 0.01 is unreachable** at n = 124; see section 4. - Calibration and risk control are evaluated over 10 outer folds rather than the full 50, a compute decision recorded in the artifacts. - The OOD escape rate is uninformative because batch 3 has almost no unfit cells. --- [← Modelling](03_modeling.md) · [README](../README.md) · [Decision framework →](05_decision_framework.md)