CellTriage / docs /04_uncertainty.md
Sarvarbek13's picture
CellTriage QC operator console - inference only, CPU-bound classical ML
749bffa verified
|
Raw
History Blame Contribute Delete
9.93 kB
# 04 - Uncertainty and risk control (RQ2)
[Back to 03: Modelling](03_modeling.md) | [README](../README.md)
Covers **Phase 7**. Every table is generated from
[`outputs/reports/`](../outputs/reports/).
**RQ2: can conformal risk control bound the escape rate at a chosen alpha, and
what does that guarantee cost in yield?** Short answer: yes in-distribution, at
a cost of about 5 percentage points of yield -- and **no under campaign shift**,
which is the more important finding.
---
## 1. The baseline being corrected is measured, not hypothetical
Phase 6 recorded the uncalibrated LightGBM quantile models achieving
**66.5% empirical coverage on a nominal 90% interval** -- a
**23.5-point shortfall**. Conformalization is fixing that specific gap.
Stating the starting point matters: "our intervals achieve 93% coverage" means
little without knowing they achieved 66.5% before calibration.
## 2. Coverage, in-distribution and out-of-distribution
Budget 100, 10 outer folds, mean ± std with bootstrap 95% CI.
| Method | Nominal | In-distribution coverage | Batch-3 OOD coverage | OOD gap |
|---|---:|---|---:|---:|
| `cqr` | 99% | 94.8% ± 5.0 [91.6, 97.6] | 75.0% | **-24.0 pts** |
| `cqr` | 95% | 95.2% ± 4.2 [92.7, 97.6] | 47.5% | **-47.5 pts** |
| `cqr` | 90% | 90.3% ± 4.7 [87.6, 93.2] | 60.0% | **-30.0 pts** |
| `split_conformal` | 99% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-29.0 pts** |
| `split_conformal` | 95% | 96.8% ± 3.7 [94.4, 98.8] | 70.0% | **-25.0 pts** |
| `split_conformal` | 90% | 90.8% ± 8.0 [86.0, 95.2] | 57.5% | **-32.5 pts** |
**In-distribution the guarantee holds**, and slightly over-covers -- expected
behaviour for a finite-sample conformal method, which is deliberately
conservative.
### The OOD result, reported first because it is unfavourable
**Coverage collapses on batch 3**, by 24 to 47 percentage points. At nominal
90%, split conformal achieves 57.5% and CQR 62.5% -- both *below* the
66.5% uncalibrated baseline the method exists to improve on.
This is not a bug, and it is the honest answer to "does your guarantee survive
contact with a new manufacturing campaign?" **It does not.**
Every conformal guarantee assumes calibration and test data are **exchangeable**.
Batch 3 was produced later under different conditions, and Phase 3 measured the
shift directly: **KS D = 0.651, p = 1.7e-11**, with per-batch Gate-1 R² falling
0.901 / 0.665 / 0.647 across campaigns. The exchangeability assumption is
violated by construction, so the guarantee has no force there.
**The operational consequence** is concrete: a conformal escape-rate guarantee
calibrated on one production campaign **cannot be trusted on the next one
without recalibration**. Phase 10 develops the monitoring trigger this implies.
### An off-by-one that a coverage check alone would not have caught
The conformal quantile was initially computed as
`np.quantile(scores, k/n, method="higher")`. NumPy maps a quantile onto index
positions `[0, n-1]`, so that expression returns the **(k+1)-th** smallest score
where the conformal definition wants the **k-th**: for n=10, alpha=0.2 it
returned 10.0 where the definition gives 9.0.
The error was **conservative** -- intervals came out wider than necessary, so
empirical coverage still met the nominal level and every coverage assertion
passed. That is precisely why it survived: *checking that coverage holds cannot
detect an interval that is too wide.* It was found by testing the quantile
definition directly against a hand-computed order statistic.
Fixing it took nominal-90% coverage from **92.8% to 90.3%** (CQR) and **95.2% to
90.8%** (split conformal) -- from over-covering to essentially exact -- and
narrowed the intervals by **4.8%** and **24.9%** respectively. Both the
superseded artifacts (`conformal_*.superseded_offbyone.csv`) and the corrected
ones are kept.
This is the third instance in this project of an acceptance check passing while
the underlying property was wrong (after the headless-figure criterion and
Gate 2's first failure), and the same lesson: **check the property, not a proxy
compatible with the failure mode.**
## 3. Interval width vs budget
A guarantee that costs an uninformatively wide interval is useless to a QC
engineer. Mean width in log10 cycle life at nominal 90%; the parenthesised
figure is the multiplicative band in raw cycle life.
| Method | N=5 | N=10 | N=20 | N=50 | N=100 |
|---|---|---|---|---|---|
| `cqr` | 0.449 (×2.81) | 0.449 (×2.81) | 0.487 (×3.07) | 0.406 (×2.55) | 0.401 (×2.52) |
| `split_conformal` | 0.349 (×2.23) | 0.328 (×2.13) | 0.302 (×2.01) | 0.244 (×1.75) | 0.205 (×1.60) |
Width narrows monotonically with budget, which is the shape RQ1 needs: more
diagnostic cycles buy a tighter statement.
**Split conformal is narrower than CQR here**, which inverts the usual
expectation that adaptive intervals beat constant-width ones. The reason is
sample size: the quantile models are fitted on roughly 64 proper-training cells,
so they are weak, and CQR must then apply a large conformal correction on top.
Adaptivity needs enough data to estimate the quantiles well, and at n = 124
there is not.
## 4. Conformal risk control -- the escape-rate guarantee
The commercially meaningful statement, bounding
**P(ship a cell whose true life < warranty target) ≤ alpha**.
| alpha | Attainable? | Escape rate (population) | Yield | Oracle ceiling | Yield cost |
|---:|---|---|---|---|---|
| 0.01 | **NO** (needs n≥99, have ~35) | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 |
| 0.05 | yes | 0.0000 ± 0.0000 | 0.564 ± 0.136 | 0.653 | 0.089 ± 0.076 |
| 0.10 | yes | 0.0000 ± 0.0000 | 0.601 ± 0.117 | 0.653 | 0.052 ± 0.033 |
### alpha = 0.01 is NOT ATTAINABLE, and saying so matters
The finite-sample conformal correction needs `ceil((n+1)(1-alpha)) ≤ n`. Below
that, the level clips to 1.0 and the method simply returns the largest
calibration score -- so **every alpha under the threshold produces an identical
interval**.
With ~35 calibration cells per fold, alpha = 0.01 requires at least
99 and is unreachable. It produces results **numerically identical** to
alpha = 0.05, which the table shows. Reporting them as two distinct operating
points would be misleading, so the attainability flag is carried in the artifact
itself.
This is a hard consequence of n = 124, not a tuning choice.
### The guarantee is doing real work, not shipping nothing
A zero escape rate is trivially achievable by rejecting everything, so the
number is only meaningful against the **zero-escape yield ceiling**: since every
cell below the warranty target must be held by any zero-escape policy, the
ceiling is the fraction of genuinely fit cells.
At alpha = 0.10 the policy ships **60.1%** against a ceiling of **65.4%** -- a
cost of **5.3 ± 3.3 percentage points of yield**. That is the answer to the
second half of RQ2.
The achieved escape rate is 0.0000 at every alpha: the bound is not merely met
but met with large margin, which is the conservatism of finite-sample conformal
methods showing through.
### A caveat on the OOD escape number
Batch 3 contains **1 cell below the warranty target out of 40**. A zero escape
rate there is therefore largely a property of the test set rather than evidence
the guarantee transfers -- unlike the coverage collapse above, which is
informative. Batch 2, by contrast, is 41/43 below target. The escape guarantee
is only meaningfully testable where unfit cells exist.
## 5. Probability calibration (secondary route)
Reliability of grade probabilities from the **secondary** direct-classification
route. The primary grading route is ordinal thresholding, whose uncertainty is
handled by the conformal methods above.
| Method | ECE | MCE | Accuracy |
|---|---|---|---|
| `isotonic` | 0.0762 ± 0.0162 | 0.4810 ± 0.1100 | 0.9277 ± 0.0524 |
| `sigmoid` | 0.1590 ± 0.0429 | 0.4798 ± 0.0992 | 0.9237 ± 0.0607 |
| `uncalibrated` | 0.0692 ± 0.0374 | 0.2915 ± 0.2151 | 0.9313 ± 0.0430 |
**Reported first because it is unfavourable: post-hoc calibration makes things
worse.** The uncalibrated model has the lowest ECE, and both isotonic and Platt
scaling roughly *double* the maximum calibration error.
The cause is sample size again. Each calibrator is fitted by inner
cross-validation over ~99 training cells, so it sees ~33 cells across three
classes -- with grade A contributing about two. Isotonic regression on that is
badly overfit, and Platt scaling assumes a sigmoid link the data cannot support.
**Conclusion: use the uncalibrated probabilities.** At this sample size,
recalibration costs more than it buys.
## 6. Does CQR inherit the linear extrapolation tail?
Checked directly, because Phase 6 documented linear models producing unbounded
predictions on 9 of 50 folds. **It does not.** The CQR interval width has a
max/median ratio across folds of **1.24×** (split conformal 1.66×), and there
were **zero quantile crossings** in 248 test predictions.
The reason is structural: the quantile models are LightGBM, and a tree cannot
extrapolate beyond its training range. The tail behaviour that disqualified
linear models does not transfer to the tree-based quantile machinery.
## 7. Limitations
- **The guarantee is in-distribution only.** Demonstrated to fail under campaign
shift; see section 2.
- **alpha = 0.01 is unreachable** at n = 124; see section 4.
- Calibration and risk control are evaluated over 10 outer folds rather than the
full 50, a compute decision recorded in the artifacts.
- The OOD escape rate is uninformative because batch 3 has almost no unfit cells.
---
[← Modelling](03_modeling.md) · [README](../README.md) · [Decision framework →](05_decision_framework.md)