File size: 9,629 Bytes
749bffa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
# Engineering patterns found while building this project

[← Back to README](../README.md)

Patterns that generalise beyond this codebase, recorded because each was found
the expensive way.

---

## Pattern 1 β€” The acceptance check that cannot fail on the defect it targets

Reviewer shorthand for this project: **R6**. Four instances, in four subsystems.

### The shape

An acceptance criterion is written in prose, then implemented as a check on
some *observable consequence* of that criterion. The failure arrives, the check
passes anyway, and it passes because **the check is compatible with the failure
mode** β€” not because the check was sloppy, but because the property it observes
is genuinely present in both the correct and the broken state.

The defining question is not *"is this check correct?"* but:

> **What failure mode is this check structurally incapable of detecting?**

A check that cannot answer that question is not yet an acceptance check.

### It happened four times here

Four independent instances, in four different subsystems, none of them
careless:

---

#### Case 1 β€” "All figures regenerate headlessly"

| | |
|---|---|
| **Criterion** | All figures regenerate headlessly |
| **Check written** | The expected PNG files appear in `outputs/figures/` |
| **Why it passed** | An **interactive** backend also writes the files β€” it just opens a blocking GUI window first |
| **Blind spot** | Everything the criterion was actually about |
| **How it surfaced** | A user closing windows by hand during a run |

The check observed a consequence (files exist) that the failure mode does not
disturb. Worse, the failure is **invisible on the machine that introduces it**
if that machine has a display, and appears only in headless CI.

**Fixed by checking the property**: assert the active backend is non-interactive,
assert a *cold* interpreter importing the plotting module stays headless, assert
no module calls `show()`, assert no module imports `pyplot` outside the one
guarded entry point, and assert figures do not leak.

---

#### Case 2 β€” Gate 2's set-size assertion

| | |
|---|---|
| **Criterion** | Reproduce the published train / primary-test / secondary-test design |
| **Check written** | `assert len(train) == 41 and len(primary) == 43 and len(secondary) == 40` |
| **Why it passed** | Batch 1 contains **exactly 41 cells** by coincidence, so the *wrong* split (train = batch1, test = batch2) produces **identical set sizes** to the correct one (train/primary interleaved over batches 1+2) |
| **Blind spot** | Which *cells* are in each set β€” the only thing that actually differs |
| **How it surfaced** | Gate 2 failed at 25.84% against a 9.1% benchmark |

This is the purest instance. The assertion was not weak β€” 41/43/40 is a precise
and non-trivial fact β€” but the numbers were a **coincidence that both the right
and wrong answers satisfy**. A check can be simultaneously specific and
uninformative.

What actually exposed it was not a check but a **diagnostic inconsistency**: the
published primary-test error (9.1%) is *lower* than the secondary-test error
(15.6%), which is only possible if the primary test is in-distribution. The
observed ordering was inverted. **A relationship between numbers caught what an
assertion on the numbers could not.**

**Fixed by checking composition, not cardinality**: assert that train *and*
primary test each span batches 1 and 2, and that secondary is batch 3 alone.

---

#### Case 3 β€” The conformal quantile off-by-one

| | |
|---|---|
| **Criterion** | Conformal intervals achieve their nominal coverage |
| **Check written** | Empirical coverage β‰₯ nominal, across folds |
| **Why it passed** | The bug made intervals **too wide**. Coverage was *higher* than nominal, so the check passed more comfortably than it would have with correct code |
| **Blind spot** | **Any conservative error whatsoever** β€” by construction |
| **How it surfaced** | Testing the quantile definition against a hand-computed order statistic |

The most instructive of the three. `np.quantile(scores, k/n, method="higher")`
maps a quantile onto index positions `[0, n-1]` and returns the **(k+1)-th**
smallest score where the conformal definition requires the **k-th**.

The error direction is what made it invisible: **a coverage check can only fail
on intervals that are too narrow.** It is structurally incapable of failing on
intervals that are too wide. Adding more coverage checks, more folds, or more
Ξ± levels would never have found it.

Fixing it moved nominal-90% coverage from 92.8% β†’ **90.3%** (CQR) and 95.2% β†’
**90.8%** (split conformal) β€” from over-covering to essentially exact β€” and
narrowed intervals by 4.8% and **24.9%**.

**Fixed by testing the definition, not its consequence**: a unit test asserting
the quantile equals a hand-computed order statistic, plus a monotonicity test in
Ξ±.

---

#### Case 4 β€” The audit sheet that displayed z-scores

| | |
|---|---|
| **Criterion** | The per-cell audit sheet explains a decision to a process engineer |
| **Check written** | Sheets generate without error; every driver named; costs consistent with the stated decision; probabilities sum to 1 |
| **Why it passed** | Every one of those things was **true**. The sheet was generated correctly β€” in the wrong units, for the wrong reader |
| **Blind spot** | Whether the content means anything to its audience |
| **How it surfaced** | Reading one |

The model pipeline ends in a `StandardScaler`, so feature values pulled from it
are **standardised**. The sheet dutifully reported *"how much the discharge
curve's shape changed: βˆ’0.8224"* β€” a z-score, presented to the one reader in
the whole system who cannot interpret one.

This is the same family as the others but with a distinct flavour: the previous
three checked a *necessary* condition and treated it as *sufficient*. This one
checked **structural** correctness and treated it as **semantic** correctness.
The artifact was internally consistent, complete, and useless.

**Fixed by quoting the physical value against a cohort reference**: *"this cell:
0.000433 (typical cell: 0.00017 β€” this one is above typical)"*.

**And the residual limitation is stated rather than closed**: no test here can
establish that an explanation is *useful*, because that requires a human reader.
The tests check completeness and internal consistency, which remains necessary
and not sufficient. Case 4 is the one instance in this list where the gap could
not be fully closed by a better check.

---

### What they have in common

| Case | Check observed | Failure mode it could not see |
|---|---|---|
| Figures | Files on disk | A backend that also writes files |
| Gate 2 | Set **sizes** | A different set **composition** with the same sizes |
| Conformal | Coverage β‰₯ nominal | Any error in the **conservative** direction |
| Audit sheet | Structural correctness | Correct content in the **wrong units for the reader** |

In the first three the check verified a *necessary* condition and treated it as
*sufficient*. The fourth is a variant: it verified **structural** correctness and
treated it as **semantic** correctness. In every case the gap between the two is
where the defect lived.

### The practice adopted from this

**Every acceptance check written in this project from Phase 8 onward states, in
its docstring, what failure mode it cannot detect.** Not as documentation β€” as a
forcing function. Writing that sentence requires enumerating the ways the
criterion could be satisfied while the underlying property is false, which is
the analysis that would have caught the first three of these before they shipped. Case 4 shows the limit of the practice: a check can be structurally sound and still miss a semantic failure that only a human reader will see.

Concretely, the questions that produce it:

1. Could a **broken** implementation satisfy this check? Describe one.
2. Is my check on a **necessary** condition or a **sufficient** one?
3. If the error were in the **conservative** direction, would this still fail?
4. Am I checking the **property**, or an **observable consequence** of it?
5. Is there a **relationship between quantities** that would break under the
   failure, even where an assertion on any single quantity would not? (This is
   what caught Case 2.)

---

## Pattern 2 β€” A result that contradicts an earlier measurement is a bug signal

Twice in this project a defect was found not by a test but by an output
**disagreeing with something already measured**:

- **`StabilitySelector` index misalignment** (Phase 4). Selection returned an
  all-*thermal* feature set. Phase 3 had already measured thermal as the
  **weakest** signal group (|ρ| β‰ˆ 0.22) and Ξ”Q(V) variance as the strongest
  (ρ = βˆ’0.90). The contradiction was the only symptom: pandas `Series.corr`
  aligns on index, and the bootstrap subset carried original row labels against
  a fresh `RangeIndex`, scoring `dq_var` at 0.023 against a true 0.908. No error
  was raised and the output looked plausible in isolation.

- **`cycle_measured` duplication after the continuation join** (Phase 3). Gate 1
  fell to RΒ² = 0.7186 with batch 1 collapsing to 0.4249, against an
  already-measured 0.8588. Validation passed, reconciliation was exact, and
  cycle lives were correct β€” only the canary caught it.

**The practice:** when a new number contradicts an established one, treat the
*new* number as the suspect until proven otherwise, and diagnose before
explaining. Both of these were initially tempting to rationalise.

---

[← Back to README](../README.md)