Spaces:
Running on Zero
Running on Zero
| # Engineering patterns found while building this project | |
| [β Back to README](../README.md) | |
| Patterns that generalise beyond this codebase, recorded because each was found | |
| the expensive way. | |
| --- | |
| ## Pattern 1 β The acceptance check that cannot fail on the defect it targets | |
| Reviewer shorthand for this project: **R6**. Four instances, in four subsystems. | |
| ### The shape | |
| An acceptance criterion is written in prose, then implemented as a check on | |
| some *observable consequence* of that criterion. The failure arrives, the check | |
| passes anyway, and it passes because **the check is compatible with the failure | |
| mode** β not because the check was sloppy, but because the property it observes | |
| is genuinely present in both the correct and the broken state. | |
| The defining question is not *"is this check correct?"* but: | |
| > **What failure mode is this check structurally incapable of detecting?** | |
| A check that cannot answer that question is not yet an acceptance check. | |
| ### It happened four times here | |
| Four independent instances, in four different subsystems, none of them | |
| careless: | |
| --- | |
| #### Case 1 β "All figures regenerate headlessly" | |
| | | | | |
| |---|---| | |
| | **Criterion** | All figures regenerate headlessly | | |
| | **Check written** | The expected PNG files appear in `outputs/figures/` | | |
| | **Why it passed** | An **interactive** backend also writes the files β it just opens a blocking GUI window first | | |
| | **Blind spot** | Everything the criterion was actually about | | |
| | **How it surfaced** | A user closing windows by hand during a run | | |
| The check observed a consequence (files exist) that the failure mode does not | |
| disturb. Worse, the failure is **invisible on the machine that introduces it** | |
| if that machine has a display, and appears only in headless CI. | |
| **Fixed by checking the property**: assert the active backend is non-interactive, | |
| assert a *cold* interpreter importing the plotting module stays headless, assert | |
| no module calls `show()`, assert no module imports `pyplot` outside the one | |
| guarded entry point, and assert figures do not leak. | |
| --- | |
| #### Case 2 β Gate 2's set-size assertion | |
| | | | | |
| |---|---| | |
| | **Criterion** | Reproduce the published train / primary-test / secondary-test design | | |
| | **Check written** | `assert len(train) == 41 and len(primary) == 43 and len(secondary) == 40` | | |
| | **Why it passed** | Batch 1 contains **exactly 41 cells** by coincidence, so the *wrong* split (train = batch1, test = batch2) produces **identical set sizes** to the correct one (train/primary interleaved over batches 1+2) | | |
| | **Blind spot** | Which *cells* are in each set β the only thing that actually differs | | |
| | **How it surfaced** | Gate 2 failed at 25.84% against a 9.1% benchmark | | |
| This is the purest instance. The assertion was not weak β 41/43/40 is a precise | |
| and non-trivial fact β but the numbers were a **coincidence that both the right | |
| and wrong answers satisfy**. A check can be simultaneously specific and | |
| uninformative. | |
| What actually exposed it was not a check but a **diagnostic inconsistency**: the | |
| published primary-test error (9.1%) is *lower* than the secondary-test error | |
| (15.6%), which is only possible if the primary test is in-distribution. The | |
| observed ordering was inverted. **A relationship between numbers caught what an | |
| assertion on the numbers could not.** | |
| **Fixed by checking composition, not cardinality**: assert that train *and* | |
| primary test each span batches 1 and 2, and that secondary is batch 3 alone. | |
| --- | |
| #### Case 3 β The conformal quantile off-by-one | |
| | | | | |
| |---|---| | |
| | **Criterion** | Conformal intervals achieve their nominal coverage | | |
| | **Check written** | Empirical coverage β₯ nominal, across folds | | |
| | **Why it passed** | The bug made intervals **too wide**. Coverage was *higher* than nominal, so the check passed more comfortably than it would have with correct code | | |
| | **Blind spot** | **Any conservative error whatsoever** β by construction | | |
| | **How it surfaced** | Testing the quantile definition against a hand-computed order statistic | | |
| The most instructive of the three. `np.quantile(scores, k/n, method="higher")` | |
| maps a quantile onto index positions `[0, n-1]` and returns the **(k+1)-th** | |
| smallest score where the conformal definition requires the **k-th**. | |
| The error direction is what made it invisible: **a coverage check can only fail | |
| on intervals that are too narrow.** It is structurally incapable of failing on | |
| intervals that are too wide. Adding more coverage checks, more folds, or more | |
| Ξ± levels would never have found it. | |
| Fixing it moved nominal-90% coverage from 92.8% β **90.3%** (CQR) and 95.2% β | |
| **90.8%** (split conformal) β from over-covering to essentially exact β and | |
| narrowed intervals by 4.8% and **24.9%**. | |
| **Fixed by testing the definition, not its consequence**: a unit test asserting | |
| the quantile equals a hand-computed order statistic, plus a monotonicity test in | |
| Ξ±. | |
| --- | |
| #### Case 4 β The audit sheet that displayed z-scores | |
| | | | | |
| |---|---| | |
| | **Criterion** | The per-cell audit sheet explains a decision to a process engineer | | |
| | **Check written** | Sheets generate without error; every driver named; costs consistent with the stated decision; probabilities sum to 1 | | |
| | **Why it passed** | Every one of those things was **true**. The sheet was generated correctly β in the wrong units, for the wrong reader | | |
| | **Blind spot** | Whether the content means anything to its audience | | |
| | **How it surfaced** | Reading one | | |
| The model pipeline ends in a `StandardScaler`, so feature values pulled from it | |
| are **standardised**. The sheet dutifully reported *"how much the discharge | |
| curve's shape changed: β0.8224"* β a z-score, presented to the one reader in | |
| the whole system who cannot interpret one. | |
| This is the same family as the others but with a distinct flavour: the previous | |
| three checked a *necessary* condition and treated it as *sufficient*. This one | |
| checked **structural** correctness and treated it as **semantic** correctness. | |
| The artifact was internally consistent, complete, and useless. | |
| **Fixed by quoting the physical value against a cohort reference**: *"this cell: | |
| 0.000433 (typical cell: 0.00017 β this one is above typical)"*. | |
| **And the residual limitation is stated rather than closed**: no test here can | |
| establish that an explanation is *useful*, because that requires a human reader. | |
| The tests check completeness and internal consistency, which remains necessary | |
| and not sufficient. Case 4 is the one instance in this list where the gap could | |
| not be fully closed by a better check. | |
| --- | |
| ### What they have in common | |
| | Case | Check observed | Failure mode it could not see | | |
| |---|---|---| | |
| | Figures | Files on disk | A backend that also writes files | | |
| | Gate 2 | Set **sizes** | A different set **composition** with the same sizes | | |
| | Conformal | Coverage β₯ nominal | Any error in the **conservative** direction | | |
| | Audit sheet | Structural correctness | Correct content in the **wrong units for the reader** | | |
| In the first three the check verified a *necessary* condition and treated it as | |
| *sufficient*. The fourth is a variant: it verified **structural** correctness and | |
| treated it as **semantic** correctness. In every case the gap between the two is | |
| where the defect lived. | |
| ### The practice adopted from this | |
| **Every acceptance check written in this project from Phase 8 onward states, in | |
| its docstring, what failure mode it cannot detect.** Not as documentation β as a | |
| forcing function. Writing that sentence requires enumerating the ways the | |
| criterion could be satisfied while the underlying property is false, which is | |
| the analysis that would have caught the first three of these before they shipped. Case 4 shows the limit of the practice: a check can be structurally sound and still miss a semantic failure that only a human reader will see. | |
| Concretely, the questions that produce it: | |
| 1. Could a **broken** implementation satisfy this check? Describe one. | |
| 2. Is my check on a **necessary** condition or a **sufficient** one? | |
| 3. If the error were in the **conservative** direction, would this still fail? | |
| 4. Am I checking the **property**, or an **observable consequence** of it? | |
| 5. Is there a **relationship between quantities** that would break under the | |
| failure, even where an assertion on any single quantity would not? (This is | |
| what caught Case 2.) | |
| --- | |
| ## Pattern 2 β A result that contradicts an earlier measurement is a bug signal | |
| Twice in this project a defect was found not by a test but by an output | |
| **disagreeing with something already measured**: | |
| - **`StabilitySelector` index misalignment** (Phase 4). Selection returned an | |
| all-*thermal* feature set. Phase 3 had already measured thermal as the | |
| **weakest** signal group (|Ο| β 0.22) and ΞQ(V) variance as the strongest | |
| (Ο = β0.90). The contradiction was the only symptom: pandas `Series.corr` | |
| aligns on index, and the bootstrap subset carried original row labels against | |
| a fresh `RangeIndex`, scoring `dq_var` at 0.023 against a true 0.908. No error | |
| was raised and the output looked plausible in isolation. | |
| - **`cycle_measured` duplication after the continuation join** (Phase 3). Gate 1 | |
| fell to RΒ² = 0.7186 with batch 1 collapsing to 0.4249, against an | |
| already-measured 0.8588. Validation passed, reconciliation was exact, and | |
| cycle lives were correct β only the canary caught it. | |
| **The practice:** when a new number contradicts an established one, treat the | |
| *new* number as the suspect until proven otherwise, and diagnose before | |
| explaining. Both of these were initially tempting to rationalise. | |
| --- | |
| [β Back to README](../README.md) | |