thundercode commited on
Commit
72833d9
·
verified ·
1 Parent(s): 496299b

release: add docs/TESTING.md

Browse files
Files changed (1) hide show
  1. docs/TESTING.md +1559 -0
docs/TESTING.md ADDED
@@ -0,0 +1,1559 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SatQuery AI — Testing, Verification and Evidence Integrity
2
+
3
+ **Status of this document.** This is the testing chapter of the public SatQuery AI release. It enumerates
4
+ the actual suites that exist, states what each one pins, and reports the full-suite result **including its
5
+ failures**. Every test file, count, command and code excerpt below was read from the repository. Where a
6
+ value could not be established it is written
7
+ `UNKNOWN — not established from the available evidence`.
8
+
9
+ **Status vocabulary** (see `DOCS_STYLE_GUIDE.md` §2): `IMPLEMENTED` · `VERIFIED` · `MEASURED` ·
10
+ `ATTEMPTED` · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` · `OPEN` · `RESOLVED` · `CLOSED`.
11
+
12
+ **The one rule this chapter follows without exception:** a green suite is not evidence about a property
13
+ nobody wrote an assertion for, and a failure is never reclassified as a pass. The project learned this the
14
+ hard way twice — a guard that pinned a *defect* as the specification
15
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5) and a guard whose assertion was satisfied by an upstream fix so it
16
+ could not see the downstream site (`F-15b`). Both are recorded below rather than smoothed over.
17
+
18
+ ---
19
+
20
+ ## Table of contents
21
+
22
+ - [0. How to read this document](#0-how-to-read-this-document)
23
+ - [1. Testing philosophy: tests pin contracts, not implementation](#1-testing-philosophy-tests-pin-contracts-not-implementation)
24
+ - [2. The suite inventory](#2-the-suite-inventory)
25
+ - [3. The doc-guard tests: keeping documentation and code in sync](#3-the-doc-guard-tests-keeping-documentation-and-code-in-sync)
26
+ - [4. The evidence engine: purity, determinism and the reproducibility property](#4-the-evidence-engine-purity-determinism-and-the-reproducibility-property)
27
+ - [5. The frontend live-wiring suite](#5-the-frontend-live-wiring-suite)
28
+ - [6. The doc/frontend suite](#6-the-docfrontend-suite)
29
+ - [7. The full `tests/unit` result, and the 5–6 environmental failures](#7-the-full-testsunit-result-and-the-56-environmental-failures)
30
+ - [8. The live-validation harness and its integrity discipline](#8-the-live-validation-harness-and-its-integrity-discipline)
31
+ - [9. Verdict-independent evidence: recomputing from raw records](#9-verdict-independent-evidence-recomputing-from-raw-records)
32
+ - [10. How to run the tests](#10-how-to-run-the-tests)
33
+ - [11. What is NOT tested](#11-what-is-not-tested)
34
+ - [12. Open items and evidence index](#12-open-items-and-evidence-index)
35
+
36
+ ---
37
+
38
+ ## 0. How to read this document
39
+
40
+ ### 0.1 The three evidence classes
41
+
42
+ The repository encodes an evidence class on every test via a pytest marker (`pytest.ini`):
43
+
44
+ ```ini
45
+ markers =
46
+ unit: fast, no I/O, no server, no network. Evidence about a component in isolation.
47
+ integration: exercises two or more real components wired together in-process.
48
+ smoke: a minimal end-to-end path run against real local artifacts, not a mock.
49
+ real_inference: ran the real model on real inputs in this environment.
50
+ ```
51
+
52
+ The comment above the markers states the purpose exactly, and it is the reason the markers exist rather
53
+ than being decorative:
54
+
55
+ > *"The markers exist to make a test's EVIDENCE CLASS explicit: a unit test never claims a deployment was
56
+ > exercised, and no test may be reported as 'real inference' unless it ran the real model on real data in
57
+ > this environment."*
58
+
59
+ And one deliberate non-marker, quoted in full because it is a subtle but important choice:
60
+
61
+ > *"`environment_blocked` is not a test marker: a blocked path is reported in the STEP 7 report rather than
62
+ > encoded as a permanently-skipped test, because **a skip can be mistaken for coverage**."*
63
+
64
+ That sentence is the spine of this chapter. A skip that looks like coverage is the failure mode the marker
65
+ design is built to avoid.
66
+
67
+ ### 0.2 What "verified" means here
68
+
69
+ | Term | Means | Example |
70
+ |---|---|---|
71
+ | **Unit** | one component, no I/O, no server, no network | `tests/unit/test_evidence_engine.py` |
72
+ | **Integration** | two or more real components wired in-process | `tests/integration/test_step7_backend_chain.py` |
73
+ | **Smoke** | a minimal end-to-end path over real local artifacts | `tests/unit/phase6_smoke.py` (standalone) |
74
+ | **Real inference** | the real model on real inputs, in this environment | — see §11 |
75
+ | **Live validation** | the deployed stack driven by a browser | §8 |
76
+
77
+ The distinction matters most at the boundary between **integration** and **live**. `docs/API_CONTRACT.md`
78
+ §8 states the boundary plainly: *"no test in this repository dials a network address"*, and it points at
79
+ `docs/ITEM5_INTEGRATION_SUITE_SCOPE.md`, which *"records what the 26-test integration suite proves (the
80
+ app's boundary, in-process) and what only a live deployment can prove (reachability, cold start, memory
81
+ ceilings)."* A green `tests/integration` run is therefore **not** evidence about a deployment.
82
+
83
+ ### 0.3 The two defects that shaped the discipline
84
+
85
+ Two incidents are referenced repeatedly below, because they are why this chapter is written the way it is:
86
+
87
+ 1. **A guard that pinned the defect.** The test
88
+ `tests/unit/test_controller.py::test_trace_records_inputs_and_query` asserted
89
+ `envelope.trace.inputs == list(request.assets)` — which, after the upload handler rewrote handles into
90
+ paths, pinned a **filesystem-path disclosure as the specification**. The lesson, recorded in
91
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §5.3: *"A guard is only as strong as the behaviour it pins, and this
92
+ one pinned the defect — a reminder that a green suite is not evidence about a property nobody wrote an
93
+ assertion for."*
94
+ 2. **A guard that could not fail.** The `to_trace()` seam's scrub had **no covering test**: falsifying it
95
+ alone left the producer-side guard **GREEN**, because the producer had already scrubbed the value the
96
+ guard inspected. Recorded as F-15b in `docs/DEPLOYMENT_ARCHITECTURE.md` §5.5, and remedied by
97
+ `test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor`, which constructs the entry
98
+ **directly** through the constructor.
99
+
100
+ Both are the same shape: *a guard whose assertion is satisfied by something other than the code under
101
+ test.*
102
+
103
+ ---
104
+
105
+ ## 1. Testing philosophy: tests pin contracts, not implementation
106
+
107
+ ### 1.1 The philosophy, as the suites state it
108
+
109
+ The repository's test docstrings state a consistent philosophy, and the consistency is the point. Four
110
+ representative statements, each quoted from a module docstring:
111
+
112
+ **Doc-guards pin examples, not prose** (`tests/unit/test_api_contract_doc.py`):
113
+
114
+ > *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
115
+ > mis-stated latency budget). What they guarantee is that a frontend developer who copies an example
116
+ > verbatim gets a request the server accepts -- which is the failure mode that actually costs time."*
117
+
118
+ **Guards pin the route, not the validity** (`tests/unit/test_frontend_live_wiring.py`):
119
+
120
+ > *"It also pins the `chang` word-boundary defect: `\bchang\b` cannot match 'changed', so the page's own
121
+ > default question ('What changed here?') routed to `vqa` instead of the change path. That is invisible to
122
+ > any test that only checks 'is this a valid enum value' — both branches produce a valid value. The test
123
+ > therefore asserts the ROUTE, not just the validity."*
124
+
125
+ **Tests read shipped source when there is no importable symbol** (`tests/unit/test_frontend_live_wiring.py`):
126
+
127
+ > *"Why these are read from the source text: the frontend is plain ES5 with no build step and no module
128
+ > exports, so there is no importable symbol. Parsing the source is the only way to assert the shipped
129
+ > bytes. The parsing is anchored on named tokens rather than line numbers, so it does not silently pass if
130
+ > the file is restructured."*
131
+
132
+ **A runbook must not assert facts the repository contradicts** (`tests/unit/test_runbook_doc.py`):
133
+
134
+ > *"A runbook is worse than useless if a copied command fails or a quoted number was never measured: the
135
+ > operator concludes the *system* is broken rather than the *document*."*
136
+
137
+ ### 1.2 The three anti-patterns the suites are written to avoid
138
+
139
+ | Anti-pattern | What it looks like | How the suites avoid it |
140
+ |---|---|---|
141
+ | **Pinning the defect** | an assertion that encodes today's buggy behaviour as correct | guards are **falsified** before being trusted: the fix is reverted and the guard is shown to fail (§1.3) |
142
+ | **Vacuous guards** | an assertion satisfied by an upstream layer, so the downstream site is untested | the seam is tested by constructing the object **directly** through its constructor (F-15b) |
143
+ | **Measurement by prose** | a text search that is tripped by documentation *about* a defect | guards walk the **parsed** examples, not the document text (`test_api_contract_doc.py::test_no_contract_example_fabricates_an_artifact_uri`) |
144
+
145
+ The third is worth quoting in full, because it is a rule about what a guard must fail on
146
+ (`tests/unit/test_api_contract_doc.py`):
147
+
148
+ > *"This walks the PARSED examples rather than the document text on purpose. The prose in section 2.4
149
+ > deliberately quotes the old `artifact://run/9f2c.../change_map.png` example when explaining why it was
150
+ > removed, so a text search would be tripped by documentation *about* the removal -- which is measuring
151
+ > the wrong thing. A guard must fail on the defect, not on a description of the defect."*
152
+
153
+ ### 1.3 Falsification is part of the discipline
154
+
155
+ A guard is not trusted until it has been shown to **fail** against the defect it guards. The pattern is
156
+ recorded with measurements throughout the project. The clearest instance is the F-15 four-site
157
+ falsification (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5), where each site was reverted individually and the
158
+ guards re-run:
159
+
160
+ | Site reverted | `test_a_construction_failure_publishes_no_filesystem_path` | `test_to_trace_scrubs_a_detail_that_arrives_from_any_constructor` |
161
+ |---|---|---|
162
+ | `registry.py:597` (`_failure_entry`) | **FAILS** | passes *(correctly — it tests the seam, not the producer)* |
163
+ | `registry.py:309` (`to_trace`) | **passes — blind** | **FAILS** |
164
+
165
+ The "passes — blind" cell is the F-15b finding: the producer guard cannot see the seam, because the
166
+ producer has already scrubbed the value. The two guards are therefore **not duplicates**, and the table is
167
+ the evidence.
168
+
169
+ A second falsification table covers the two controller sites (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.5):
170
+
171
+ | Site reverted | `test_a_failed_step_publishes_no_filesystem_path` — assertion that fires |
172
+ |---|---|
173
+ | `controller.py:801` (`_failure_evidence`) | line 590, `payload["message"]` |
174
+ | `controller.py:969` (`_warnings`) | line 582, `warnings[]` |
175
+
176
+ And a third, for the trace-path fix (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.4):
177
+
178
+ | Step | Result |
179
+ |---|---|
180
+ | Revert only the `PARSE` site, run the module | **1 failed, 55 passed** at `tests/unit/test_controller.py:1153` |
181
+ | Restore byte-exact, verify hash | `0c00ae6f54300668122728dc706344044bc19d48328ceead30323c7147e04397` |
182
+ | Module re-run | **56 passed** |
183
+ | Is the guard vacuous? | **No** — the shape check only discriminates because the fixture writes real GeoTIFFs, and the name-equality check rejects the revert independently |
184
+
185
+ The "restore byte-exact, verify hash" row is the anti-tamper step: after reverting a site to falsify a
186
+ guard, the file is restored and its hash checked, so the falsification itself cannot leave the tree in a
187
+ modified state.
188
+
189
+ ### 1.4 Determinism is asserted, not hoped for
190
+
191
+ Several suites pin **determinism** as a first-class property rather than a nice-to-have:
192
+
193
+ | Suite | Determinism property |
194
+ |---|---|
195
+ | `tests/unit/test_evidence_engine.py` | the same claims produce a byte-identical, identically-ordered collection even when uuids differ between processes (§4) |
196
+ | `tests/unit/test_router_threshold_sweep.py` | validation-only sweep contracts |
197
+ | `tests/unit/test_change_threshold_sweep.py` | validation-only sweep contracts |
198
+ | `tests/unit/test_report_generator.py` | *"pure, deterministic, fixtures only"* |
199
+ | `tests/unit/test_vlm_evaluate.py` | the predeclared acceptance rule |
200
+ | `tests/unit/test_fusion_seed_variance.py` | the predeclared seed-variance pass |
201
+
202
+ The determinism claim that matters most is the evidence engine's, because it is what makes the live
203
+ validation's verdict-recomputation property hold (§9).
204
+
205
+ ---
206
+
207
+ ## 2. The suite inventory
208
+
209
+ ### 2.1 The layout
210
+
211
+ `pytest.ini` sets the roots:
212
+
213
+ ```ini
214
+ [pytest]
215
+ testpaths = tests
216
+ pythonpath = .
217
+ addopts = -q --tb=short
218
+ filterwarnings =
219
+ ignore::DeprecationWarning
220
+ ignore::PendingDeprecationWarning
221
+ ignore::rasterio.errors.NotGeoreferencedWarning
222
+ ```
223
+
224
+ `tests/` contains eight directories and two top-level modules:
225
+
226
+ | Location | Contents | Evidence class |
227
+ |---|---|---|
228
+ | `tests/unit/` | 98 Python files — 94 collected test modules + `__init__.py`, `conftest.py`, `_qa_probe_tokens.py`, `phase6_smoke.py` | `unit` (mostly) |
229
+ | `tests/e2e/` | `test_demo.py`, `test_gate4_e2e.py`, `_demo_probe.py` | end-to-end |
230
+ | `tests/geospatial/` | `test_transform.py` | `unit` |
231
+ | `tests/integration/` | `test_step7_backend_chain.py`, `test_step7_r02_serving.py` | `integration` |
232
+ | `tests/leakage/` | `test_leakage.py` | `unit` |
233
+ | `tests/model/` | `test_vlm_contract.py`, `test_vqa_quality_gate.py` | `unit` / `integration` |
234
+ | `tests/routing/` | `test_router.py` | `unit` |
235
+ | `tests/` | `test_config.py`, `test_schemas.py` | `unit` |
236
+
237
+ `tests/unit` totals **49,673 lines** across its 98 Python files. Two files there are **deliberately not
238
+ collected**:
239
+
240
+ | File | Why it is not collected |
241
+ |---|---|
242
+ | `tests/unit/_qa_probe_tokens.py` | *"QA scratch probe (NOT collected by pytest -- underscore prefix)."* |
243
+ | `tests/unit/phase6_smoke.py` | *"Phase 6 real CPU smoke test -- standalone, NOT collected by pytest."* |
244
+ | `tests/unit/conftest.py` | *"Shared fixtures for the Phase 6 VLM unit tests."* — a fixture module, not a test module |
245
+ | `tests/unit/__init__.py` | package marker |
246
+
247
+ ### 2.2 The full file inventory, with what each covers
248
+
249
+ The descriptions below are the **first line of each module's own docstring** — the module's statement of
250
+ what it pins, not a paraphrase.
251
+
252
+ #### 2.2.1 Gateway, deployment and configuration
253
+
254
+ | File | What it covers (module docstring) |
255
+ |---|---|
256
+ | `test_gateway_app.py` | *"The gateway ASGI layer: its contract, its allowlist, and real HTTP execution."* |
257
+ | `test_gateway_assets.py` | *"The ephemeral asset store: opaque handles, three obligations, real lifetimes."* |
258
+ | `test_gateway_policy.py` | *"The gateway's decisions must be correct, cheap, and never reach the Space on a rejection."* |
259
+ | `test_gateway_responsibilities.py` | *"STEP 7 §C — the eighteen gateway responsibilities, as UNIT tests."* |
260
+ | `test_space_app.py` | *"The Space entrypoint must be cheap, honest, and hash-preserving."* |
261
+ | `test_deployment_adapter.py` | *"The deployment capability adapter: registry-authoritative, contract-shaped."* |
262
+ | `test_deploy_config.py` | *"Tests for the Phase 18 deploy packaging manifest and its validator."* |
263
+ | `test_deploy_requests.py` | *"Phase 18 §79 DEPLOYMENT: cold-start and sequential-request behaviour."* |
264
+ | `test_config.py` | *"Phase 1 tests — configuration registry and frozen contract guards."* |
265
+ | `test_code_revision.py` | *"Tests for `core.code_revision` — provenance defect section 6b."* |
266
+ | `test_packaging_script.py` | *"Regression tests for `scripts/package_kaggle_code.py`."* |
267
+
268
+ #### 2.2.2 Controller, planner, registry and schemas
269
+
270
+ | File | What it covers |
271
+ |---|---|
272
+ | `test_controller.py` | *"Tests for `core.controller` — dispatch, partial failure and assembly (Phase 15)."* |
273
+ | `test_controller_seam.py` | *"The controller's asset-resolution seam — regression tests for a real defect."* |
274
+ | `test_planner.py` | *"Tests for `core.planner` — the deterministic policy planner (Phase 15)."* |
275
+ | `test_registry.py` | *"Tests for `core.registry` — the specialist capability registry (Phase 15)."* |
276
+ | `test_schemas.py` | *"Phase 1 tests — canonical schema contract."* |
277
+ | `test_errors.py` | *"Tests for `core.errors` — the taxonomy, and the F-15 path scrubber."* |
278
+ | `test_evidence_engine.py` | *"Tests for the evidence engine — the aggregator (Phase 13)."* (§4) |
279
+ | `test_report_generator.py` | *"Unit tests for `reports.generator` — pure, deterministic, fixtures only."* |
280
+
281
+ #### 2.2.3 Router
282
+
283
+ | File | What it covers |
284
+ |---|---|
285
+ | `test_router_threshold_sweep.py` | *"Phase 4/13 — contracts for the validation-only ROUTER threshold sweep."* |
286
+ | `tests/routing/test_router.py` | *"Phase 4 tests — the intent router."* |
287
+
288
+ #### 2.2.4 Change detection and change-VQA (the largest subsystem by file count)
289
+
290
+ | File | What it covers |
291
+ |---|---|
292
+ | `test_change.py` | *"Phase 9 tests — change detection model, dataset, and post-processing."* |
293
+ | `test_change_levir_real_layout.py` | *"Phase 9 — the LEVIR-CD loader against the REAL dataset layout."* |
294
+ | `test_change_phase9_decontamination.py` | *"Phase 9 -- pins that stop the change-detection module from being re-contaminated."* |
295
+ | `test_change_specialist.py` | *"Phase 9 — the change specialist serving contract."* |
296
+ | `test_change_threshold_sweep.py` | *"Phase 9 — contracts for the validation-only threshold sweep."* |
297
+ | `test_change_train_script_contract.py` | *"Phase 9 — contracts for `scripts/train_change.py`."* |
298
+ | `test_change_vqa_amp.py` | *"R-02 — the mixed-precision training contract."* |
299
+ | `test_change_vqa_dataset.py` | *"R-02 test areas E-J — the CDVQA dataset layer."* |
300
+ | `test_change_vqa_eval_cli.py` | *"R-02 — the evaluation CLI's empty-split diagnosis."* |
301
+ | `test_change_vqa_features.py` | *"R-02 test areas K-L — frozen change features and the feature caches."* |
302
+ | `test_change_vqa_head.py` | *"R-02 test areas M-Q — batching, the reasoning head, and the metrics."* |
303
+ | `test_change_vqa_integration.py` | *"R-02 test area R — the change-VQA specialist and its planner/registry wiring."* |
304
+ | `test_change_vqa_kaggle_notebook.py` | *"R-02 — the Kaggle notebook's discovery contract."* |
305
+ | `test_change_vqa_prepare_script.py` | *"R-02 — `prepare_record.json` must be the provenance of the whole directory."* |
306
+ | `test_change_vqa_promotion.py` | *"STEP 1 — the promoted R-02 change-VQA head at its canonical serving path."* |
307
+ | `test_change_vqa_smoke.py` | *"R-02 — the local smoke training run, and the artifact it leaves behind."* |
308
+ | `test_change_vqa_vocab.py` | *"R-02 test areas A-D — the CDVQA answer space and question ontology."* |
309
+ | `test_cdvqa_adapter.py` | *"Tests for the CDVQA reader, join and example construction."* |
310
+ | `test_cdvqa_second_overlap.py` | *"Tests for scripts/check_cdvqa_second_overlap.py."* |
311
+
312
+ #### 2.2.5 Optical-SAR
313
+
314
+ | File | What it covers |
315
+ |---|---|
316
+ | `test_optical_sar_croma.py` | *"CROMA wrapper — the interface verified against the real model."* |
317
+ | `test_optical_sar_croma_geometry.py` | *"Phase 11 geometry guards — the four requirements that were *measured but…"* |
318
+ | `test_optical_sar_fusion_head.py` | *"Optical-SAR fusion head — the frozen concatenation and channel dropout."* |
319
+ | `test_optical_sar_prompts.py` | *"Optical-SAR prompts — the narration boundary."* |
320
+ | `test_optical_sar_radiometry.py` | *"CROMA encoder-input radiometry — the DEV-2 ruling, as executable guards."* |
321
+ | `test_optical_sar_sensor_adapter.py` | *"Optical-SAR sensor adapter — the two hard rules, enforced."* |
322
+ | `test_optical_sar_specialist.py` | *"Phase 11/12 — the optical-SAR specialist serving contract."* |
323
+ | `test_app_serving_optical_sar.py` | *"Regression guards for the optical-SAR serving wiring (Phase 14 / Pass 16)."* |
324
+ | `test_eval_fusion_115.py` | *"Tests for the pre-registered 11.5 metric tool."* |
325
+ | `test_fusion_extraction.py` | *"Tests for the Phase 12 frozen-feature extraction pipeline (fixtures only)."* |
326
+ | `test_fusion_seed_variance.py` | *"Tests for the predeclared seed-variance pass (Gate F D-03 / D-05)."* |
327
+ | `test_fusion_training.py` | *"Tests for the Phase 12 fusion-head training loop (fixtures only)."* |
328
+ | `test_reben_adapter.py` | *"Tests for the reBEN v2 -> PairedSample adapter (Phase 12)."* |
329
+ | `test_analyze_reben_labels.py` | *"Tests for the reBEN label analyser."* |
330
+
331
+ #### 2.2.6 Grounding
332
+
333
+ | File | What it covers |
334
+ |---|---|
335
+ | `test_grounding_dataset.py` | *"Tests for the Phase 8 grounding dataset."* |
336
+ | `test_grounding_degenerate_fallback.py` | *"Regression tests for the degenerate fallback-box guard (Phase 8)."* |
337
+ | `test_grounding_feature_extraction.py` | *"Tests for the frozen-feature cache."* |
338
+ | `test_grounding_head_wiring.py` | *"Wiring the trained RemoteCLIP grounding head into production (work order §5)."* |
339
+ | `test_grounding_metrics.py` | *"Tests for the grounding metrics."* |
340
+ | `test_grounding_nms.py` | *"Unit tests for grounding NMS (specialists/grounding/head.py::nms)."* |
341
+ | `test_vrsbench_loader.py` | *"Tests for the VRSBench referring-expression loader."* |
342
+ | `test_vrsbench_lazy_images.py` | *"Tests for the lazy-image mode of the VRSBench loader."* |
343
+
344
+ #### 2.2.7 VLM training and evaluation
345
+
346
+ | File | What it covers |
347
+ |---|---|
348
+ | `test_vlm_artifact.py` | *"Tests for the Phase 6 adapter artifact (`training.vlm.artifact`)."* |
349
+ | `test_vlm_collate.py` | *"Tests for Phase 6 batch assembly (`training.vlm.collate`)."* |
350
+ | `test_vlm_config.py` | *"Tests for the Phase 6 training recipe (`training.vlm.config`)."* |
351
+ | `test_vlm_dataset.py` | *"Tests for the Phase 6 BigEarthNet -> SmolVLM instruction corpus."* |
352
+ | `test_vlm_evaluate.py` | *"Tests for the predeclared acceptance rule (`training.vlm.evaluate`)."* |
353
+ | `test_vlm_evaluate_cache.py` | *"Tests for the resumable per-question answer cache in `training.vlm.evaluate`."* |
354
+ | `test_vlm_evaluate_decision_split.py` | *"Tests for `decide_acceptance(..., decision_split=...)`."* |
355
+ | `test_vlm_formatting.py` | *"Tests for Phase 6 prompt construction and loss masking (`training.vlm.formatting`)."* |
356
+ | `test_vlm_lora.py` | *"Tests for Phase 6 LoRA injection (`training.vlm.lora`)."* |
357
+ | `tests/model/test_vlm_contract.py` | *"Phase 5 tests — SmolVLM contract, prompts, and the VQA specialist."* |
358
+ | `tests/model/test_vqa_quality_gate.py` | *"Integration tests: the F5-5 quality gate blocks the VLM before it runs."* |
359
+
360
+ #### 2.2.8 Datasets, metrics and evaluation
361
+
362
+ | File | What it covers |
363
+ |---|---|
364
+ | `test_bigearthnet_blocks.py` | *"Tests for the T2 spatial-block scene key (Phase 12)."* |
365
+ | `test_bigearthnet_pairing.py` | *"Tests for the BigEarthNet optical<->SAR pairing step (Phase 12)."* |
366
+ | `test_bigearthnet_prep.py` | *"Tests for the BigEarthNet reader and instruction-pair construction."* |
367
+ | `test_prepare_script.py` | *"Tests for the BigEarthNet preparation CLI."* |
368
+ | `test_caption_metrics.py` | *"Caption metrics (ruling R-16): BLEU, ROUGE-L, CIDEr, BERTScore."* |
369
+ | `test_vqa_metrics.py` | *"Tests for the Tier-1 VQA / Change-VQA text metrics (`evaluation.metrics.vqa`)."* |
370
+ | `test_eval_normalize.py` | *"Tests for `evaluation.normalize` — the plan section 63 metric normaliser."* |
371
+ | `test_evaluation_runner.py` | *"STEP 5 — the evaluation runner."* |
372
+ | `test_calibration.py` | *"STEP 2 — the calibration fitter, the artifact, and the consumer wiring."* |
373
+ | `test_quality_gate.py` | *"Tests for the deterministic input-quality gate (finding F5-5)."* |
374
+ | `test_benchmark_adapters.py` | *"STEP 5 — benchmark adapters: contract, registry, and the non-invention guard."* |
375
+ | `test_benchmark_adapters_real.py` | *"The four real-corpus benchmark adapters: correctness against synthetic fixtures."* |
376
+ | `test_public_test_corpus.py` | *"STEP 5 — the immutable public-test corpus (plan section 37, Test Set Firewall)."* |
377
+ | `test_manifest_freeze.py` | *"Tests for the frozen dataset-manifest record (`evaluation/manifest_freeze.json`)."* |
378
+ | `test_prompt_freeze.py` | *"Tests for the frozen prompt-set record (`evaluation/prompt_freeze.json`)."* |
379
+ | `test_run_manifest_population.py` | *"Tests for STEP 4 — runtime population of the run manifest."* |
380
+ | `test_provenance_fixes.py` | *"Tests for the R-02 provenance fixes: sections 6a, 6b and 6c."* |
381
+ | `test_diagnose_feature_cache.py` | *"Tests for `scripts/diagnose_feature_cache.py`'s exit-code contract."* |
382
+
383
+ #### 2.2.9 Serving composition and the app boundary
384
+
385
+ | File | What it covers |
386
+ |---|---|
387
+ | `test_app_serving.py` | *"Regression tests for `app.serving` — the public serving composition root."* |
388
+ | `test_phase6_closure.py` | *"Tests for the Phase 6 closure record."* |
389
+
390
+ #### 2.2.10 Contract-conformance, frontend and documentation guards
391
+
392
+ | File | What it covers |
393
+ |---|---|
394
+ | `test_api_contract_doc.py` | *"The API contract must not drift from the schemas it documents."* |
395
+ | `test_frontend_guide_doc.py` | *"The contract documents must stay consistent with each other and the schemas."* |
396
+ | `test_frontend_live_wiring.py` | *"Regression tests for the frontend <-> orchestrator live wiring."* (§5) |
397
+ | `test_runbook_doc.py` | *"The deployment runbook must not assert facts the repository contradicts."* |
398
+ | `test_deploy_config.py` | *"Tests for the Phase 18 deploy packaging manifest and its validator."* |
399
+ | `test_step7_api_contract_conformance.py` | *"STEP 7 section D -- the API contract, tested against the schema that implements it."* |
400
+ | `test_safe_delete_shim.py` | *"The Windows safe-delete shim: verbatim prefixes, and the shell's return code."* (§7) |
401
+
402
+ #### 2.2.11 The other directories
403
+
404
+ | File | What it covers |
405
+ |---|---|
406
+ | `tests/geospatial/test_transform.py` | *"Phase 2 tests — geospatial transforms and the raster input contract."* |
407
+ | `tests/leakage/test_leakage.py` | *"Phase 3 tests — Gate 1: dataset manifests and scene-level leakage isolation."* |
408
+ | `tests/integration/test_step7_backend_chain.py` | *"STEP 7 sections H, I, M -- the real inference service, driven over HTTP."* |
409
+ | `tests/integration/test_step7_r02_serving.py` | *"STEP 7 section I -- the R-02 head, verified through the real serving path."* |
410
+ | `tests/e2e/test_demo.py` | *"End-to-end tests for the Phase-19 demonstration driver (`demo.run_demo`)."* |
411
+ | `tests/e2e/test_gate4_e2e.py` | *"Gate-4 end-to-end test — the full control tier in its real local state."* |
412
+ | `tests/e2e/_demo_probe.py` | *"Subprocess probe: run the demo in a FRESH interpreter and report which…"* (helper) |
413
+
414
+ ### 2.3 What the inventory says about coverage shape
415
+
416
+ Three observations a reader should draw from the inventory:
417
+
418
+ 1. **The largest single cluster is change detection and change-VQA** (19 files), which matches the
419
+ project's stated engineering priority ordering (`docs/MASTER_ARCHITECTURE_PLAN.md` §2.1: optical-SAR
420
+ first, change second).
421
+ 2. **Every capability has a *serving contract* test, separate from its *model* tests.** For example
422
+ `test_change.py` covers the model/dataset/post-processing while `test_change_specialist.py` covers
423
+ *"the change specialist serving contract"*; the same split exists for optical-SAR
424
+ (`test_optical_sar_fusion_head.py` vs `test_optical_sar_specialist.py`) and grounding. The split is
425
+ deliberate: a model that works in a notebook and a specialist that serves the contract are different
426
+ claims.
427
+ 3. **Documentation is tested as code.** Five of the 94 modules are doc-guards (§3), and they are counted in
428
+ the 183-passed doc/frontend suite (§6) rather than being an afterthought.
429
+
430
+ ### 2.4 The integration suite's documented scope
431
+
432
+ `docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` is the record of what the integration suite proves. The contract
433
+ summarises it (`docs/API_CONTRACT.md` §8):
434
+
435
+ > *"`docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` records what the 26-test integration suite proves (the app's
436
+ > *boundary*, in-process) and what only a live deployment can prove (reachability, cold start, memory
437
+ > ceilings). Read it before treating a green `tests/integration` run as evidence about a deployment —
438
+ > **no test in this repository dials a network address**, including this contract's own `/v1/*`
439
+ > examples."*
440
+
441
+ The exact number of tests **collected** in a single `tests/integration` invocation is
442
+ `UNKNOWN — not established from the available evidence`. The documented figure is 26; the evidence this
443
+ chapter can establish is the two module files and their stated scope.
444
+
445
+ ---
446
+
447
+ ## 3. The doc-guard tests: keeping documentation and code in sync
448
+
449
+ ### 3.1 Why documentation is tested
450
+
451
+ Five modules treat documentation as a **tested artefact**. The rationale is stated in each module's
452
+ docstring, and the common thread is that a documentation defect costs a real debugging session:
453
+
454
+ | Guard | The failure it prevents |
455
+ |---|---|
456
+ | `test_api_contract_doc.py` | a frontend builds against an example that does not validate, and the cause looks like a frontend bug |
457
+ | `test_frontend_guide_doc.py` | a value a frontend must send or read is documented wrong, producing a `422` and a wasted session |
458
+ | `test_runbook_doc.py` | an operator copies a command that fails, or quotes a number that was never measured, and concludes the *system* is broken |
459
+ | `test_deploy_config.py` | the deploy manifest silently moves the frozen `Config.hash` |
460
+ | `test_step7_api_contract_conformance.py` | the contract and the schema that implements it drift apart |
461
+
462
+ `test_api_contract_doc.py` records what the guards actually catch, with a list of four real defects the
463
+ act of writing the contract produced:
464
+
465
+ > *"Writing this contract produced exactly that failure four times over -- `Box` fields were nested instead
466
+ > of flat, `Evidence` used `summary`/`value`/`source` instead of
467
+ > `coordinates`/`score`/`source_specialist`, `CoordinateSystem` values were shorthand rather than the real
468
+ > `normalized_0_1`/`pixel`/`geo`, and `ExecutionTrace.inputs`/`outputs` were dicts rather than lists. Each
469
+ > was caught by validating the examples against the real models, which is what these tests do
470
+ > permanently."*
471
+
472
+ ### 3.2 `test_api_contract_doc.py`
473
+
474
+ **Scope:** validate every fenced ```json block in `docs/API_CONTRACT.md` against the real Pydantic models.
475
+
476
+ | Mechanism | Detail |
477
+ |---|---|
478
+ | Fixture `contract_text` | asserts the document exists, then reads it |
479
+ | Fixture `json_blocks` | regex-extracts every ```json block and parses it; a block that is not valid JSON raises `AssertionError` naming the block index |
480
+ | `_find(blocks, predicate)` | locates an example by shape rather than by position |
481
+
482
+ | Test | What it pins |
483
+ |---|---|
484
+ | `test_every_json_block_in_the_contract_parses` | *"A malformed example is unusable; this fails loudly rather than silently."* |
485
+ | `test_the_result_envelope_example_validates` | the envelope example validates against `ResultEnvelope`; `schema_version == SCHEMA_VERSION`; `confidence.method in ("uncalibrated", "temperature_scaling")` |
486
+ | `test_no_contract_example_fabricates_an_artifact_uri` | no parsed example contains an `artifact://` URI or a filesystem path in `change_map` / `artifact_ref` (walks the **parsed** examples, §1.2) |
487
+ | (further tests) | documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded |
488
+
489
+ The `confidence.method` assertion is the honest-signal guard: it pins that the contract's example cannot
490
+ claim a calibration method the engine does not have. §4 shows the engine side of the same rule.
491
+
492
+ ### 3.3 `test_frontend_guide_doc.py`
493
+
494
+ **Scope:** `docs/API_CONTRACT.md` and `docs/FRONTEND_INTEGRATION.md` must agree with each other **and** with
495
+ `core/schemas.py`.
496
+
497
+ It imports the real models — `Box, CoordinateSystem, Evidence, EvidenceType, Region, Task` — and compares
498
+ the prose against them. The docstring lists the five defects that motivated it:
499
+
500
+ > *"Writing them produced, in order: `CoordinateSystem` documented as `normalized`/`geographic` (real:
501
+ > `normalized_0_1`/`geo`), `Box` geometry documented as nested (real: flat), `Evidence` fields named
502
+ > `summary`/`value`/`source` (real: `coordinates`/`score`/`source_specialist`),
503
+ > `ExecutionTrace.inputs`/`outputs` as objects (real: lists), and `Task.unsupported` missing from both
504
+ > tables."*
505
+
506
+ The representative test asserts a **bidirectional** requirement:
507
+
508
+ ```python
509
+ def test_every_task_is_documented_in_both_documents(api_text, frontend_text):
510
+ for task in Task:
511
+ assert f"`{task.value}`" in api_text, f"Task.{task.value} missing from API contract"
512
+ assert f"`{task.value}`" in frontend_text, (
513
+ f"Task.{task.value} missing from the frontend guide; a value the client "
514
+ f"may receive but cannot find documented is a crash waiting to happen"
515
+ )
516
+ ```
517
+
518
+ The message on the second assertion is the rationale: *"a value the client may receive but cannot find
519
+ documented is a crash waiting to happen."*
520
+
521
+ The suite also pins: no login screen, no secrets in the browser, the confidence-property trap, and
522
+ serialized analyses (`docs/EVALUATION.md` §7.2).
523
+
524
+ ### 3.4 `test_runbook_doc.py`
525
+
526
+ **Scope:** `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` must not assert facts the repository contradicts.
527
+
528
+ The module names **two failure modes**:
529
+
530
+ > 1. *"**Invented measurements.** The runbook quotes command output. If it quotes a byte size, a
531
+ > capability list, a function name or an exit code that the repository does not produce, the quote is
532
+ > fabricated. Where a fact is cheap to verify locally, this module verifies it."*
533
+ > 2. *"**Invented API surface.** The runbook calls into the codebase (`build_serving_registry`,
534
+ > `available()`, `CHANGE_CHECKPOINT`, `CHANGE_VQA_HEAD`). A runbook that names a function the code does
535
+ > not export is a runbook whose commands raise `AttributeError` on the operator's first copy-paste."*
536
+
537
+ It reads three documents — the runbook, `docs/DEPLOYMENT_ARCHITECTURE.md` and `docs/API_CONTRACT.md` — so
538
+ that a claim contradicted *across* documents is caught, not only a claim contradicted by code.
539
+
540
+ **The most important class of assertion in this module** is stated in its docstring:
541
+
542
+ > *"The absence of a deployment cannot be tested here, so these tests assert the runbook's *local* claims
543
+ > and, separately, assert that the document itself labels its unimplemented sections as unimplemented. The
544
+ > second class matters more than it looks: a runbook that describes a live deployment without having
545
+ > performed one is the exact overclaim the work order forbids."*
546
+
547
+ So the guard pins **honesty about status**, not merely factual accuracy: an "undecided" item must be
548
+ declared undecided, and a "designed" item must not be described as "verified" (`docs/EVALUATION.md` §7.2).
549
+
550
+ ### 3.5 `test_deploy_config.py`
551
+
552
+ **Scope:** the Phase 18 deploy manifest and its validator, and the **config-hash regression guard**.
553
+
554
+ The docstring states the two properties it proves:
555
+
556
+ > *"These tests prove the manifest is INERT — it is never merged into the config registry and cannot move
557
+ > `Config.hash` — and that the validator rejects each failure mode. Negative cases use in-memory dicts via
558
+ > `validate_documents`, so the real config files are never mutated."*
559
+
560
+ Key elements:
561
+
562
+ | Element | Detail |
563
+ |---|---|
564
+ | `REGISTRY_MARKER` | `test_deploy_manifest_exists_and_declares_the_non_registry_marker` asserts `doc[REGISTRY_MARKER] is False` |
565
+ | `test_validator_reports_no_errors_on_the_real_files` | `assert validate() == []` |
566
+ | `test_validator_cli_exits_zero_as_a_standalone_script` | drives the real entry point **in a child interpreter** (the same one running the suite) *"to prove the module's own `sys.path` bootstrap makes `core` importable WITHOUT pytest's `pythonpath = .` shortcut"* |
567
+ | `GPU_DURATION_KEYS` | `gpu_duration_vqa`, `gpu_duration_grounding`, `gpu_duration_change`, `gpu_duration_optical_sar` |
568
+ | `FROZEN_CONFIG_HASH` | imported from `tests.test_config` rather than re-declared — *"We import it rather than re-declare it, so the two cannot drift."* |
569
+
570
+ The `FROZEN_CONFIG_HASH` import is the anti-drift move applied to the guard itself: the deploy-config test
571
+ does not carry its own copy of `78f1e3700da15aa1`, it imports the authority.
572
+
573
+ The suite also pins the C-8 constraints: `torch_compile` false, CPU mode required, lazy load, single model
574
+ cache (`docs/EVALUATION.md` §7.2).
575
+
576
+ ### 3.6 `test_step7_api_contract_conformance.py`
577
+
578
+ *"STEP 7 section D -- the API contract, tested against the schema that implements it."* This is the
579
+ schema-level counterpart to `test_api_contract_doc.py`: where the doc-guard validates the *examples* in the
580
+ markdown, this module tests the *contract* against the implementing schema directly.
581
+
582
+ ### 3.7 What the doc-guards cannot do
583
+
584
+ The modules say so themselves, and the honesty is worth reproducing because it bounds what a green run
585
+ means:
586
+
587
+ > *"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a
588
+ > decision. What is asserted is that every value a frontend must send or read is the real value, which is
589
+ > the class of error that produces a 422 and a wasted debugging session."*
590
+ > — `tests/unit/test_frontend_guide_doc.py`
591
+
592
+ > *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
593
+ > mis-stated latency budget)."*
594
+ > — `tests/unit/test_api_contract_doc.py`
595
+
596
+ So a doc-guard guarantees **example-level and value-level** conformance. It does not guarantee that a
597
+ sentence is true. That distinction is the reason `test_runbook_doc.py` exists separately — it targets a
598
+ different class of claim (numbers, names, and status honesty) that the example-validators cannot reach.
599
+
600
+ ---
601
+
602
+ ## 4. The evidence engine: purity, determinism and the reproducibility property
603
+
604
+ ### 4.1 Why this suite exists at all
605
+
606
+ The evidence engine is the component that turns a `SpecialistResult` into a *proof*: an ordered,
607
+ identified, digestible `EvidenceCollection`. Everything the release claims about "evidence" downstream —
608
+ the trace, the audit trail, the ability to recompute a verdict from a recorded run — rests on one
609
+ property of this component: **the same inputs must produce the same output, byte for byte.**
610
+
611
+ If that property fails, then two runs of the same request can produce two different evidence sets, and
612
+ no verdict can be recomputed from a record. The evidence engine is therefore the load-bearing component
613
+ behind §9's "verdict-independent evidence" claim, and its suite is the guard on that load-bearing
614
+ property.
615
+
616
+ The suite is `tests/unit/test_evidence_engine.py` — **73 test functions** in 877 lines
617
+ (`tests/unit/test_evidence_engine.py`).
618
+
619
+ ### 4.2 The contract, in the module's own words
620
+
621
+ The module docstring states exactly two failure modes it is written against, and a third honesty rule:
622
+
623
+ > *"The engine's whole value is that it is *reproducible* and *lossless*. So the tests are mostly about
624
+ > the two ways that can silently break:*
625
+ >
626
+ > *determinism — the same claims must produce the same ordered, identified collection even when
627
+ > specialists finish in a different order, or when the uuids differ between processes.*
628
+ > *no loss — no specialist's evidence may vanish without being counted, and agreement between
629
+ > specialists must be recorded rather than collapsed with one name thrown away.*
630
+ >
631
+ > *Plus the honesty rule on confidence: an uncalibrated engine must pass the raw score through and SAY
632
+ > it is uncalibrated, never manufacture a fitted mapping."*
633
+ > — `tests/unit/test_evidence_engine.py`
634
+
635
+ Three properties, then: **determinism**, **no loss**, and **confidence honesty**. The suite is
636
+ organised around them.
637
+
638
+ ### 4.3 The determinism family
639
+
640
+ The core purity claim is a single test whose docstring is the whole argument:
641
+
642
+ ```python
643
+ def test_aggregate_is_deterministic_across_repeat_calls() -> None:
644
+ """Same input -> byte-identical output. The core purity claim."""
645
+ ```
646
+
647
+ It asserts three things across two calls with the same input — identical `ids()`, identical
648
+ `evidence_digest()`, and identical `model_dump()` lists. "Byte-identical output" is not a figure of
649
+ speech: the digest is the serialised form, so a difference anywhere in the collection changes the digest.
650
+
651
+ The determinism family, read from the file:
652
+
653
+ | Test | What it pins |
654
+ |---|---|
655
+ | `test_aggregate_is_deterministic_across_repeat_calls` | the core purity claim (line 132) |
656
+ | `test_order_is_independent_of_input_order` | forward vs. backward input order give the same digest |
657
+ | `test_equal_scores_order_deterministically_by_coordinates` | *"Ties must not fall back to insertion order — that is input-order dependence wearing a disguise"* |
658
+ | `test_higher_score_sorts_first_within_a_type` | the ordering rule inside an evidence type |
659
+ | `test_unscored_items_sort_after_scored_items` | a total order even when some items carry no score |
660
+ | `test_types_are_grouped_together` | type grouping is stable |
661
+ | `test_ids_are_sequential_and_zero_padded` | the identifier format |
662
+ | `test_ids_restart_from_one_for_each_aggregation` | ids are a function of the collection, not global state |
663
+ | `test_source_results_are_not_mutated` | purity — computing evidence does not write back |
664
+ | `test_result_objects_are_not_mutated` | purity, at the result level |
665
+
666
+ The two mutation tests are the *purity* half of "purity and determinism": the engine is a function, not
667
+ a transformer of its inputs. A caller's `SpecialistResult` is the same object after `aggregate` as
668
+ before.
669
+
670
+ ### 4.4 The no-loss family
671
+
672
+ Losslessness is the harder property to test, because the ways evidence can vanish are subtle. The
673
+ deduplication tests are where it is pinned:
674
+
675
+ | Test | What it pins |
676
+ |---|---|
677
+ | `test_payload_only_differences_are_deduplicated_as_one_claim` | two items differing only in payload are one claim |
678
+ | `test_identical_items_from_one_specialist_collapse` | intra-specialist dedup |
679
+ | `test_deduplication_ignores_random_uuid` | dedup is on content, not on the process-local uuid |
680
+ | `test_deduplication_ignores_payload_differences_and_merges_them` | payloads merge rather than one being dropped |
681
+ | `test_different_coordinates_are_not_deduplicated` | dedup does not over-collapse |
682
+ | `test_different_coordinate_systems_are_not_deduplicated` | a normalised box and a pixel box are different claims |
683
+ | `test_same_claim_from_two_specialists_is_kept_and_marked_corroborated` | **agreement is recorded, not collapsed** — the second specialist's name is kept as corroboration |
684
+ | `test_unrelated_items_carry_no_corroboration_marker` | the marker is meaningful, not always-on |
685
+ | `test_multiple_results_aggregate_without_loss` | the headline no-loss assertion |
686
+ | `test_aggregation_does_not_double_count_a_specialists_own_evidence` | lossless ≠ duplicating |
687
+
688
+ The corroboration tests are the direct expression of the docstring's *"agreement between specialists
689
+ must be recorded rather than collapsed with one name thrown away."* This is a design decision encoded
690
+ as a test: when two specialists agree, the collection records **both** names and marks the claim
691
+ corroborated, rather than keeping one and discarding the other.
692
+
693
+ The cap tests pin that truncation is *counted*, not silent:
694
+
695
+ | Test | What it pins |
696
+ |---|---|
697
+ | `test_zero_max_items_is_rejected` | an invalid cap is refused |
698
+ | `test_cap_is_respected_and_the_drop_is_counted` | items past the cap are dropped **and counted** |
699
+ | `test_no_truncation_flag_when_under_the_cap` | the truncation flag is truthful |
700
+ | `test_default_cap_matches_the_config_value` | `DEFAULT_MAX_ITEMS` tracks the config |
701
+ | `test_sources_are_reported_before_the_cap_is_applied` | *which* specialists contributed is recorded before truncation |
702
+
703
+ That last test is a specific honesty property: if the cap drops a specialist's only item, the collection
704
+ still records that the specialist contributed. Losslessness at the *source* level survives truncation at
705
+ the *item* level.
706
+
707
+ The empty/degenerate cases are pinned too: `test_empty_input_yields_an_empty_collection`,
708
+ `test_a_result_with_no_evidence_contributes_nothing`, `test_accepts_a_single_result_without_wrapping`,
709
+ and the two argument-error tests `test_both_results_and_evidence_is_rejected` /
710
+ `test_neither_results_nor_evidence_is_rejected`.
711
+
712
+ ### 4.5 The confidence honesty rule
713
+
714
+ This is where the project's style guide (`DOCS_STYLE_GUIDE.md` §3) meets a test. The rule the engine
715
+ must obey: an uncalibrated engine **passes the raw score through and says it is uncalibrated**; it never
716
+ manufactures a fitted mapping. The tests that pin this:
717
+
718
+ | Test | What it pins |
719
+ |---|---|
720
+ | `test_no_calibration_passes_raw_through_and_says_so` | the default is identity + an explicit "uncalibrated" label |
721
+ | `test_uncalibrated_is_never_labelled_as_fitted` | the label cannot be silently upgraded |
722
+ | `test_identity_temperature_is_reported_as_uncalibrated` | `T = 1.0` is *not* calibration |
723
+ | `test_engine_without_calibration_does_not_claim_calibration` | the engine-level claim |
724
+ | `test_raw_outside_unit_range_is_clamped_not_rejected` | a robustness decision, pinned |
725
+ | `test_non_finite_raw_is_refused` | NaN/Inf are refused, not clamped |
726
+
727
+ This matters because the project's calibration genuinely made the metric **worse** (ECE
728
+ `0.013755 → 0.014929`, `DOCS_STYLE_GUIDE.md` §3) and is retained only because it is in the frozen
729
+ config. The engine must therefore never present calibration as an improvement, and these tests are what
730
+ enforce that at the code level. The `METHOD_TEMPERATURE` / `METHOD_UNCALIBRATED` constants are the two
731
+ labels, and `test_uncalibrated_is_never_labelled_as_fitted` is the guard that the first is never used
732
+ for the second.
733
+
734
+ The fitted path is tested separately and is deterministic:
735
+ `test_temperature_scaling_is_applied_when_fitted`,
736
+ `test_temperature_above_one_softens_and_below_one_sharpens`,
737
+ `test_temperature_scaling_is_monotonic`, `test_temperature_scaling_is_deterministic`,
738
+ `test_endpoint_scores_do_not_produce_nan`, `test_invalid_temperatures_are_rejected`, and
739
+ `test_specialist_degradation_survives_calibration` / `test_specialist_components_are_preserved_through_calibration`.
740
+
741
+ ### 4.6 Calibration artifact I/O
742
+
743
+ The engine reads a calibration artifact from disk, and the tests pin the failure modes explicitly rather
744
+ than letting them degrade silently:
745
+
746
+ | Test | What it pins |
747
+ |---|---|
748
+ | `test_artifact_round_trips_through_json` | write→read is lossless |
749
+ | `test_nested_artifact_form_is_accepted` | both artifact shapes are accepted |
750
+ | `test_missing_artifact_degrades_instead_of_failing` | a missing artifact is a *degradation* by default |
751
+ | `test_missing_artifact_raises_when_required` | …but raises when the caller says it is required |
752
+ | `test_malformed_artifact_is_not_silently_swallowed` | a corrupt artifact is an error, not a shrug |
753
+ | `test_artifact_without_a_temperature_is_rejected` | the required field is required |
754
+ | `test_load_calibration_respects_the_master_switch` | the config switch is honoured |
755
+ | `test_load_calibration_reads_the_configured_file` | the configured path is the one read |
756
+ | `test_load_calibration_without_a_file_configured_is_none` | no config ⇒ `None`, not a guess |
757
+ | `test_engine_from_config_uses_the_configured_cap` / `..._falls_back_to_the_default_cap` | the cap wiring |
758
+
759
+ The pairing of `degrades_instead_of_failing` with `raises_when_required` is the pattern this repository
760
+ uses throughout: a soft default for a convenience path, and a hard error for the path where silence
761
+ would be a lie.
762
+
763
+ ### 4.7 The scratch-root workaround, and why it is documented in the test file
764
+
765
+ The suite does **not** use pytest's built-in `tmp_path`. The module says why, and the reason is a
766
+ sandbox property, not a preference:
767
+
768
+ > *"Scratch root, repo-local. The built-in `tmp_path` fixture cannot finalize under this machine's
769
+ > sandbox, so artifact tests use this instead. It is the same directory `--basetemp=.pytest_tmp` points
770
+ > pytest at, so nothing is written outside the workspace."*
771
+ > — `tests/unit/test_evidence_engine.py`
772
+
773
+ ```python
774
+ SCRATCH_ROOT = Path(__file__).resolve().parents[2] / ".pytest_tmp"
775
+ ```
776
+
777
+ And the module docstring carries the operational requirement:
778
+
779
+ > *"The `--basetemp=.pytest_tmp` flag is mandatory (see the project handoff): the default pytest temp
780
+ > root triggers a sandbox denial on this machine."*
781
+
782
+ The `scratch` fixture is deliberately **not** cleaned up on teardown, and the docstring says why: *"the
783
+ sandbox's safe-delete guard rejects the recursive delete, and a test that writes into the repo's own
784
+ ignored scratch directory is harmless. Each test gets a unique path so no state leaks between them."*
785
+ This is the same guard that causes §7's four `test_safe_delete_shim` failures — documented here as a
786
+ design constraint rather than hidden as a quirk.
787
+
788
+ ### 4.8 The public surface
789
+
790
+ `test_module_public_surface_is_declared` pins the exported names, and
791
+ `test_serialises_through_the_schema_serialiser` pins that the collection serialises through the *schema*
792
+ serialiser (not a bespoke one), so the wire form and the evidence form cannot drift. The shorthand and
793
+ helpers are pinned by `test_aggregate_evidence_shorthand_matches_the_engine`,
794
+ `test_collection_helpers_filter_correctly`, `test_collection_summary_reports_the_observable_facts`,
795
+ `test_digest_changes_when_a_claim_changes`, `test_digest_is_stable_across_regenerated_uuids`,
796
+ `test_collection_is_iterable_and_sized`.
797
+
798
+ `test_digest_is_stable_across_regenerated_uuids` is the one that makes §9 possible: because the digest
799
+ ignores process-local uuids, a recorded run's evidence digest can be recomputed in a *different process*
800
+ and still match. Without it, "recompute the verdict from the record" would be a claim you could not test.
801
+
802
+ ### 4.9 What the evidence-engine suite does NOT establish
803
+
804
+ - It does **not** establish that the specialists produce correct evidence — only that the *aggregator*
805
+ is deterministic and lossless given whatever they produce.
806
+ - It does **not** establish end-to-end reproducibility of a live run; that is §8's job.
807
+ - It does **not** establish that the calibration is *good* — the tests pin that calibration is
808
+ *honestly labelled*, not that it improves anything (it does not; ECE got worse).
809
+
810
+ ---
811
+
812
+ ## 5. The frontend live-wiring suite
813
+
814
+ ### 5.1 What it proves, and why it reads source text
815
+
816
+ `tests/unit/test_frontend_live_wiring.py` is the regression net for the frontend↔orchestrator boundary.
817
+ Its docstring states two things it proves, **both of which were broken before the change**:
818
+
819
+ > *"A. **CORS.** The deployment allowed exactly one origin (`https://satquery.pages.dev`), so the
820
+ > frontend could not be driven from a local development server at all: every request from
821
+ > `http://localhost:8080` was refused with `Disallowed CORS origin`, and a developer had to deploy to
822
+ > Cloudflare to test a one-line JavaScript change. The fix adds explicit localhost origins while keeping
823
+ > production listed and keeping `*` rejected.*
824
+ >
825
+ > *B. **The task vocabulary.** The page's router and the server's `Task` enum must agree. They are two
826
+ > independently written vocabularies in two languages, and nothing previously checked that a value the
827
+ > frontend would send is a value the server accepts. This module reads BOTH and compares them."*
828
+ > — `tests/unit/test_frontend_live_wiring.py`
829
+
830
+ And it pins a third, subtler defect — one invisible to a validity check:
831
+
832
+ > *"It also pins the `chang` word-boundary defect: `\bchang\b` cannot match 'changed', so the page's own
833
+ > default question ('What changed here?') routed to `vqa` instead of the change path. That is invisible
834
+ > to any test that only checks 'is this a valid enum value' — both branches produce a valid value. The
835
+ > test therefore asserts the ROUTE, not just the validity."*
836
+
837
+ That last sentence is the philosophical centre of the whole suite: **assert the route, not the
838
+ validity.** A test that asks "is `vqa` a valid task?" passes for both the correct and the defective
839
+ router, because both emit valid tasks. Only a test that asks "does *this question* route to *this
840
+ task*?" can see the defect.
841
+
842
+ **Why source-text parsing.** The module reads the frontend source rather than importing it, and says so:
843
+
844
+ > *"the frontend is plain ES5 with no build step and no module exports, so there is no importable
845
+ > symbol. Parsing the source is the only way to assert the shipped bytes. The parsing is anchored on
846
+ > named tokens rather than line numbers, so it does not silently pass if the file is restructured."*
847
+
848
+ The anchors are module-level path constants — `REPO_ROOT`, `DEPLOY_RENDER`, `FRONTEND_JS`, `MISSION_JS`,
849
+ `LIVE_JS`, `CORE_JS`, `MISSION_HTML` — and the production origin is a single constant,
850
+ `PRODUCTION_ORIGIN = "https://satquery.pages.dev"`. The CORS half imports the orchestrator with a clean
851
+ environment via `_render_module()`, because `_allowed_origins` reads `os.environ` at **call** time while
852
+ `_DEV_ORIGINS` is built at **import** time — a distinction the module's helper comment records
853
+ explicitly.
854
+
855
+ ### 5.2 The measured result
856
+
857
+ **106 passed** (`release/CURRENT_RELEASE_STATE.md:121`; `README.md` §The test suites;
858
+ `docs/EVALUATION.md` §7.1). The command is:
859
+
860
+ ```bash
861
+ .venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
862
+ ```
863
+
864
+ This is a **re-run-this-session** figure, and it is the number this chapter quotes.
865
+
866
+ ### 5.3 The fifteen test classes
867
+
868
+ The classes were read from the file. Each name is a claim:
869
+
870
+ | # | Class | What it pins |
871
+ |---|---|---|
872
+ | 1 | `TestTheCorsAllowlistKeepsProductionAndAddsDevelopment` | production origin always allowed, not duplicated, dev origins added, `*` rejected even when hidden in a list |
873
+ | 2 | `TestTheCorsDecisionMatchesTheAllowlist` | the response-leg decision matches the allowlist (echo, no-headers, lookalike-host rejection) |
874
+ | 3 | `TestTheFrontendSpeaksTheServersTaskVocabulary` | every mapped value is a real server task; the mapping covers every task the router can emit; specialist names are human labels, not tasks |
875
+ | 4 | `TestTheDefaultQuestionReachesTheChangePath` | the page's own default question routes to the change path; the change stem matches its inflections; the temporal slot is required for a change question |
876
+ | 5 | `TestALocationQuestionIsNotAChangeQuestion` | a "where" question reads as grounding, with one or two assets |
877
+ | 6 | `TestTheArchitecturePolicyDoesNotReadALocationAsAChange` | the architecture sample routes to grounding; a place word in a "where" question is not a change |
878
+ | 7 | `TestTheTaskRespectsThePairRequirement` | paired tasks substitute correctly when given one asset; single-asset tasks are never substituted |
879
+ | 8 | `TestTheLiveClientUsesTheRealIngestionPath` | the client targets the orchestrator's routes; uploads send raw bytes with a derived content type; the infer body carries asset ids under `assets`; the client never sends a filesystem path |
880
+ | 9 | `TestThePageLoadsTheLiveClient` | `mission.html` includes live JS before mission JS; the declared API base is an absolute origin, not the relative fallback, not the frontend origin; no fixture image is wired into the analysis path |
881
+ | 10 | `TestTheLivePathKeepsTheHonestFailureContract` | a failure is attributed to the step that failed; the preview is kept for the no-file case |
882
+ | 11 | `TestTheCaptionStopsCallingARealUploadIllustrative` | a successful run rewrites the leading word and the alt text; the failure path does not claim an analysis happened; a new file clears the previous run's claim |
883
+ | 12 | `TestTheFrontendSendsOnlyTheAssetsTheTaskRequires` | a single-asset task with a pair uploaded uses only `t1`; the old wiring would have sent both; change/optical-SAR send the pair when present |
884
+ | 13 | `TestDescriptiveQueriesRouteToCaption` | descriptive queries reach caption; a description is not mistaken for change; caption is single-asset |
885
+ | 14 | `TestTheOpticalSarInputValidation` | a real optical+GeoTIFF SAR pair is ok; two plain photos warn; a missing SAR image is an error; the warning names the optical+radar expectation |
886
+ | 15 | `TestTheServerErrorTranslation` | invalid request gets asset guidance; 422 gets malformed guidance; recoverable `false` gets restart guidance; the server's message and detail are preserved; a null error is handled; a recoverable transport code is not called unrecoverable |
887
+
888
+ Two of these deserve a note because they encode *counter-intuitive* decisions:
889
+
890
+ **`TestTheFrontendSendsOnlyTheAssetsTheTaskRequires` (class 12).** The class contains
891
+ `test_the_old_wiring_would_have_sent_both_files_for_vqa` and
892
+ `test_the_fix_sends_one_file_where_the_old_sent_two`. This is a *differential* test: it asserts not just
893
+ that the new code is right, but that the *old* code was wrong in the specific way claimed. That is the
894
+ falsification discipline (§1.3) applied inside a suite — the fix is only meaningful if the defect it
895
+ fixes is real, and the test proves the defect by asserting the old behaviour differs.
896
+
897
+ **`TestTheCaptionStopsCallingARealUploadIllustrative` (class 11).** This is a *copy-honesty* test. The
898
+ page used to caption a real analysis "illustrative"; the suite asserts the leading word and the alt text
899
+ are rewritten on success, and — crucially — that the **failure path does not claim an analysis
900
+ happened** (`test_the_failure_path_does_not_claim_an_analysis_happened`). A UI that says "analysis
901
+ complete" on a failure is a claim defect, and it is tested as one.
902
+
903
+ ### 5.4 The historical count: 94 → 100 → 106
904
+
905
+ A reviewer may see a different number. `docs/FINAL_DELIVERY_REPORT.md §7` records **94 passed** for this
906
+ file, while the current figure is **106**. The difference is the suite growing after the delivery report
907
+ was written: the regression tests added for the router fix took this file from **100 → 106** tests
908
+ (`DELIVERY_REPORT_2026-09-25.md`, regression-tests note). This chapter quotes **106** and cites both
909
+ figures so neither looks like a contradiction (`docs/REPRODUCIBILITY.md` §5.2).
910
+
911
+ ### 5.5 What the frontend live-wiring suite does NOT establish
912
+
913
+ - It does **not** run the browser. It asserts the *shipped bytes* of the frontend against the
914
+ *orchestrator module* in-process. Whether the page actually loads and runs is §8's job (live
915
+ validation).
916
+ - It does **not** prove the router's *quality* — only that specific questions route to the specific
917
+ tasks asserted. The router's residual misroutes (`"What is the new runway?"` → `change`) are real and
918
+ recorded (`docs/LIMITATIONS.md`; `LIVE_VALIDATION_POSTFIX.md`), not hidden by a green suite.
919
+ - It does **not** test the inference service. It tests the boundary the frontend speaks to, which is the
920
+ orchestrator.
921
+
922
+ ---
923
+
924
+ ## 6. The doc/frontend suite
925
+
926
+ ### 6.1 The five files, and the measured result
927
+
928
+ The "doc/frontend suite" is a named grouping of **five files** that together report **183 passed**
929
+ (`docs/FINAL_DELIVERY_REPORT.md §7`; `docs/EVALUATION.md` §7.2; `docs/REPRODUCIBILITY.md` §5.3). The
930
+ five are:
931
+
932
+ | File | What it pins |
933
+ |---|---|
934
+ | `tests/unit/test_frontend_guide_doc.py` | both documents agree on tasks, coordinate systems, evidence types, geometry shapes; no login screen; no secrets in the browser; the confidence-property trap; serialized analyses |
935
+ | `tests/unit/test_frontend_live_wiring.py` | the frontend↔orchestrator boundary — see §5 |
936
+ | `tests/unit/test_api_contract_doc.py` | every JSON block parses and validates; no contract example fabricates an artifact URI; documented task/coordinate-system values are the real ones; every error code documented; the frozen config hash is recorded; the calibration caveat is recorded |
937
+ | `tests/unit/test_runbook_doc.py` | the runbook names only public serving entrypoints; quoted artifact sizes match the real files; the capability list matches the registry; env vars are the architecture ones; undecided items declared undecided; **no claim that a deployment happened**; verified distinguished from designed |
938
+ | `tests/unit/test_deploy_config.py` | the deploy manifest declares the non-registry marker; the validator reports no errors; the **config-hash regression guard**; the C-8 constraints (`torch_compile` false, cpu mode required, lazy load, single model cache) |
939
+
940
+ These are **documentation conformance tests**: they fail if a doc drifts from the code it describes.
941
+ They are the mechanism that keeps this release's documentation honest — the reason a doc claim can be
942
+ trusted is that a test would go red if the claim stopped matching the code.
943
+
944
+ ### 6.2 Why four of the five are doc-guards, and one is not
945
+
946
+ Four of the five (`test_frontend_guide_doc`, `test_api_contract_doc`, `test_runbook_doc`,
947
+ `test_deploy_config`) are doc-guards, analysed individually in §3. The fifth
948
+ (`test_frontend_live_wiring`) is not a doc-guard — it is a code-behaviour suite that happens to be
949
+ grouped here because it reads frontend source text. The grouping is by *audience* (frontend +
950
+ documentation) rather than by mechanism, and that is why the count 183 covers both.
951
+
952
+ ### 6.3 The doc-guard guarantee, restated
953
+
954
+ The doc-guards guarantee **example-level and value-level** conformance, not prose truth (§3.7). The two
955
+ module docstrings state the limit themselves:
956
+
957
+ > *"These tests validate the EXAMPLES, not the prose. They cannot detect a wrong sentence (e.g. a
958
+ > mis-stated latency budget)."* — `tests/unit/test_api_contract_doc.py`
959
+
960
+ > *"Prose can still be wrong in ways these tests cannot see -- a mis-stated latency, a wrong reason for a
961
+ > decision."* — `tests/unit/test_frontend_guide_doc.py`
962
+
963
+ `test_runbook_doc.py` reaches a different class of claim — numbers, names, and **status honesty** — and
964
+ its most important assertion is that the document itself *labels its unimplemented sections as
965
+ unimplemented* (§3.4). That is the guard against the exact overclaim the style guide forbids.
966
+
967
+ ### 6.4 The config-hash regression guard
968
+
969
+ `test_deploy_config.py` carries the anti-drift move that makes it more than a validator: it does **not**
970
+ re-declare the frozen config hash. It imports it:
971
+
972
+ > *"We import it rather than re-declare it, so the two cannot drift."*
973
+ > — `tests/unit/test_deploy_config.py`
974
+
975
+ The authority is `tests.test_config`, and the value is the frozen hash `78f1e3700da15aa1`
976
+ (`DOCS_STYLE_GUIDE.md` §3). A second copy of the hash anywhere in the tree would be a latent
977
+ contradiction; the import makes drift structurally impossible.
978
+
979
+ The same file proves the manifest is **INERT** — *"it is never merged into the config registry and
980
+ cannot move `Config.hash`"* — and that negative cases use in-memory dicts via `validate_documents` *"so
981
+ the real config files are never mutated."* It also drives the validator CLI **in a child interpreter**
982
+ to prove the module's own `sys.path` bootstrap makes `core` importable *"WITHOUT pytest's `pythonpath =
983
+ .` shortcut"* (`test_validator_cli_exits_zero_as_a_standalone_script`).
984
+
985
+ ### 6.5 What the doc/frontend suite does NOT establish
986
+
987
+ - It does **not** establish that any deployment exists. `test_runbook_doc.py` explicitly asserts the
988
+ opposite direction: that the document does **not** claim a deployment happened (§3.4).
989
+ - It does **not** validate prose (§6.3).
990
+ - It does **not** check that the *code* is correct — only that the docs match the code. If the code and
991
+ the docs are wrong together, a doc-guard is green.
992
+
993
+ ---
994
+
995
+ ## 7. The full `tests/unit` result, and the 5–6 environmental failures
996
+
997
+ This is the section where the release's honesty discipline is most visible. Running the entire unit tree
998
+ produces failures. They are reported, classified, and *not* reclassified as passes.
999
+
1000
+ ### 7.1 The result
1001
+
1002
+ Running `python -m pytest tests/unit` in the authoring sandbox produces **5–6 failures**. Every one is
1003
+ **environmental or ordering**-related, not a regression in shipped code. The classification, exactly as
1004
+ `docs/FINAL_DELIVERY_REPORT.md §7` records it:
1005
+
1006
+ | # | Failure | Attribution | Regression? |
1007
+ |---|---|---|---|
1008
+ | 1–4 | `test_safe_delete_shim` — **4 failures** | the sandbox's bulk-**delete guard** (Windows verbatim-path behaviour) | **No** |
1009
+ | 5 | one **ordering flake** in the router route test | test **ordering** (passes in isolation) | **No** |
1010
+ | 6 | one **stale adapter test** (`optical_sar` absent when CROMA unshipped) | stale test — **CROMA is now shipped** | **No** |
1011
+
1012
+ ### 7.2 The evidence that they are not regressions
1013
+
1014
+ The evidence is the **re-run**. Re-running the affected files **together** gives **137 passed**
1015
+ (`docs/FINAL_DELIVERY_REPORT.md §7`; `docs/EVALUATION.md` §7.3–7.4). The logic is stated plainly:
1016
+
1017
+ > *"if the failures were real regressions in the code under test, re-running those files together would
1018
+ > still fail. They do not — which is what distinguishes an environmental/ordering failure from a
1019
+ > regression."*
1020
+
1021
+ This is the falsification discipline (§1.3) applied to a *test result* rather than to a *guard*: the
1022
+ claim "these are not regressions" is itself tested, by a re-run designed to fail if it were false.
1023
+
1024
+ ### 7.3 The four `test_safe_delete_shim` failures, in detail
1025
+
1026
+ `tests/unit/test_safe_delete_shim.py` (338 lines, `pytestmark = pytest.mark.unit`) guards two defects in
1027
+ the WorkBuddy Windows safe-delete shim (`cli/vendor/shim/sitecustomize.py`) that produced **phantom test
1028
+ failures** in this repository. The module's docstring records both with measurements.
1029
+
1030
+ **Defect 1 — a verbatim temp path was not recognised as a temp path.**
1031
+
1032
+ > *"`_path_for_compare` returned `normcase(realpath(abspath(path)))`, which preserves the Windows verbatim
1033
+ > prefix `\\?\`. `os.path.relpath` cannot relate `\\?\c:\...` to `c:\...`, so `_is_under_root` returned
1034
+ > False and the OS-temp exemption in `_should_bypass_safe_delete` did not apply."*
1035
+
1036
+ Measured, before the fix:
1037
+
1038
+ ```
1039
+ plain temp subdir -> bypass True
1040
+ verbatim temp subdir -> bypass False <-- the bug
1041
+ non-temp subdir -> bypass False
1042
+ ```
1043
+
1044
+ That is why routine pytest `garbage-*` collection — which walks
1045
+ `\\?\C:\Users\...\Temp\pytest-of-anish\garbage-*` — reached the bulk guard, tripped `confirmRequired` at
1046
+ **69 entries against a threshold of 50**, and **latched a rejection that then blocked every delete in
1047
+ the conversation**.
1048
+
1049
+ **Defect 2 — a successful delete was reported as a failure.**
1050
+
1051
+ > *"`_platform_trash` raised whenever `SHFileOperationW` returned non-zero."*
1052
+
1053
+ Measured on the authoring host with `FOF_ALLOWUNDO|FOF_NOCONFIRMATION|FOF_NOERRORUI|FOF_SILENT`, calling
1054
+ shell32 directly (the shim not involved):
1055
+
1056
+ ```
1057
+ series non-temp rc temp rc
1058
+ 4 regions x 3 interleaved rounds 2, 2, 2, ... 0, 0, 0
1059
+ 6 processes x 20 deletes 2 (120/120) 0 (120/120)
1060
+ ```
1061
+
1062
+ The target is removed in **every** case and the user's Recycle Bin is populated and active (1,300+
1063
+ `$I`/`$R` entries), so the return code is **not** a reliable "could not delete" signal. The module is
1064
+ explicit that the code is **intermittent** and that the precise Windows-internal trigger is not known:
1065
+
1066
+ > *"across pytest invocations the same non-temp delete sometimes returned 0, and in one run a *temp*
1067
+ > delete returned non-zero. The precise Windows-internal trigger is **NOT established** and is not
1068
+ > claimed here."*
1069
+
1070
+ This is a place where this chapter writes **`UNKNOWN — not established from the available evidence`**,
1071
+ following the module's own honesty.
1072
+
1073
+ **What the guard asserts** (the `WHAT IS ASSERTED` block of the docstring):
1074
+
1075
+ - A verbatim-prefixed path inside the OS temp root compares equal to its plain form and is exempted.
1076
+ - A verbatim-prefixed path *outside* the temp root is still guarded — *"the fix normalises the prefix, it
1077
+ does not widen the exemption."*
1078
+ - Deleting through a verbatim temp path succeeds end to end.
1079
+ - A non-zero shell code on a delete whose target is genuinely gone does not raise (stubbed shell, stubbed
1080
+ `lexists` — deterministic, deletes nothing).
1081
+ - A non-zero shell code on a delete whose target **survives** still raises — *"Fail-closed is preserved;
1082
+ only the false positive is gone."*
1083
+ - Against the real shell, a non-temp delete does not raise (behavioural confirmation, not the guard).
1084
+
1085
+ **Why the guard is stub-based.** The module states the reason, and it is a direct consequence of defect
1086
+ 2's intermittency:
1087
+
1088
+ > *"the guard for this defect stubs the shell: a test that depends on the real return code would
1089
+ > sometimes pass against the broken shim. The stub tests fail on the original shim on every run."*
1090
+
1091
+ That is the difference between a test that *sometimes* catches a defect and a test that *always* does.
1092
+
1093
+ **The rule for extending the file** is stated in the docstring and is worth reproducing, because it is
1094
+ the exact failure this file exists to prevent:
1095
+
1096
+ > *"**Do not perform a real delete outside the OS temp root.** … An earlier draft of the defect-2 test
1097
+ > did exactly that, tripped `confirmRequired` at the threshold and latched a rejection that blocked every
1098
+ > subsequent delete in the session — the very failure this file exists to prevent."*
1099
+
1100
+ The module also **skips cleanly** when the shim is absent or disabled, via
1101
+ `pytest.importorskip("sitecustomize", ...)` plus `skipif` markers (`needs_shim_helpers`, `windows_only`).
1102
+ So on a normal user environment these tests skip; in the authoring sandbox they reach the guard and 4 of
1103
+ the 8 fail. The constants involved are `VERBATIM = "\\\\?\\"` and a `NON_TEMP` probe path that is *"never
1104
+ created"* — the predicates tested are pure.
1105
+
1106
+ ### 7.4 The ordering flake
1107
+
1108
+ One router route test fails only when the full suite runs in a particular order, and **passes in
1109
+ isolation**. That is the definition of an ordering flake: the failure is a function of collection order,
1110
+ not of the code. `docs/REPRODUCIBILITY.md` §5.4 records it in the same row-set as the shim failures.
1111
+
1112
+ ### 7.5 The stale adapter test
1113
+
1114
+ One test asserts `optical_sar` is **absent when CROMA is unshipped**. CROMA is now shipped
1115
+ (`CURRENT_RELEASE_STATE.md` §1 capability contract: `optical_sar` `available: true`), so the assertion's
1116
+ premise has changed. This is a **stale test**, not a code defect: the test encodes an assumption about
1117
+ the environment that the environment has outgrown.
1118
+
1119
+ ### 7.6 A different environment state: the 2,237-passed run
1120
+
1121
+ `docs/PHASE12_115_METRIC_COMPUTED.md` §8 records, for **2026-09-22**, a full unit suite result of
1122
+ **2,237 passed, 0 failed (622.68 s)**. That is a real, measured result in a *different* environment state
1123
+ (the safe-delete shim's guard state and the CROMA shipping status differ). It is recorded rather than
1124
+ suppressed, *"because the two results are both true and the difference is exactly the
1125
+ environmental/ordering story"* (`docs/EVALUATION.md` §7.3).
1126
+
1127
+ The same document records a correction that is itself an honesty lesson: an earlier revision claimed
1128
+ `2,179 passed` computed as `2,171 + 8`, *"presented as a check but the total was never measured — it was
1129
+ inferred from a stale baseline."* The measured figure was **2,237**. This is the style guide's rule
1130
+ (`DOCS_STYLE_GUIDE.md` §1.3) in the wild: an inferred number presented as a measurement was corrected to
1131
+ the measured one.
1132
+
1133
+ ### 7.7 The UNKNOWN: the collected count
1134
+
1135
+ The exact **collected** test count for the current-session full-suite run is
1136
+ **`UNKNOWN — not established from the available evidence`**. The record gives the failure classification
1137
+ and the 137-passed re-run, but not a collected total. Per-suite counts are known (106, 183, 137, 73, 51);
1138
+ the single collected total is not.
1139
+
1140
+ ### 7.8 Why this is reported rather than hidden
1141
+
1142
+ The style guide's first rule is `DO NOT FABRICATE`, and its status rule is *"a mixed result is never 'all
1143
+ work perfectly'"* (`DOCS_STYLE_GUIDE.md` §1.7). A full-suite run with 5–6 failures is a mixed result.
1144
+ Presenting only the 106 and 183 would be exactly the "all work perfectly" overclaim the guide forbids.
1145
+ The correct presentation is the three-row table (§7.1) plus the re-run evidence (§7.2) plus the explicit
1146
+ UNKNOWN (§7.7) — which is what this section is.
1147
+
1148
+ ---
1149
+
1150
+ ## 8. The live-validation harness and its integrity discipline
1151
+
1152
+ ### 8.1 What live validation is, and why it is separate from the suites
1153
+
1154
+ Every suite in §2–§7 runs in-process. `docs/API_CONTRACT.md` §8 states the boundary plainly: *"no test in
1155
+ this repository dials a network address."* A green `tests/integration` run is therefore **not** evidence
1156
+ about a deployment. **Live validation** is the separate activity that drives the *deployed* stack — the
1157
+ real Cloudflare Pages frontend, the real Render orchestrator, the real tunnel, the real Codespace — with
1158
+ a browser, and records what happened.
1159
+
1160
+ The record lives at `.workbuddy-ai/scratch/live_validation/` — `LIVE_VALIDATION_POSTFIX.md`, the raw
1161
+ harness outputs (`run_output.txt`, `run_final2.txt`, `run_final3.txt`), the recomputed verdicts
1162
+ (`results_final.json`, `results_pass3.json`), and eight per-case screenshots (`A1_vqa.png` … `B2_newairport.png`).
1163
+
1164
+ ### 8.2 The three passes, summarised
1165
+
1166
+ | Pass | HEAD | Harness | Result |
1167
+ |---|---|---|---|
1168
+ | 1 | `ff46eba42b18` + `d413d3672311` | v1 (`fill_input`) | 8/8 |
1169
+ | 2 | `2d7ae53b482d` | v2 asserting | 8/8 |
1170
+ | 3 | `2d7ae53b482d` | v2 asserting (pre-discriminator-fix) | 8/8 (recomputed) |
1171
+
1172
+ **24 live runs, 24 correct dispatches, 0 mock nodes, no run id repeated across passes.** The trace fill
1173
+ was **94.4444 %** on every case, and every case carries a real `run_*` id, the Hugging Face link in the
1174
+ DOM, and the `capabilities → assets → infer` call sequence all addressed to
1175
+ `satquery-backend-m4yv.onrender.com` (two `assets` calls for the pair tasks).
1176
+
1177
+ The eight cases are Phase A (six regression cases: `vqa`, `caption`, `grounding`, `change`,
1178
+ `change_vqa`, `optical_sar`) and Phase B (the two defect cases: `"Where are the built-up areas in this
1179
+ image?"` and `"Where is the new airport?"`, both expected to reach `grounding` with a **single** asset —
1180
+ the exact condition under which the old router collapsed to `vqa` and answered "River").
1181
+
1182
+ ### 8.3 The CDP synthetic-key-event trap
1183
+
1184
+ This is the integrity lesson that shaped the harness, and it is the most transferable finding in this
1185
+ chapter.
1186
+
1187
+ **The shape of the false pass.** Pass 1's harness drove the query box with `fill_input()`, which types
1188
+ with **real CDP key events**. A re-run attempt failed on case 1 with `run_id=0002`, `mock_nodes=9`,
1189
+ `answer="No answer yet"`, and only a `capabilities` call — the **mock path**. The root cause, measured
1190
+ directly:
1191
+
1192
+ > *"`fill_input` types with **real CDP key events**, and Chrome **drops synthesized key events when the
1193
+ > browser window does not hold OS focus**; it has **no assertion**, so it clicked Run with the page's
1194
+ > default query still in the box. Measured directly: with Chrome backgrounded, `press_key("Z")` left
1195
+ > `#qtext.value` unchanged, while `type_text("Q")` (CDP `Input.insertText`, not focus-gated) inserted
1196
+ > fine."*
1197
+ > — `LIVE_VALIDATION_POSTFIX.md`
1198
+
1199
+ The measurement is recorded in `diag_focus.harness`, which prints the value of `#qtext` after
1200
+ `press_key("Z")` and after `type_text("Q")` — showing the first is a no-op and the second works.
1201
+
1202
+ **The generalisation.** A harness that types into a form and then reads a *result* cannot tell the
1203
+ difference between:
1204
+
1205
+ - "the query was entered, and the run used it", and
1206
+ - "the query was never entered, and the run used the page's default".
1207
+
1208
+ Both produce a result. The result of the second is a *plausible* answer to a *different* question, and a
1209
+ harness without an assertion records it as the verdict for the first. That is a **silent false pass**,
1210
+ and it is the reason the fix is not "use a different typing helper" but "**assert the form state before
1211
+ dispatch**".
1212
+
1213
+ ### 8.4 The fix: deterministic query entry plus pre-dispatch assertions
1214
+
1215
+ The harness was rebuilt as `run_all_postfix2.harness`, whose header states the change:
1216
+
1217
+ > *"v2 sets the query deterministically (js value + input event, falling back to `Input.insertText`) and
1218
+ > ASSERTS the form state before clicking, recording `q_ok` / `obs_ok` / `t0_ok` and a computed verdict
1219
+ > per case."*
1220
+
1221
+ The three assertions, and the failure each prevents:
1222
+
1223
+ | Assertion | Checks | Failure it prevents |
1224
+ |---|---|---|
1225
+ | `q_ok` | `#qtext.value` holds the intended query | the silent-drop false pass |
1226
+ | `obs_ok` | `#obsTail == 'ready'` **and** one file on `#fileInput` | an upload that did not land |
1227
+ | `t0_ok` | both frames present, for pair tasks | a pair task run on one asset |
1228
+ | `no_mock_nodes` | `mock_nodes == 0` | the preview path being recorded as live |
1229
+
1230
+ The rule, stated in the session handoff (`HANDOFF_NEXT_AGENT.md` §5.2) and reproduced in
1231
+ `docs/architecture/09-frontend.md` §10.3:
1232
+
1233
+ > *"**Do NOT use `fill_input()` or `press_key()` to enter the query.** They type with real CDP key events,
1234
+ > which Chrome **silently drops when the browser window does not hold OS focus** … Use `js()` to set
1235
+ > `#qtext.value` (plus `input`/`change` events) and/or `type_text()` (CDP `Input.insertText`, not
1236
+ > focus-gated). **Always assert the form state before clicking Run** … or a no-op will be recorded as a
1237
+ > pass."*
1238
+
1239
+ The `set_query()` function in the harness implements exactly this: it focuses the box, clears it, tries
1240
+ `type_text(q)` (the non-focus-gated CDP path), reads the value back, and **falls back** to a direct
1241
+ `js` set with `input`/`change` events if the read-back does not match. It returns the value actually read
1242
+ back, and `q_ok = (qgot == c["q"])`. The assertion is on the *read-back*, not on the *attempt*.
1243
+
1244
+ `obs_ok`'s second half is a real DOM fact: `#obsTail` reads `none` in the markup (`mission.html:67`) and
1245
+ the live driver flips it to `ready` on upload, so `obs_ok` proves the upload landed rather than merely
1246
+ that a click happened.
1247
+
1248
+ ### 8.5 The three discriminators that proved the earlier 8/8 was clean
1249
+
1250
+ The response to a silent-false-pass risk is not to re-run and hope; it is to check whether the *earlier*
1251
+ result was contaminated, using evidence the harness recorded. The report does exactly that:
1252
+
1253
+ > *"I then checked whether the earlier 8/8 run was infected by the same silent failure. It was not:*
1254
+ >
1255
+ > * *its recorded intents are **query-specific** — A1 reads `taskvqa…temporalnone`, whereas the default
1256
+ > query "What changed here?" would read `taskchange…temporalrequired` (exactly what the failed run
1257
+ > showed);*
1258
+ > * *its answers **embed the query text** — e.g. `[grounding] Located 6 candidate region(s) for 'Where
1259
+ > are the built-up areas in this image?'`;*
1260
+ > * *A6 required two files (`optical 4/12 + SAR 2/2` channels), which only the uploaded pair supplies.*
1261
+ >
1262
+ > *So the 8/8 result is a valid measurement."*
1263
+ > — `DELIVERY_REPORT_2026-09-25.md` §3
1264
+
1265
+ The three discriminators generalise into a reusable checklist:
1266
+
1267
+ | Discriminator | What it proves |
1268
+ |---|---|
1269
+ | the recorded **intent** is query-specific | the query reached the router |
1270
+ | the **answer** embeds the query text | the server received the intended query |
1271
+ | a case **requires an artefact** only the setup supplies | the setup really happened |
1272
+
1273
+ This is why the harness records the full per-case payload — `intent`, `answer`, `files_t1`, `files_t0`,
1274
+ `obs_tail`, `mock_nodes`, `api_calls` — rather than just a pass/fail bit. The bit can be wrong; the raw
1275
+ record can be re-examined.
1276
+
1277
+ ### 8.6 The two discriminator bugs (and why they were false *failures*, not false passes)
1278
+
1279
+ Two harness bugs were found and fixed, both causing **false failures** — the safe direction:
1280
+
1281
+ 1. **The answer `[task]` tag exists only for region tasks.** `vqa`/`caption` answers are bare
1282
+ ("Grassland", prose), so a tag-based verdict heuristic reported them as failures. The fix computes the
1283
+ dispatched task as `answer_tag` when present, else the intent's reading.
1284
+ 2. **The intent panel is a concatenated string.** `task([a-z_]+)` must be matched **non-greedily** up to
1285
+ `modality`; a greedy match swallows the whole string. The fix is
1286
+ `re.search(r"task([a-z_]+?)modality", intent)`.
1287
+
1288
+ The harness header states the asymmetry: *"Two further harness bugs were found and fixed, both causing
1289
+ **false failures**."* A false failure costs a re-run; a false pass costs the integrity of the result. The
1290
+ harness was built so that its bugs err toward false failure, and §9 makes even the false-failure case
1291
+ recoverable without a re-run.
1292
+
1293
+ ### 8.7 Liveness traps discovered during the runs
1294
+
1295
+ Two operational traps are recorded because they look exactly like stalls:
1296
+
1297
+ - **`browser-use` block-buffers stdout even when redirected to a file**, so the output file stays at
1298
+ **0 bytes until the process exits** — indistinguishable from a stall. The harness now forces
1299
+ line-buffering (`sys.stdout.reconfigure(line_buffering=True)`) and prints `CASE_START <id>` per case,
1300
+ so the file grows case by case.
1301
+ - **`grep` block-buffers when piped**, so piping the harness through `grep` swallows all output if the
1302
+ pipeline is killed. Redirect to a file instead.
1303
+
1304
+ The harness also records `ALL_DONE` as a final marker, so a truncated run is detectable from the output
1305
+ file alone.
1306
+
1307
+ ### 8.8 What the live validation does NOT establish
1308
+
1309
+ - It is **not** the independent audit. `LIVE_VALIDATION_POSTFIX.md` opens by saying so: *"This is a
1310
+ *re-validation*, not the independent audit: the prior independent audit is
1311
+ `../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md`."*
1312
+ - It does **not** establish model *quality*. The two verdicts are kept separate: **Deployment/integration:
1313
+ PASS** and **Model quality: MIXED** (caption and grounding meaningful; change/change_vqa plausible; VQA
1314
+ weak-but-related — A1 answers "Grassland"; optical-SAR returns a bare class index `class_18`, not a
1315
+ human label).
1316
+ - It does **not** establish load behaviour, latency ceilings, or cold-start times. It is eight cases, run
1317
+ three times.
1318
+ - It does **not** clear B-07: transient tunnel-agent gaps remain **OPEN**, and the patch is **prepared,
1319
+ NOT deployed** (`CURRENT_RELEASE_STATE.md` §6).
1320
+
1321
+ ---
1322
+
1323
+ ## 9. Verdict-independent evidence: recomputing from raw records
1324
+
1325
+ ### 9.1 The property
1326
+
1327
+ The live-validation harness has a `verdict` field. The project's design decision is that **the verdict
1328
+ must not be the primary record.** The recorded evidence — `q_ok`, `obs_ok`, `t0_ok`, `run_id`,
1329
+ `mock_nodes`, `intent`, `answer`, `api_calls` — is recorded *independently of* the verdict computation,
1330
+ so a harness bug in the verdict logic can never silently turn a real failure into a pass, and can be
1331
+ corrected without a re-run.
1332
+
1333
+ ### 9.2 The recomputation script
1334
+
1335
+ `.workbuddy-ai/scratch/recompute_verdicts.py` implements the property. Its docstring states the case that
1336
+ motivated it:
1337
+
1338
+ > *"The v2 harness recorded, per case: `q_ok`, `obs_ok`, `t0_ok`, `run_id`, `mock_nodes`, `intent`,
1339
+ > `answer`, `api_calls` -- all independent of the verdict computation. Its verdict field was wrong for
1340
+ > vqa/caption because the tag heuristic assumed every answer starts with `"[task]"`; vqa/caption answers
1341
+ > are bare. This recomputes the verdict from the intent's `task<name>` slot (primary) plus the answer tag
1342
+ > (cross-check only when present)."*
1343
+
1344
+ The script is read-only over the recorded run and writes `results_final.json`. The core of it:
1345
+
1346
+ ```python
1347
+ mi = re.search(r"task([a-z_]+?)modality", intent)
1348
+ intent_task = mi.group(1) if mi else ""
1349
+ tag = answer[1:answer.find("]")] if (answer.startswith("[") and "]" in answer) else ""
1350
+ expect = d["expect"]
1351
+
1352
+ # The DISPATCHED task is what the server actually ran. The answer carries a "[task]"
1353
+ # prefix for the region tasks (grounding/change/change_vqa/optical_sar); vqa and
1354
+ # caption answers are bare, so fall back to the intent's reading there.
1355
+ # NOTE: A5 is a legitimate case where the two differ -- the router READS "change" but
1356
+ # dispatches "change_vqa" (the documented quantifier upgrade). That is not a failure.
1357
+ dispatched = tag if tag else intent_task
1358
+
1359
+ checks = {
1360
+ "query_in_box": bool(d.get("q_ok")),
1361
+ "asset_attached": bool(d.get("obs_ok")),
1362
+ "pair_attached": bool(d.get("t0_ok")),
1363
+ "live_run": str(d.get("run_id", "")).startswith("run_"),
1364
+ "no_mock_nodes": d.get("mock_nodes") == 0,
1365
+ "task_dispatched": dispatched == expect,
1366
+ }
1367
+ verdict = "PASS" if all(checks.values()) else "FAIL"
1368
+ ```
1369
+
1370
+ Every input to the verdict is a *recorded observation*; the verdict is a pure function of them. The
1371
+ script also records `read_vs_dispatched_differs` — so the one legitimate read≠dispatch case (A5, the
1372
+ documented quantifier upgrade) is **flagged, not failed**.
1373
+
1374
+ ### 9.3 Pass 3 is the proof of the property
1375
+
1376
+ Pass 3 was launched with the harness build that still carried the two discriminator bugs (§8.6), so its
1377
+ raw output says **`SUMMARY 0/8`**. The verdicts were recomputed from the recorded evidence by
1378
+ `recompute_verdicts.py` and are **8/8 PASS** (`results_pass3.json`). The record states the significance:
1379
+
1380
+ > *"That is the intended safety property — **the recorded evidence (run id, `mock_nodes`, intent, answer,
1381
+ > assertions) is independent of the verdict computation**, so a harness bug can never silently turn a
1382
+ > real failure into a pass, and never costs a re-run to correct."*
1383
+ > — `LIVE_VALIDATION_POSTFIX.md`
1384
+
1385
+ This is the strongest form of the claim: the property was not merely designed, it was **exercised**. A
1386
+ run with a broken verdict computation was corrected *post hoc* from its own raw record, without touching
1387
+ the deployment or re-running the browser.
1388
+
1389
+ ### 9.4 The generalisation
1390
+
1391
+ The pattern is: **record observations, compute verdicts separately, keep the observations.** It is the
1392
+ same shape as the evidence engine's digest (§4.8): the *record* is content-addressed and process-independent,
1393
+ and the *interpretation* is a function over it. Both make the same guarantee — that a conclusion can be
1394
+ re-derived from the evidence rather than trusted as an assertion.
1395
+
1396
+ ---
1397
+
1398
+ ## 10. How to run the tests
1399
+
1400
+ ### 10.1 The interpreter
1401
+
1402
+ Use the repository virtual environment, not a bare `python`. On Windows:
1403
+
1404
+ ```bash
1405
+ .venv/Scripts/python.exe -m pytest <target> -q
1406
+ ```
1407
+
1408
+ `docs/REPRODUCIBILITY.md` §10.2 records this as the invocation that reproduces the published results. On
1409
+ Windows, pytest must be given `-p no:cacheprovider` because the sandbox refuses `.pytest_cache` writes
1410
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §0).
1411
+
1412
+ ### 10.2 Run targeted files, not the whole tree
1413
+
1414
+ **A full `tests/unit` run trips the sandbox's bulk-delete guard** (the 4× `test_safe_delete_shim`
1415
+ failures of §7.3). The workaround is to run targeted suites:
1416
+
1417
+ ```bash
1418
+ .venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
1419
+ .venv/Scripts/python.exe -m pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp -q
1420
+ .venv/Scripts/python.exe -m pytest tests/unit/test_deploy_config.py -p no:cacheprovider -q
1421
+ .venv/Scripts/python.exe -m pytest tests/unit/test_api_contract_doc.py tests/unit/test_frontend_guide_doc.py -p no:cacheprovider -q
1422
+ ```
1423
+
1424
+ The runbook notes that **multi-suite invocations in one command have been refused by the environment
1425
+ before; single suites are reliable** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §2.6). The `--basetemp`
1426
+ flag is **mandatory** for the evidence-engine suite (§4.7): the default pytest temp root triggers a
1427
+ sandbox denial on this machine, and the module's `scratch` fixture writes to a repo-local `.pytest_tmp`
1428
+ instead.
1429
+
1430
+ ### 10.3 `pytest.ini`
1431
+
1432
+ The configuration is four lines plus the marker registrations:
1433
+
1434
+ ```ini
1435
+ [pytest]
1436
+ testpaths = tests
1437
+ pythonpath = .
1438
+ addopts = -q --tb=short
1439
+ ```
1440
+
1441
+ (`pytest.ini:1-4`). `pythonpath = .` is what makes `from core.config import ...` resolve from the repo
1442
+ root without an installed package — and `test_deploy_config.py` deliberately proves the module *also*
1443
+ bootstraps `sys.path` itself, *"WITHOUT pytest's `pythonpath = .` shortcut"* (§6.4). The markers are
1444
+ registered at `pytest.ini:23-27` (§0.1), and `filterwarnings` silences three warning classes including
1445
+ `rasterio.errors.NotGeoreferencedWarning` (`pytest.ini:5-8`).
1446
+
1447
+ ### 10.4 The output-buffering trap
1448
+
1449
+ **Do not pipe pytest through `grep`** in the authoring sandbox — output is block-buffered and a killed
1450
+ pipeline swallows it. Redirect to a file instead (`docs/REPRODUCIBILITY.md` §5.6). The same trap applies
1451
+ to the live-validation harness (§8.7).
1452
+
1453
+ ### 10.5 The expected results
1454
+
1455
+ | Suite | Command | Expected | Status |
1456
+ |---|---|---|---|
1457
+ | Frontend live-wiring | `pytest tests/unit/test_frontend_live_wiring.py -q` | **106 passed** | `VERIFIED` |
1458
+ | Doc/frontend suite | `pytest` on the 5 doc/frontend files | **183 passed** | `VERIFIED` |
1459
+ | Evidence engine | `pytest tests/unit/test_evidence_engine.py --basetemp=.pytest_tmp` | **73 passed** | `VERIFIED` |
1460
+ | Gateway policy | `pytest tests/unit/test_gateway_policy.py` | **51 passed** | `VERIFIED` |
1461
+ | Full unit suite | `pytest tests/unit` | **5–6 environmental/ordering failures**, rest pass; a re-run passes **137** | `MEASURED` |
1462
+
1463
+ (`release/repo/README.md` §The test suites; `docs/PHASE19_FINAL_HARDENING.md` §6;
1464
+ `docs/REPRODUCIBILITY.md` §5.1.) The precise full-suite **collected** count is
1465
+ `UNKNOWN — not established from the available evidence` (§7.7).
1466
+
1467
+ ---
1468
+
1469
+ ## 11. What is NOT tested
1470
+
1471
+ A green suite is not evidence about a property nobody wrote an assertion for (`DOCS_STYLE_GUIDE.md` §1.6;
1472
+ §0.3). This section lists the properties the release does **not** have test evidence for, so that no
1473
+ reader mistakes coverage for completeness.
1474
+
1475
+ | Property | State | Why / what exists instead |
1476
+ |---|---|---|
1477
+ | **Continuous integration** | **does not exist** | there is **no `.github/` directory** and no CI workflow in the repository. Every test result in this chapter was produced by a **manual** run in the authoring sandbox. Nothing runs the suites automatically on push. |
1478
+ | **End-to-end system benchmark** | **NOT RUN (none exists)** | `CURRENT_RELEASE_STATE.md` §4: *"System-level end-to-end benchmark — **NOT RUN (none exists)**."* No system-level accuracy is claimed anywhere in the release. |
1479
+ | **Load / soak / concurrency** | **NOT RUN** | no test exercises sustained load, concurrent users, or a long-running process. The per-IP rate limiter is a **fairness** control, not a security or capacity control (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2). |
1480
+ | **Adversarial / fuzz testing** | **NOT RUN** | no fuzzing or adversarial-input suite. Input validation is tested by example (§2.2.9), not by search. |
1481
+ | **Cross-dataset generalization** | **NOT RUN** | the measured metrics are per-dataset (LEVIR-CD-256, VRSBench, held-out test sets); no test measures transfer to an unseen dataset. |
1482
+ | **Browser / device matrix** | **NOT RUN** | live validation ran in **headed Chromium** only. No Firefox/Safari/WebKit, no mobile, no accessibility audit. |
1483
+ | **Deployment verification in CI** | **NOT RUN** | the runbook guard asserts the document does **not** claim a deployment happened (§3.4). Deployment reachability is established only by the manual live validation (§8). |
1484
+ | **Router test split** | **NOT RUN** | `CURRENT_RELEASE_STATE.md` §4: router accuracy `0.965116` is **validation, ungated, n = 86**; *"the test split was NOT RUN"* (`DOCS_STYLE_GUIDE.md` §3). |
1485
+ | **Calibration as an improvement** | **not established (and measured the other way)** | ECE went `0.013755 → 0.014929` — **worse**. The engine suite pins that calibration is *honestly labelled* (§4.5), not that it helps. |
1486
+ | **Optical-SAR / change-VQA acceptance** | **OPEN** | the rulings are OPEN (`CURRENT_RELEASE_STATE.md` §4; `DOCS_STYLE_GUIDE.md` §3). A suite passing is not a model acceptance. |
1487
+ | **VLM adapter acceptance** | **REJECTED** | metrics `USABLE_VERIFIED` (exact_match 0.963) but status **ACCEPTANCE-REJECTED**. USABLE ≠ ACCEPTED (`DOCS_STYLE_GUIDE.md` §3). |
1488
+ | **Tunnel reliability (B-07)** | **OPEN** | transient tunnel-agent gaps; patch **prepared, NOT deployed**. No test can cover a defect that is not fixed in production. |
1489
+ | **The inference service's internals** | **UNKNOWN** | the tunnel-agent source is not in the monorepo; `docs/architecture/02-deployment-topology.md` §3.3/§9.4 documents the topology as a **measured** fact, but the agent's internals are `UNKNOWN — not established from the available evidence`. |
1490
+ | **Credential write permission (HF)** | **UNKNOWN** | `CURRENT_RELEASE_STATE.md` §7: HF token identity `thundercode`, role `fineGrained`; *"write permission not yet proven."* |
1491
+
1492
+ Two of these deserve emphasis:
1493
+
1494
+ - **No CI is the largest gap.** Every count in this chapter is a manual measurement. A reader should not
1495
+ read "106 passed" as "the suite passes on every commit" — it means "it passed when it was run, in the
1496
+ recorded environment." There is no automation that would catch a regression on push.
1497
+ - **No end-to-end benchmark is the second.** The individual task metrics are artifact-backed and
1498
+ measured, but there is no measurement of the *system* — router + specialist + envelope + frontend — on
1499
+ a held-out end-to-end corpus. The live validation (§8) proves the *pipeline runs and dispatches
1500
+ correctly*; it does not measure system accuracy.
1501
+
1502
+ ---
1503
+
1504
+ ## 12. Open items and evidence index
1505
+
1506
+ ### 12.1 Open / blocked / not-run items for this chapter's topic
1507
+
1508
+ Following the style guide's requirement that every doc end with an explicit list
1509
+ (`DOCS_STYLE_GUIDE.md` §4):
1510
+
1511
+ | Item | State | Note |
1512
+ |---|---|---|
1513
+ | B-07 — transient tunnel-agent gaps | **OPEN** | patch **prepared, NOT deployed**; root shape measured (in `auto` mode a tunnel timeout falls through to the forward path, burning `wake_timeout_s=120` on a 302 ≈ 249 s) |
1514
+ | B-02 — `codespace_name` trailing `\n` | **OPEN (cosmetic)** | confirmed still live; the wake path strips it |
1515
+ | No CI | **does not exist** | no `.github/` directory; all results are manual |
1516
+ | System-level end-to-end benchmark | **NOT RUN (none exists)** | no system-level accuracy claimed |
1517
+ | Router test split | **NOT RUN** | val-only, n = 86 |
1518
+ | Full-suite collected test count | **UNKNOWN** | per-suite counts known; the single collected total is not |
1519
+ | `test_safe_delete_shim` failures | **environmental** | sandbox delete-guard; not regressions; precise Windows trigger **UNKNOWN** |
1520
+ | Ordering flake | **environmental** | passes in isolation |
1521
+ | Stale adapter test | **stale** | asserts CROMA unshipped; CROMA is shipped |
1522
+ | Live validation | **MEASURED** | 3 passes × 8 cases, 8/8 each, 24 runs, 0 mock nodes, trace fill 94.4444 % |
1523
+ | Optical-SAR / change-VQA rulings | **OPEN** | 0.931 / macro-F1 0.434161; 0.697626 / 0.378373 |
1524
+ | VLM adapter acceptance | **REJECTED** | metrics usable; acceptance rejected |
1525
+ | Calibration | **not an improvement** | ECE worse (0.013755 → 0.014929); retained only because it is in the frozen config |
1526
+
1527
+ ### 12.2 Where the evidence lives
1528
+
1529
+ | What | Where |
1530
+ |---|---|
1531
+ | Test configuration and markers | `pytest.ini` |
1532
+ | Frontend live-wiring suite | `tests/unit/test_frontend_live_wiring.py` — **106 passed** |
1533
+ | Evidence-engine suite | `tests/unit/test_evidence_engine.py` — **73 passed** |
1534
+ | Safe-delete shim guard | `tests/unit/test_safe_delete_shim.py` — 338 lines; skips cleanly off-sandbox |
1535
+ | Doc-guards | `tests/unit/test_api_contract_doc.py`, `test_frontend_guide_doc.py`, `test_runbook_doc.py`, `test_deploy_config.py`, `test_step7_api_contract_conformance.py` |
1536
+ | Integration suite scope | `tests/integration/test_step7_backend_chain.py`, `test_step7_r02_serving.py`; `docs/ITEM5_INTEGRATION_SUITE_SCOPE.md` |
1537
+ | The suite results, reported in full | `docs/EVALUATION.md` §7.1–7.5; `docs/REPRODUCIBILITY.md` §5; `release/repo/README.md` §The test suites |
1538
+ | The full-suite failure classification | `docs/FINAL_DELIVERY_REPORT.md` §7; `docs/PHASE12_115_METRIC_COMPUTED.md` §8 (the 2,237-passed run) |
1539
+ | Live-validation record | `.workbuddy-ai/scratch/live_validation/LIVE_VALIDATION_POSTFIX.md` |
1540
+ | Live-validation raw output | `.workbuddy-ai/scratch/live_validation/run_output.txt`, `run_final2.txt`, `run_final3.txt` |
1541
+ | Live-validation screenshots | `.workbuddy-ai/scratch/live_validation/A1_vqa.png` … `B2_newairport.png` (8) |
1542
+ | Live-validation harness (v1) | `.workbuddy-ai/scratch/run_all_postfix.harness` |
1543
+ | Live-validation harness (v2, asserting) | `.workbuddy-ai/scratch/run_all_postfix2.harness` |
1544
+ | CDP focus diagnostic | `.workbuddy-ai/scratch/diag_focus.harness` |
1545
+ | Verdict recomputation | `.workbuddy-ai/scratch/recompute_verdicts.py` |
1546
+ | Prior independent audit | `../2026-09-25-00-37-41/live_validation/INDEPENDENT_LIVE_VALIDATION.md` |
1547
+ | Release-state inventory | `release/CURRENT_RELEASE_STATE.md` |
1548
+
1549
+ ### 12.3 Cross-references
1550
+
1551
+ - The trust model, secret custody, CORS allowlist, size limits and rate-limiting-as-fairness that the
1552
+ frontend suite asserts against: [`SECURITY.md`](SECURITY.md).
1553
+ - The API contract the doc-guards validate: [`API_CONTRACT.md`](API_CONTRACT.md).
1554
+ - The deployment topology the live validation drives: [`DEPLOYMENT.md`](DEPLOYMENT.md),
1555
+ [`architecture/02-deployment-topology.md`](architecture/02-deployment-topology.md).
1556
+ - The measured metrics this chapter does **not** re-derive: [`EVALUATION.md`](EVALUATION.md).
1557
+ - The reproduction commands and their expected outputs: [`REPRODUCIBILITY.md`](REPRODUCIBILITY.md) §5.
1558
+ - The router defect the live validation re-validates: [`architecture/04-router.md`](architecture/04-router.md),
1559
+ [`architecture/09-frontend.md`](architecture/09-frontend.md) §10.