thundercode commited on
Commit
280cc90
Β·
verified Β·
1 Parent(s): af0705e

release: add docs/DEVELOPMENT.md

Browse files
Files changed (1) hide show
  1. docs/DEVELOPMENT.md +1330 -0
docs/DEVELOPMENT.md ADDED
@@ -0,0 +1,1330 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Development Guide
2
+
3
+ **Status tags:** `IMPLEMENTED` Β· `VERIFIED` Β· `MEASURED` Β· `ATTEMPTED` Β· `NOT RUN` Β· `BLOCKED` Β·
4
+ `DEFERRED` Β· `OPEN` Β· `RESOLVED` Β· `BY DESIGN`.
5
+
6
+ This is the contributor-facing guide to the SatQuery AI codebase. It answers: *what do I need to
7
+ install*; *how do I run it*; *what is each directory for*; *what rules will break the build if I break
8
+ them*; *how do I add a component*; *how do I test*; and *which environment traps will cost me an hour if
9
+ nobody tells me about them*.
10
+
11
+ Every command, path, timeout and rule below comes from a file that was read. Where the evidence does not
12
+ settle a question, the text says `UNKNOWN β€” not established from the available evidence` rather than
13
+ guessing.
14
+
15
+ > **Read this first.** `configs/base.yaml` is the single registry, and its hash β€” **`78f1e3700da15aa1`**
16
+ > β€” is frozen. **Editing the config moves the hash and invalidates every artifact keyed to it.** This is
17
+ > the one rule in the repository whose violation is not locally reversible (Β§7).
18
+
19
+ ---
20
+
21
+ ## Table of contents
22
+
23
+ **Part I β€” Prerequisites and setup**
24
+ 1. Prerequisites
25
+ 2. Local setup
26
+ 3. Running the service locally
27
+
28
+ **Part II β€” Repository layout**
29
+ 4. The top-level directory map
30
+ 5. Where the contracts live
31
+ 6. The package layout inside `core/`, `specialists/`, `app/`
32
+
33
+ **Part III β€” The configuration discipline**
34
+ 7. "No magic numbers in Python"
35
+ 8. What happens if you break the rule
36
+ 9. The hash-exempt path: environment overrides
37
+ 10. The invariants the loader enforces
38
+
39
+ **Part IV β€” Contract-first development**
40
+ 11. `core/schemas.py` is the binding contract
41
+ 12. The specialist interface
42
+ 13. How to add a new specialist
43
+ 14. How evidence is emitted
44
+
45
+ **Part V β€” The test workflow**
46
+ 15. Running tests
47
+ 16. The evidence-class markers
48
+ 17. The known failures, and why they are not regressions
49
+
50
+ **Part VI β€” Scripts and notebooks**
51
+ 18. The class of scripts under `scripts/`
52
+ 19. `notebooks/`
53
+
54
+ **Part VII β€” Training entry points**
55
+ 20. Local training (router, grounding, change, optical-SAR)
56
+ 21. External GPU training (change-VQA, VLM LoRA)
57
+ 22. Calibration
58
+
59
+ **Part VIII β€” Coding conventions**
60
+ 23. What the code actually does
61
+
62
+ **Part IX β€” Known development traps**
63
+ 24. The traps, in one table
64
+ 25. Trap 1 β€” the stale untracked `deploy/`
65
+ 26. Trap 2 β€” the dead sandbox proxy
66
+ 27. Trap 3 β€” pytest only in the venv
67
+ 28. Trap 4 β€” Chrome drops synthetic CDP key events
68
+ 29. Trap 5 β€” Cloudflare `_headers` concatenate
69
+ 30. Trap 6 β€” the annotation-scope trap
70
+ 31. Trap 7 β€” the full-suite bulk-delete guard
71
+ 32. Trap 8 β€” the stale serve process
72
+
73
+ **Part X β€” Status and evidence**
74
+ 33. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for development
75
+ 34. Where the evidence lives
76
+
77
+ ---
78
+
79
+ # Part I β€” Prerequisites and setup
80
+
81
+ ## 1. Prerequisites
82
+
83
+ | Requirement | Value | Source |
84
+ |---|---|---|
85
+ | Python | **3.11+** | `release/repo/README.md` Β§Installation |
86
+ | Device | **a CPU is sufficient**; no CUDA requirement | `release/repo/README.md` Β§Installation |
87
+ | GPU | not required | `docs/DEPLOYMENT_DECISION.md` Β§5 |
88
+ | Git | needed to clone | `release/repo/README.md` Β§Installation |
89
+
90
+ The codebase is CPU-first and the deployment is CPU-only. This is not a fallback β€” it is the design:
91
+
92
+ - `core/config.py:86-91` β€” `device_preference` honours the `SATQUERY_DEVICE` env override, else
93
+ `"cuda" if torch.cuda.is_available() else "cpu"`.
94
+ - Every specialist defaults to `device="cpu"` (`docs/DEPLOYMENT_DECISION.md` Β§5).
95
+ - **No `.cuda()` call exists anywhere**; all placement is `.to(device)` (`docs/DEPLOYMENT_DECISION.md`
96
+ Β§5).
97
+ - `configs/base.yaml:293` β€” `cpu_mode_required: true`.
98
+
99
+ The development container that runs the inference tier pins Python 3.12
100
+ (`.devcontainer/devcontainer.json:3`, `image: mcr.microsoft.com/devcontainers/python:3.12`). The
101
+ repository's own guidance is 3.11+; a 3.11 or 3.12 interpreter both work.
102
+
103
+ **Not implemented** (noted, not needed): no thread capping (`torch.set_num_threads`) and no
104
+ quantisation. `lazy_load` and `cache_max_models` are only *reported* (`app/deployment.py:609-610`,
105
+ `:896-897`); **UNVERIFIED** whether any code enforces a one-model cache (`docs/DEPLOYMENT_DECISION.md`
106
+ Β§5).
107
+
108
+ ## 2. Local setup
109
+
110
+ ```bash
111
+ git clone https://github.com/Anish-lab-blip/SatQuery-AI
112
+ cd SatQuery-AI
113
+ python -m venv .venv
114
+ source .venv/Scripts/activate # Windows git-bash; use .venv/bin/activate on Linux/macOS
115
+ pip install -r requirements.txt
116
+ ```
117
+
118
+ (`release/repo/README.md` Β§Installation)
119
+
120
+ ### 2.1 The dependency manifest and its two profiles
121
+
122
+ `requirements.txt` is **frozen at architecture v1.0** and declares two install profiles in its header
123
+ comment (`requirements.txt:1-8`):
124
+
125
+ ```
126
+ # Two install profiles:
127
+ # CPU (local dev / schema / geo / unit tests):
128
+ # pip install -r requirements.txt
129
+ # GPU (Kaggle T4x2 / HF ZeroGPU): torch is preinstalled on both.
130
+ # Do NOT pin torch here β€” Kaggle and HF ship their own builds.
131
+ ```
132
+
133
+ **torch is deliberately not pinned** (`requirements.txt:21`). The platform supplies it. This is why the
134
+ local CPU install works on a machine with no CUDA and why the Kaggle/ZeroGPU targets get their own
135
+ builds.
136
+
137
+ The manifest's sections, with the contract notes the file itself carries:
138
+
139
+ | Section | Packages | Note |
140
+ |---|---|---|
141
+ | core | `numpy`, `pyyaml`, `pydantic` | β€” |
142
+ | geospatial | `rasterio`, `pyproj`, `opencv-python-headless` | β€” |
143
+ | models | `transformers>=4.52`, `open-clip-torch>=2.24`, `sentence-transformers>=2.7`, `peft>=0.10`, `huggingface_hub>=0.23`, `safetensors>=0.4`, `einops>=0.7` | see below |
144
+ | ui + reporting | `gradio>=4.44`, `reportlab>=4.1` | β€” |
145
+ | dev | `pytest>=8.0`, `pytest-cov>=5.0` | β€” |
146
+
147
+ Three contract notes in the file are load-bearing (`requirements.txt:43-53`):
148
+
149
+ - `transformers>=4.52` β€” SmolVLM via `AutoModelForImageTextToText`; `AutoModelForVision2Seq` is
150
+ deprecated (finding C-2).
151
+ - `open-clip-torch` β€” the RemoteCLIP checkpoint is loaded via `pretrained=<path>` so
152
+ `load_checkpoint()` runs its state-dict fixups (C-4).
153
+ - `huggingface_hub` β€” used to pin the RemoteCLIP / SmolVLM / CROMA revisions.
154
+
155
+ > **The `einops` lesson.** `einops>=0.7` is **not optional**. It is required by the vendored
156
+ > `specialists/optical_sar/vendor/use_croma.py` (`from einops import rearrange`). It was found missing on
157
+ > 2026-09-18 by executing the vendored module: it raised `ModuleNotFoundError`, so the CROMA path was
158
+ > blocked on a dependency that no document declared (`requirements.txt:28-33`).
159
+
160
+ ### 2.2 Device selection
161
+
162
+ Set `SATQUERY_DEVICE` to force a device:
163
+
164
+ ```bash
165
+ export SATQUERY_DEVICE=cpu # cpu | cuda | mps | null
166
+ ```
167
+
168
+ `device_preference` reads the env var first and only falls back to a torch probe if it is unset
169
+ (`core/config.py:87-91`). The value is read **without importing torch** on the metadata path, which is
170
+ what keeps the health route cheap (`docs/DEPLOYMENT_TOPOLOGY.md` Β§3.3).
171
+
172
+ ## 3. Running the service locally
173
+
174
+ The inference service is served by the same launcher the Codespace runs
175
+ (`release/repo/README.md` Β§Local development):
176
+
177
+ ```bash
178
+ # Inference service, CPU (this is the launcher the Codespace runs)
179
+ PORT=8000 python deploy/codespace/serve.py
180
+
181
+ # Health
182
+ curl localhost:8000/v1/health
183
+ ```
184
+
185
+ `serve.py` builds the app through `build_space_app()` and binds it with uvicorn on `$PORT` (default 8000)
186
+ (`deploy/codespace/serve.py:16-25`). It imports cheaply β€” FastAPI is imported inside `build_space_app()`
187
+ and no model is loaded at module scope β€” so it stays import-safe on a CPU host with no GPU and no weights
188
+ present (`deploy/codespace/serve.py:1-13`).
189
+
190
+ The server answers four routes (`app/space_app.py`):
191
+
192
+ | Route | Line | Purpose |
193
+ |---|---|---|
194
+ | `GET /v1/health` | `app/space_app.py:521` | liveness, derived status, torch-free device probe |
195
+ | `GET /v1/capabilities` | `app/space_app.py:549` | the six declared capabilities |
196
+ | `POST /v1/assets` | `app/space_app.py:555` | handle-based upload (off unless enabled) |
197
+ | `POST /v1/analyze` | `app/space_app.py:661` | the analysis path |
198
+
199
+ `POST /v1/assets` is **off unless explicitly enabled** (`SATQUERY_ASSET_ENABLED` **and**
200
+ `SATQUERY_ASSET_DIR` must both be set; otherwise it answers `503` rather than defaulting to a temp
201
+ directory) (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§3.1.1).
202
+
203
+ ### 3.1 Serving the frontend locally
204
+
205
+ The frontend is fully static and needs no build step (`release/repo/README.md` Β§Local development):
206
+
207
+ ```bash
208
+ python -m http.server 5500 --directory frontend
209
+ ```
210
+
211
+ The gateway's dev-origin allowlist includes `localhost` and `127.0.0.1` on ports 3000, 5500, 5173, 8000
212
+ and 8080 (`deploy/render/main.py:139-143`), so a local dev server on any of those ports is accepted
213
+ without an env-var change.
214
+
215
+ ### 3.2 Running the gateway locally
216
+
217
+ ```bash
218
+ PORT=10000 \
219
+ GITHUB_TOKEN=<token> \
220
+ CODESPACE_NAME=<codespace> \
221
+ SATQUERY_ALLOWED_ORIGINS="https://satquery.pages.dev" \
222
+ uvicorn deploy.render.main:app --port 10000
223
+ ```
224
+
225
+ (`deploy/render/README.md` Β§Running locally β€” the placeholder is the repository's own.)
226
+
227
+ Then exercise it:
228
+
229
+ ```bash
230
+ curl http://localhost:10000/api/health
231
+ curl -X POST http://localhost:10000/api/infer -H 'content-type: application/json' -d '{"query":"..."}'
232
+ curl http://localhost:10000/api/capabilities
233
+ ```
234
+
235
+ > **The monorepo `deploy/` is stale** (Β§25). Use it for local experimentation only; it is **not** the
236
+ > deployed source and it lacks the tunnel code the live service runs.
237
+
238
+ ---
239
+
240
+ # Part II β€” Repository layout
241
+
242
+ ## 4. The top-level directory map
243
+
244
+ | Directory / file | Purpose | Source |
245
+ |---|---|---|
246
+ | `app/` | the serving composition root and the FastAPI entrypoint | `app/serving.py`, `app/space_app.py`, `app/deployment.py` |
247
+ | `core/` | the config loader, the typed schemas, the controller, the planner, the registry, the error taxonomy | `core/config.py`, `core/schemas.py`, `core/controller.py`, `core/planner.py`, `core/registry.py`, `core/errors.py`, `core/code_revision.py` |
248
+ | `router/` | the MiniLM intent router: encoder, classifier, adapter, training, label space, lexical fallback | `router/encoder.py`, `router/classifier.py`, `router/adapter.py`, `router/train.py`, `router/dataset.py`, `router/fallback.py`, `router/label_space.py` |
249
+ | `specialists/` | the six specialist implementations behind one interface | `specialists/base.py`, `specialists/vqa/`, `specialists/grounding/`, `specialists/change/`, `specialists/optical_sar/` |
250
+ | `preprocessing/` | raster loading, imagery, quality checks | `preprocessing/raster.py`, `preprocessing/imagery.py`, `preprocessing/quality.py` |
251
+ | `geospatial/` | CRS handling and transforms | `geospatial/crs.py`, `geospatial/transform.py` |
252
+ | `evidence/` | the evidence engine and confidence | `evidence/engine.py`, `evidence/confidence.py` |
253
+ | `evaluation/` | manifests, leakage, metrics, normalisation, the runner, benchmark adapters, the frozen prompt set | `evaluation/manifests.py`, `evaluation/leakage.py`, `evaluation/metrics/`, `evaluation/normalize.py`, `evaluation/runner.py`, `evaluation/benchmark_adapters/`, `evaluation/prompt_freeze.json`, `evaluation/manifest_freeze.json` |
254
+ | `training/` | the training loops and data adapters, split by task | `training/router/`, `training/grounding/`, `training/change/`, `training/change_vqa/`, `training/fusion/`, `training/vlm/`, `training/calibration/`, `training/data/` |
255
+ | `configs/` | the frozen registry and the frozen deploy manifest | `configs/base.yaml`, `configs/deploy.yaml` |
256
+ | `gateway/` | the standalone gateway: policy (decisions) + app (HTTP plumbing) | `gateway/policy.py`, `gateway/app.py`, `gateway/assets.py` |
257
+ | `deploy/` | the deployment launchers β€” **stale and untracked in the monorepo** | `deploy/codespace/`, `deploy/render/` (Β§25) |
258
+ | `tests/` | the suites, split by concern and evidence class | `tests/unit/`, `tests/integration/`, `tests/routing/`, `tests/geospatial/`, `tests/leakage/`, `tests/model/`, `tests/e2e/` |
259
+ | `scripts/` | the flat set of executable helpers | Β§18 |
260
+ | `notebooks/` | the Kaggle notebooks | Β§19 |
261
+ | `artifacts/` | trained heads, checkpoints, caches, evidence archives | ~3.7 GB total (`release/CURRENT_RELEASE_STATE.md` Β§3) |
262
+ | `frontend/` | the static site | staged by `scripts/stage_pages.mjs` |
263
+ | `docs/` | the project's own engineering records | ~60 files |
264
+ | `hf/` | the Hugging Face project card / model cards | β€” |
265
+ | `demo/`, `benchmark/`, `reports/`, `data/`, `logs/` | supporting material | β€” |
266
+
267
+ > **`training/router/` is empty** in the working copy; the router's training lives in `router/train.py`
268
+ > and `router/dataset.py` and is driven by `scripts/train_router.py` (Β§20).
269
+
270
+ > **The repository has no `pyproject.toml` and no `setup.py`.** `app` is a plain package, so the repo root
271
+ > must be on `sys.path` for `from app.space_app import ...` to resolve
272
+ > (`deploy/codespace/launch.sh:31-34`; `pytest.ini` sets `pythonpath = .`).
273
+
274
+ ## 5. Where the contracts live
275
+
276
+ Three files are authoritative, and the code β€” not a document β€” wins when they disagree:
277
+
278
+ | Contract | File | What it fixes |
279
+ |---|---|---|
280
+ | the config registry | `configs/base.yaml` | every tunable value; hashed (Β§7) |
281
+ | the wire/data schemas | `core/schemas.py` | the request, result, evidence, trace and health shapes |
282
+ | the specialist interface | `specialists/base.py` | the four-method contract every specialist implements |
283
+ | the capability table | `core/registry.py` | which capabilities exist and how to build them |
284
+ | the planner's mapping | `core/planner.py` | `TASK_CAPABILITY` and `CAPABILITY_ASSETS` |
285
+
286
+ When `docs/API_CONTRACT.md` and `core/schemas.py` disagree, the contract document's own authority clause
287
+ resolves it: *"the request/response shapes are not invented here. They are the existing, tested Pydantic
288
+ models in `core/schemas.py`. This document describes them; it does not declare new ones."*
289
+ (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15). The known instance of this β€” the forward-compatibility
290
+ promise versus `extra="forbid"` β€” is recorded as **C-2** and left as a documented contradiction rather
291
+ than silently resolved (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15).
292
+
293
+ ## 6. The package layout inside `core/`, `specialists/`, `app/`
294
+
295
+ ### 6.1 `core/`
296
+
297
+ | File | Role |
298
+ |---|---|
299
+ | `core/config.py` | the loader, the validator, the hash (Β§7, Β§10) |
300
+ | `core/schemas.py` | every Pydantic model (Β§11) |
301
+ | `core/controller.py` | the nine-state controller that runs a plan and produces the trace |
302
+ | `core/planner.py` | the policy layer: `TASK_CAPABILITY`, `CAPABILITY_ASSETS`, `plan()` |
303
+ | `core/registry.py` | the capability registry: spec table, lazy construction, degradation states |
304
+ | `core/errors.py` | the typed error taxonomy and `scrub_paths` |
305
+ | `core/code_revision.py` | the code revision recorded in a run |
306
+
307
+ The three-tier control split is deliberate: the **registry** knows *which* specialists exist and *how* to
308
+ construct them; the **planner** *decides* what runs; the **controller** executes. The registry's own
309
+ docstring states it: *"It never decides what runs β€” that is `core.planner`'s job alone"*
310
+ (`core/registry.py:6-8`).
311
+
312
+ ### 6.2 `specialists/`
313
+
314
+ | Package | Files | Capability |
315
+ |---|---|---|
316
+ | `specialists/vqa/` | `inference.py`, `model.py`, `prompts.py` | `vqa`, `caption` |
317
+ | `specialists/grounding/` | `specialist.py`, `remoteclip.py`, `head.py`, `inference.py` | `grounding` |
318
+ | `specialists/change/` | `specialist.py`, `stanet.py`, `vqa_specialist.py`, `postprocess.py` | `change`, `change_vqa` |
319
+ | `specialists/optical_sar/` | `specialist.py`, `croma.py`, `fusion_head.py`, `sensor_adapter.py`, `radiometry.py`, `inference.py`, `prompts.py`, `vendor/` | `optical_sar` |
320
+
321
+ ### 6.3 `app/`
322
+
323
+ | File | Role |
324
+ |---|---|
325
+ | `app/serving.py` | the composition root β€” wires checkpoints through the registry's `builders=` override |
326
+ | `app/space_app.py` | `build_space_app()` β€” the FastAPI app and the four routes |
327
+ | `app/deployment.py` | the deployment description and device resolution |
328
+
329
+ ---
330
+
331
+ # Part III β€” The configuration discipline
332
+
333
+ ## 7. "No magic numbers in Python"
334
+
335
+ `configs/base.yaml` opens with the rule, in the file itself:
336
+
337
+ ```yaml
338
+ # RULE: no magic numbers anywhere in Python. Everything tunable lives here.
339
+ # Every value below is loaded, validated and hashed by core/config.py.
340
+ ```
341
+
342
+ (`configs/base.yaml:4-5`)
343
+
344
+ The loader reinforces it: *"One config system. No duplicated constants. Every value in configs/base.yaml
345
+ is loaded, validated against the frozen architecture, and hashed so evaluation runs are reproducible."*
346
+ (`core/config.py:1-5`).
347
+
348
+ **Access pattern.** Never read the YAML directly; import the singleton:
349
+
350
+ ```python
351
+ from core.config import get_config
352
+
353
+ cfg = get_config()
354
+ cfg.get("croma.image_resolution") # dotted-path access
355
+ cfg.require("change.encoder") # raises ConfigError if missing
356
+ cfg.seed # project.seed, default 42
357
+ ```
358
+
359
+ `get_config()` is an `lru_cache(maxsize=1)` singleton β€” *"Import this, do not re-read YAML"*
360
+ (`core/config.py:270-273`).
361
+
362
+ ## 8. What happens if you break the rule
363
+
364
+ Two distinct failures, and both are loud.
365
+
366
+ ### 8.1 The loader fails startup
367
+
368
+ `Config.__init__` calls `_validate()`, which collects **every** violation and raises a single
369
+ `ConfigError` naming them all (`core/config.py:46-49,94-222`):
370
+
371
+ ```python
372
+ if errors:
373
+ raise ConfigError(
374
+ "configuration failed frozen-architecture validation:\n - "
375
+ + "\n - ".join(errors)
376
+ )
377
+ ```
378
+
379
+ The header states the intent: *"if a config tries to violate a frozen decision (e.g. bf16 on T4,
380
+ torch.compile on ZeroGPU, a CROMA image_resolution that is not a multiple of 8), it fails loudly rather
381
+ than at runtime"* (`core/config.py:6-8`).
382
+
383
+ ### 8.2 Editing the config MOVES the hash
384
+
385
+ The hash is a sha256 over the whole registry, truncated to 16 hex characters
386
+ (`core/config.py:76-80`):
387
+
388
+ ```python
389
+ @property
390
+ def hash(self) -> str:
391
+ """Stable hash of the whole registry. Recorded in every evaluation run."""
392
+ blob = json.dumps(self._data, sort_keys=True, default=str).encode()
393
+ return hashlib.sha256(blob).hexdigest()[:16]
394
+ ```
395
+
396
+ Because it is computed over the entire registry, **any** edit to `configs/base.yaml` changes it. The
397
+ frozen value is **`78f1e3700da15aa1`** (`release/repo/README.md` Β§Reproducibility;
398
+ `docs/PHASE18_DEPLOYMENT_PACKAGING.md` Β§3). Every artifact records the hash it was produced against, so a
399
+ hash change **invalidates every artifact keyed to it**.
400
+
401
+ > **This is the one non-reversible action in the repository.** `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§6.1
402
+ > lists it as the single row in the rollback table whose answer to "Reversible?" is **no**: *"The recorded
403
+ > benchmark hash is gone; the frozen benchmark no longer matches."*
404
+
405
+ **Verify the hash after any config-adjacent change:**
406
+
407
+ ```bash
408
+ $PY -c "from core.config import get_config; print(get_config().hash)"
409
+ # expected: 78f1e3700da15aa1
410
+ ```
411
+
412
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§1.1, where `$PY` is `$REPO/.venv/Scripts/python.exe` on Windows.)
413
+
414
+ **Two things that do *not* move the hash** β€” both verified:
415
+
416
+ 1. **`configs/deploy.yaml` is inert.** It carries a `registry: false` marker and is never merged into the
417
+ registry, so editing it cannot move the hash. Verify with `$PY scripts/validate_deploy_config.py` β†’
418
+ exit 0, `"deploy.yaml is inert (not in the registry) and C-8-consistent."`
419
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§2.4). **But** `scripts/validate_deploy_config.py` hard-fails if
420
+ the `deployment:` block in `deploy.yaml` differs key-for-key from `base.yaml`'s, so a change to one
421
+ alone fails the validator (`docs/DEPLOYMENT_DECISION.md` Β§4).
422
+ 2. **The serving wiring uses `builders=`.** `app/serving.py` wires the change and change-VQA heads through
423
+ the registry's `builders=` override instead of config, which is *"the seam that keeps `Config.hash`
424
+ unchanged while still pointing serving at the trained artifacts"* (`docs/PHASE19_FINAL_HARDENING.md`
425
+ Β§3.4).
426
+
427
+ ## 9. The hash-exempt path: environment overrides
428
+
429
+ Two registry values can be overridden from the environment without editing the YAML
430
+ (`core/config.py:261-265`):
431
+
432
+ | Variable | Effect |
433
+ |---|---|
434
+ | `SATQUERY_PRECISION` | overrides `training.precision` |
435
+ | `SATQUERY_TORCH_COMPILE` | overrides `deployment.torch_compile` (`"true"` β†’ `True`) |
436
+
437
+ **Both are still validated.** Setting `SATQUERY_TORCH_COMPILE=true` **fails startup** because finding C-8
438
+ forbids `torch.compile` (`release/repo/docs/DEPLOYMENT.md` Β§6.3; `core/config.py:114-119`).
439
+
440
+ This is the general pattern for deployment state that must not move the hash: read it from the
441
+ environment. The asset-store capacity, TTL and per-file cap follow the same rule
442
+ (`release/repo/README.md` Β§Installation).
443
+
444
+ ## 10. The invariants the loader enforces
445
+
446
+ `core/config.py::_validate` is not documentation β€” it is a check that raises `ConfigError`. Each invariant
447
+ exists because a specific finding proved the failure mode.
448
+
449
+ | Invariant | Why it exists | Finding |
450
+ |---|---|---|
451
+ | `croma.image_resolution % 8 == 0` | CROMA asserts this; native 120 β†’ 225 patches | C-7 |
452
+ | `training.precision ∈ {fp16, bf16, fp32}` | the T4 is SM 7.5, so bf16 is unavailable | C-6 |
453
+ | `deployment.torch_compile is not true` | ZeroGPU does not support `torch.compile` | C-8 |
454
+ | `vlm.processor_longest_edge ≀ image.tile_size` | the processor's default `longest_edge` is 2048, which upscales a 512 px tile 4Γ— and then splits it into **17** sub-images β€” a ~17Γ— overrun, not the 4Γ— the plan estimated | F5-2 |
455
+ | `vlm.prompt_must_use_chat_template is true` | SmolVLM raises `ValueError` on prompts lacking one `<image>` token per image | F5-3 |
456
+ | `fusion.input_dim == 3*encoder_dim + optical_channels + sar_channels` (= 2318) | CROMA emits optical/SAR/joint GAP vectors; the availability mask is consumed by the head | C-1 |
457
+ | `croma.optical_channels == 12` and `croma.sar_channels == 2` | CROMA's `s2_channels` / `s1_channels` are fixed | β€” |
458
+ | `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (= 2048) | a mismatch is a **silent** shape error β€” torch raises only at the similarity step, after patch features are already cached | P7-1 |
459
+ | `router.tasks` includes `unsupported` and `router.num_tasks == len(router.tasks)` | the ontology and its declared size cannot drift apart | β€” |
460
+ | `change.sa_mode ∈ {BAM, PAM}` and `change.encoder` is set | the change architecture is not implicit | C-9 |
461
+ | `image.top_k_tiles ≀ image.max_tiles` | the dispatch ceiling cannot exceed the examination ceiling | β€” |
462
+
463
+ (`core/config.py:94-222`; `release/repo/README.md` Β§Installation.)
464
+
465
+ > **Why the grounding-head guard matters most.** A mismatch there is a *silent* shape error: torch raises
466
+ > only at the similarity step, by which point the patch features have already been computed and cached β€”
467
+ > *"the failure surfaces far from its cause"* (`core/config.py:170-196`).
468
+
469
+ ---
470
+
471
+ # Part IV β€” Contract-first development
472
+
473
+ ## 11. `core/schemas.py` is the binding contract
474
+
475
+ `core/schemas.py` (462 lines) holds every model the system exchanges. The rule is simple: **no specialist
476
+ may invent its own result shape.** Every specialist returns exactly `SpecialistResult`
477
+ (`core/schemas.py:325-406`), whose docstring says so: *"Every specialist returns exactly this. No
478
+ exceptions."*
479
+
480
+ The models, in order:
481
+
482
+ | Model | Line | Role |
483
+ |---|---|---|
484
+ | `Task` (enum) | 35 | the six-task ontology (`vqa`, `caption`, `grounding`, `change`, `optical_sar`, `change_vqa`) + `unsupported` |
485
+ | `Modality` (enum) | 50 | `optical` / `sar` / `joint` |
486
+ | `CoordinateSystem` (enum) | 57 | `normalized_0_1` / `geo` |
487
+ | `EvidenceType` (enum) | 65 | the evidence classes |
488
+ | `ControllerState` (enum) | 79 | the nine-state spine |
489
+ | `Intent` | 94 | the router's output |
490
+ | `GeoMetadata` | 120 | CRS / transform metadata |
491
+ | `SensorDescriptor` | 138 | sensor identity |
492
+ | `AssetMetadata` | 153 | one input asset |
493
+ | `Box` | 171 | a flat box (see the flat-vs-nested trap below) |
494
+ | `Region` / `ChangeRegion` | 189 / 202 | spatial outputs |
495
+ | `Evidence` | 216 | one observable artefact |
496
+ | `ConfidenceBreakdown` | 259 | measurable confidence |
497
+ | `TraceStep` / `ModelRef` / `ExecutionTrace` | 279 / 288 / 296 | the trace |
498
+ | `SpecialistResult` | 325 | the master contract |
499
+ | `AnalysisRequest` | 412 | the request |
500
+ | `ResultEnvelope` | 421 | the response wrapper |
501
+ | `HealthStatus` | 430 | the health shape |
502
+
503
+ **Every model sets `model_config = ConfigDict(extra="forbid")`.** This is deliberate and has a documented
504
+ consequence (C-2, Β§5): the models reject a body carrying an unknown key, which contradicts
505
+ `API_CONTRACT.md` Β§1.1's forward-compatibility promise. The code is authoritative; the document is the
506
+ inaccurate half (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§15).
507
+
508
+ > **The flat-vs-nested trap.** `Box` is **flat**, not nested. An earlier documentation draft described it
509
+ > with nested geometry, and *"every box the frontend drew would have been at the origin"*
510
+ > (`docs/PHASE19_FINAL_HARDENING.md` Β§3.6). Documentation is validated against the real models by tests
511
+ > (`test_api_contract_doc.py`, `test_frontend_guide_doc.py`), which is what caught it.
512
+
513
+ ### 11.1 The two cross-field validators you must not break
514
+
515
+ `SpecialistResult` carries a `_task_output_consistency` validator (`core/schemas.py:353-406`) with two
516
+ rules:
517
+
518
+ - **Grounding with no localisation is degraded, not a crash.** If `task == GROUNDING` and there are no
519
+ boxes or regions, the validator appends a warning and sets `degraded = True`.
520
+ - **Change-VQA with no answer text is degraded.** If `task == CHANGE_VQA` and the answer is blank, the
521
+ validator marks it degraded β€” but only if the specialist has not already set `degraded` itself.
522
+
523
+ The `CHANGE` clause that once lived there was **removed** (F-16c): it had collapsed to `not regions`, and
524
+ a *successful* no-change analysis began reporting `degraded: true`. **CHANGE is the one task whose
525
+ `degraded` flag is now set entirely by its specialist** (`core/schemas.py:362-395`).
526
+
527
+ ## 12. The specialist interface
528
+
529
+ Every specialist implements exactly one abstract base class, `Specialist`, in `specialists/base.py`. The
530
+ interface is **four methods**, and the split is deliberate (`specialists/base.py:1-17`):
531
+
532
+ ```
533
+ validate_request -> can this specialist serve this request at all?
534
+ execute -> do the work, return a SpecialistResult
535
+ produce_evidence -> what observable artefacts support the result?
536
+ estimate_confidence -> what measurable signals support the score?
537
+ ```
538
+
539
+ The class attributes a specialist must set (`specialists/base.py:48-58`):
540
+
541
+ | Attribute | Meaning |
542
+ |---|---|
543
+ | `name` | stable machine-readable name, used in traces and evidence sources |
544
+ | `version` | semantic version; *"bump when behaviour changes, not when code moves"* |
545
+ | `capabilities` | the capability strings this specialist serves, e.g. `("vqa", "caption")` |
546
+
547
+ The four abstract methods (`specialists/base.py:62-96`):
548
+
549
+ | Method | Contract |
550
+ |---|---|
551
+ | `validate_request(request)` | raise the **most specific** typed error available (`InvalidRequestError`, `PairMisalignmentError`, …), never a bare `Exception` |
552
+ | `execute(request)` | return a normalised `SpecialistResult`; *"Never returns None."* |
553
+ | `produce_evidence(result)` | return `list[Evidence]`; *"Must not invent anything the specialist did not actually compute."* |
554
+ | `estimate_confidence(result)` | return a `ConfidenceBreakdown`; *"Never an LLM utterance."* |
555
+
556
+ The base class also provides helpers a specialist should use rather than re-implement
557
+ (`specialists/base.py:98-134`): `supports()`, `require_assets()` (asserts an exact asset count and raises
558
+ a typed error), `require_capability()`, `model_refs()` (for the trace), and `describe()`.
559
+
560
+ `SpecialistRequest` (`specialists/base.py:34-45`) is the input: `assets`, `query`, `params`, `run_id`.
561
+ Its docstring states the isolation rule: *"Everything a specialist is given. No specialist reads global
562
+ state."*
563
+
564
+ ## 13. How to add a new specialist
565
+
566
+ This is the verified procedure, read from the code. There are **five** steps, and skipping any one fails
567
+ loudly (which is the design).
568
+
569
+ ### Step 1 β€” implement the `Specialist` ABC
570
+
571
+ Subclass `specialists.base.Specialist`, set `name`, `version` and `capabilities`, and implement the four
572
+ methods. The builder for the class is a module-level function
573
+ `build_<x>_specialist(config, *, device, **kwargs)` β€” the same shape as the four existing builders
574
+ (`core/registry.py:185-259` names them: `build_vqa_specialist`, `build_grounding_specialist`,
575
+ `build_change_specialist`, `build_change_vqa_specialist`, `build_optical_sar_specialist`).
576
+
577
+ ### Step 2 β€” add a row to the spec table
578
+
579
+ Add a `SpecialistSpec` to `default_specs()` in `core/registry.py:171-260`. The spec's fields
580
+ (`core/registry.py:121-165`):
581
+
582
+ | Field | Meaning |
583
+ |---|---|
584
+ | `name` | the registry key β€” **the capability string**, never the specialist's own `name` |
585
+ | `capabilities` | the tuple the built object must declare; **asserted** after construction |
586
+ | `module` | dotted module path containing the builder |
587
+ | `builder` | builder function name inside `module` |
588
+ | `requires_assets` | exact asset count, or `None` for "any" |
589
+ | `config_keys` | `{builder_kwarg: dotted.config.key}` β€” only keys that resolve are passed |
590
+ | `optional_config_keys` | as above, but a missing key contributes nothing |
591
+ | `failure_states` | registry state to use when construction raises, keyed by exception class name |
592
+
593
+ The `SpecialistSpec.__post_init__` refuses a spec whose `name` is not in its own `capabilities`
594
+ (`core/registry.py:158-165`): *"the registry keys on capability, so this spec would be unreachable."*
595
+
596
+ ### Step 3 β€” understand the capability-vs-name asymmetry
597
+
598
+ **This is the trap the registry exists to encode** (`core/registry.py:30-45`). The VQA specialist declares:
599
+
600
+ ```python
601
+ name = "vlm" # specialists/vqa/inference.py:114
602
+ capabilities = ("vqa", "caption") # specialists/vqa/inference.py:116
603
+ ```
604
+
605
+ Every other specialist's `name` equals its single capability. A registry keyed on `name` would make `vqa`
606
+ permanently unfindable while every other specialist kept working. So the registry keys on **capability**
607
+ and **asserts** the capability tuple against the constructed object
608
+ (`core/registry.py:497-515`). If your spec's `capabilities` disagrees with what your specialist declares,
609
+ construction fails with a `SpecialistError` naming both tuples.
610
+
611
+ ### Step 4 β€” register the task (only if it is a *new* task)
612
+
613
+ If the new specialist serves an existing task, nothing else is needed. If it is a new task, add it to:
614
+
615
+ - `Task` in `core/schemas.py:35-49`, and
616
+ - `TASK_CAPABILITY` in `core/planner.py:121-128`, and
617
+ - `CAPABILITY_ASSETS` in `core/planner.py:133-142`.
618
+
619
+ `CAPABILITY_ASSETS` mirrors each specialist's own `validate_request`, which **stays authoritative** β€” the
620
+ planner's copy is a *"cheap precondition so it can refuse before construction is attempted"*
621
+ (`core/planner.py:130-132`). A test asserts `CAPABILITY_ASSETS` equals `SpecialistSpec.requires_assets`
622
+ in both directions, and asserts the gateway carries **no third copy** of the asset-count table
623
+ (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§3).
624
+
625
+ > **Adding a task touches `router.tasks`.** `router.num_tasks` must equal `len(router.tasks)`
626
+ > (`core/config.py:199-206`), so a new task means a config edit β€” **which moves the hash** (Β§8.2). This is
627
+ > the one part of "add a specialist" that has a global consequence.
628
+
629
+ ### Step 5 β€” return the right shapes
630
+
631
+ `execute` must return a `SpecialistResult`; `produce_evidence` must return `list[Evidence]`;
632
+ `estimate_confidence` must return a `ConfidenceBreakdown`. The registry will construct your specialist
633
+ lazily and record its state.
634
+
635
+ ### What the registry does with your specialist
636
+
637
+ | Behaviour | Mechanism |
638
+ |---|---|
639
+ | lazy construction | the builder is imported via `importlib.import_module` **inside** `build()` β€” no module-level specialist import (`core/registry.py:15-25,462-475`) |
640
+ | memoisation | the second `build(cap)` returns the cached entry (`core/registry.py:432-434`) |
641
+ | three states | `AVAILABLE` / `DEGRADED` / `UNAVAILABLE` (`core/registry.py:102-107`) |
642
+ | degradation detection | duck-typed on `has_checkpoint` / `has_head` / `has_encoder` / `model is None` (`core/registry.py:527-548`) |
643
+ | failure is retained | a construction failure returns an `UNAVAILABLE` entry rather than raising; **only an unknown capability raises** (`core/registry.py:418-450`) |
644
+ | corrupt β‰  missing | `ModelLoadError` / `ModelUnavailableError` map to `UNAVAILABLE`, never retried as `DEGRADED` (`core/registry.py:550-612`) |
645
+ | path scrubbing | the client-visible `detail` is `scrub_paths(...)`; the raw string goes to the log only (`core/registry.py:560-598`) |
646
+
647
+ > **The corrupt-vs-missing rule is load-bearing.** *"silently running an untrained model because a real
648
+ > checkpoint failed to load would be the worst outcome"* (`specialists/change/specialist.py:834-838`,
649
+ > quoted in `core/registry.py:63-75`). Do not add a fallback that re-adds a degradation the builder
650
+ > refused.
651
+
652
+ ## 14. How evidence is emitted
653
+
654
+ Evidence is produced by `Specialist.produce_evidence(result)` and returned as `list[Evidence]`. The
655
+ `Evidence` model (`core/schemas.py:216-256`):
656
+
657
+ | Field | Type | Note |
658
+ |---|---|---|
659
+ | `evidence_id` | `str` | auto-generated `ev_<hex>`; **must be unique within a result** |
660
+ | `type` | `EvidenceType` | the evidence class |
661
+ | `source_specialist` | `str` | which specialist computed it |
662
+ | `coordinate_system` | `CoordinateSystem \| None` | required for spatial evidence |
663
+ | `coordinates` | `list[float] \| None` | the geometry |
664
+ | `score` | `float \| None` | `0.0 ≀ score ≀ 1.0` |
665
+ | `artifact_ref` | `str \| None` | **never a filesystem path** (F-16) |
666
+ | `payload` | `dict[str, Any]` | structured detail |
667
+
668
+ Two validators enforce correctness:
669
+
670
+ - `SpecialistResult._unique_evidence_ids` rejects a result with duplicate `evidence_id` values
671
+ (`core/schemas.py:345-351`).
672
+ - `Evidence._spatial_needs_crs` rejects spatial evidence (`BOUNDING_BOX`, `MASK`, `CHANGE_MAP`, `TILE`,
673
+ `IMAGE_CROP`, `JOINT_FEATURE_REGION`) that carries coordinates but **no** `coordinate_system`
674
+ (`core/schemas.py:241-255`).
675
+
676
+ > **`artifact_ref` is never a filesystem path.** v1 exposes no artifact-serving endpoint, so it is null
677
+ > unless a deployment supplies a client-fetchable reference. *"An artifact may still be written
678
+ > server-side where configured; being written is not the same as being retrievable."*
679
+ > (`core/schemas.py:227-238`.)
680
+
681
+ ---
682
+
683
+ # Part V β€” The test workflow
684
+
685
+ ## 15. Running tests
686
+
687
+ ### 15.1 Always use the repository venv interpreter
688
+
689
+ pytest is installed **only** in the repository virtualenv. Invoking the system `pytest` fails or resolves
690
+ to a different interpreter (`release/repo/docs/REPRODUCIBILITY.md` Β§10.2):
691
+
692
+ ```bash
693
+ .venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q
694
+ ```
695
+
696
+ On Windows, pytest must be given `-p no:cacheprovider` because the sandbox refuses `.pytest_cache` writes
697
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§0).
698
+
699
+ ### 15.2 Run targeted files, not the whole tree
700
+
701
+ **A full `tests/unit` run trips the sandbox's bulk-delete guard** (4Γ— `test_safe_delete_shim` failures)
702
+ (`release/repo/docs/REPRODUCIBILITY.md` Β§10.3). The workaround is to run the targeted suites:
703
+
704
+ ```bash
705
+ .venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q # 106 passed
706
+ ```
707
+
708
+ The runbook notes that **multi-suite invocations in one command have been refused by the environment
709
+ before; single suites are reliable** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` Β§2.6):
710
+
711
+ ```bash
712
+ $PY -m pytest tests/unit/test_deploy_config.py -p no:cacheprovider -q
713
+ $PY -m pytest tests/unit/test_app_serving.py -p no:cacheprovider -q
714
+ $PY -m pytest tests/unit/test_api_contract_doc.py tests/unit/test_frontend_guide_doc.py -p no:cacheprovider -q
715
+ ```
716
+
717
+ > **Do not pipe pytest through `grep`** in the authoring sandbox β€” output is block-buffered and a killed
718
+ > pipeline swallows it. Redirect to a file instead (`release/repo/docs/REPRODUCIBILITY.md` Β§5.6).
719
+
720
+ ### 15.3 `pytest.ini`
721
+
722
+ ```ini
723
+ [pytest]
724
+ testpaths = tests
725
+ pythonpath = .
726
+ addopts = -q --tb=short
727
+ ```
728
+
729
+ (`pytest.ini:1-4`)
730
+
731
+ `pythonpath = .` is what makes `from core.config import ...` resolve from the repo root without an
732
+ installed package.
733
+
734
+ ## 16. The evidence-class markers
735
+
736
+ `pytest.ini` registers four markers whose purpose is to make a test's **evidence class** explicit
737
+ (`pytest.ini:23-27`):
738
+
739
+ | Marker | Meaning |
740
+ |---|---|
741
+ | `unit` | fast, no I/O, no server, no network. Evidence about a component in isolation. |
742
+ | `integration` | exercises two or more real components wired together in-process. |
743
+ | `smoke` | a minimal end-to-end path run against real local artifacts, not a mock. |
744
+ | `real_inference` | ran the real model on real inputs in this environment. |
745
+
746
+ The file's comment states the discipline:
747
+
748
+ > *"a unit test never claims a deployment was exercised, and no test may be reported as 'real inference'
749
+ > unless it ran the real model on real data in this environment."* (`pytest.ini:13-17`)
750
+
751
+ **`environment_blocked` is deliberately not a marker.** *"a blocked path is reported in the STEP 7 report
752
+ rather than encoded as a permanently-skipped test, because a skip can be mistaken for coverage"*
753
+ (`pytest.ini:19-21`). An unregistered mark is an **error** rather than a silent no-op (`pytest.ini:14-15`).
754
+
755
+ ## 17. The known failures, and why they are not regressions
756
+
757
+ Running the entire unit tree trips **5–6** failures, classified by cause
758
+ (`release/repo/docs/REPRODUCIBILITY.md` Β§5.4):
759
+
760
+ | # | Failure | Cause | Regression? |
761
+ |---|---|---|---|
762
+ | 1–4 | `test_safe_delete_shim` (Γ—4) | the sandbox's bulk-**delete guard** (Windows verbatim-path behaviour) | **No** |
763
+ | 5 | one ordering flake in the router route test | passes in isolation; order/collection-dependent | **No** |
764
+ | 6 | one stale adapter test | asserts `optical_sar` absent when CROMA is *unshipped* β€” **CROMA is now shipped** | **No** |
765
+
766
+ Re-running the affected files together passes **137** tests, which is what isolates them as environmental
767
+ rather than behavioural (`release/repo/docs/REPRODUCIBILITY.md` Β§5.4).
768
+
769
+ **Known per-suite results:**
770
+
771
+ | Suite | Command | Expected |
772
+ |---|---|---|
773
+ | Frontend live-wiring | `pytest tests/unit/test_frontend_live_wiring.py` | **106 passed** |
774
+ | Doc/frontend suite | `pytest` on the 5 doc/frontend files | **183 passed** |
775
+ | Full unit suite | `pytest tests/unit` | 5–6 **environmental** failures |
776
+ | Gateway policy | `pytest tests/unit/test_gateway_policy.py` | **51 passed** |
777
+ | Evidence engine | `pytest tests/unit/test_evidence_engine.py` | **73 passed** |
778
+
779
+ (`release/repo/README.md` Β§The test suites; `docs/PHASE19_FINAL_HARDENING.md` Β§6.)
780
+
781
+ > **The precise full-suite *collected* count is `UNKNOWN β€” not established from the available evidence`.**
782
+ > Per-suite counts are known; the single collected total is not (`release/repo/docs/REPRODUCIBILITY.md`
783
+ > Β§5.4).
784
+
785
+ ---
786
+
787
+ # Part VI β€” Scripts and notebooks
788
+
789
+ ## 18. The class of scripts under `scripts/`
790
+
791
+ `scripts/` is a **flat** directory of executable helpers β€” no sub-packages. The listing holds **51 script
792
+ files** (47 `.py`, 3 `.ps1`, 1 `.mjs`) plus `__init__.py`. They fall into clear classes:
793
+
794
+ | Class | Examples | Purpose |
795
+ |---|---|---|
796
+ | **Training entry points** | `train_router.py`, `train_grounding.py`, `train_change.py`, `train_fusion.py`, `train_change_vqa.py` | thin CLIs over the `training/` modules (Β§20) |
797
+ | **Evaluation** | `eval_change.py`, `eval_fusion_115.py`, `eval_grounding_head.py`, `evaluate_change_vqa.py` | per-task evaluation |
798
+ | **Data preparation** | `prepare_bigearthnet.py`, `prepare_change_vqa.py`, `select_bigearthnet_slice.py` | build corpora and selection manifests |
799
+ | **Contract probes** | `probe_grounding_head_contract.py`, `probe_remoteclip_contract.py`, `probe_vlm_contract.py` | measure a real model's contract before relying on it |
800
+ | **Smoke tests** | `smoke_test_change.py`, `smoke_test_grounding_head.py`, `smoke_test_vlm.py` | a minimal real path per specialist |
801
+ | **Verification / gates** | `verify_gate1.py`, `verify_croma_forward.py`, `verify_grounding_e2e.py`, `verify_levir_real.py`, `verify_cdvqa_imagery.py` | prove a property end to end |
802
+ | **Threshold / hyperparameter sweeps** | `sweep_change_threshold.py`, `sweep_router_threshold.py`, `fusion_seed_variance.py` | sweep a frozen knob |
803
+ | **Phase-6 (VLM) tooling** | `phase6_train_vlm.py`, `phase6_adjudicate_test.py`, `phase6_close.py`, `phase6_rerule.py`, `phase6_recover_baseline_test.py` | the VLM acceptance workflow |
804
+ | **Phase-12 (fusion) tooling** | `p12_integrity_verify.py`, `p12_preflight_verify.py`, `run_phase12_extraction.ps1`, `run_phase12_resume_ref.ps1`, `run_phase12_resume_s1_ref.ps1` | the fusion feature-extraction workflow (PowerShell drivers) |
805
+ | **Calibration** | `fit_calibration.py` | fit the temperature-scaling artifact on validation (Β§22) |
806
+ | **Kaggle packaging / rehearsal** | `package_kaggle_code.py`, `rehearse_kaggle_notebook.py`, `rehearse_change_notebook.py` | package and dry-run the notebooks |
807
+ | **Deployment validation** | `validate_deploy_config.py` | prove `configs/deploy.yaml` is inert |
808
+ | **Frontend staging** | `stage_pages.mjs` | build the Cloudflare Pages bundle |
809
+ | **Environment / diagnostics** | `check_env.py`, `diagnose_feature_cache.py`, `croma_normalisation_arm_probe.py`, `analyze_grounding_resolution.py`, `analyze_reben_labels.py`, `check_cdvqa_second_overlap.py`, `establish_cdvqa_temporal_order.py`, `exp_grounding_resolution.py`, `extract_fusion_features.py` | environment and dataset diagnostics |
810
+
811
+ **The training entry points share one convention**, stated in their docstrings: *"A THIN CLI over
812
+ `training/<task>/train.py` … Everything that computes lives in the module; this file parses arguments,
813
+ reports the environment, prints the accounting, and returns an exit code."* (`scripts/train_change.py:1-6`,
814
+ `scripts/train_fusion.py:1-6`). The exit codes are a contract (`scripts/train_change.py:17-24`):
815
+
816
+ ```
817
+ 0 training completed (or --dry-run validated a present dataset)
818
+ 2 the dataset is missing, empty, or cannot be split -- NOT a crash
819
+ 3 the run started and the trainer raised a typed error
820
+ ```
821
+
822
+ > **Re-exported names are part of the contract.** Several scripts re-export their trainer's symbols at
823
+ > module scope, and a test asserts it. E.g. `tests/unit/test_change_train_script_contract.py` asserts
824
+ > `change_loss`, `train_change_head` and `evaluate` are reachable as `scripts.train_change.*`
825
+ > (`scripts/train_change.py:26-30`).
826
+
827
+ ## 19. `notebooks/`
828
+
829
+ | Notebook | Purpose |
830
+ |---|---|
831
+ | `kaggle_change_vqa_train.ipynb` | R-02 change-VQA training (Β§21) |
832
+ | `kaggle_phase6_vlm_lora.ipynb` | Phase-6 SmolVLM LoRA training (Β§21) |
833
+ | `kaggle_change_training.ipynb` | the change **detector** β€” *"frozen and out of scope"* (`docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4) |
834
+ | `kaggle_grounding_resolution.ipynb` | Phase-7 grounding resolution β€” unrelated to R-02 |
835
+
836
+ > **Do not confuse the two change notebooks.** For change-VQA use `kaggle_change_vqa_train.ipynb`; do
837
+ > **not** use `kaggle_change_training.ipynb` (that one trains the change *detector*) or
838
+ > `kaggle_grounding_resolution.ipynb` (Phase 7) (`docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4).
839
+
840
+ Two `.bak-*` files sit beside the Phase-6 notebook; they are editor backups, not runnable notebooks.
841
+
842
+ ---
843
+
844
+ # Part VII β€” Training entry points
845
+
846
+ The six artifacts train in two places: four locally, two on an external GPU. This is a **deliberate
847
+ boundary** β€” the release ships *"frozen artifacts with provenance, not a retraining harness"*
848
+ (`release/repo/docs/REPRODUCIBILITY.md` Β§8.5).
849
+
850
+ | Artifact | Where it trains | Reproducible from this release? |
851
+ |---|---|---|
852
+ | router adapter | local CPU | **yes** β€” `configs/base.yaml` Β§`router.training` |
853
+ | grounding head | local | **yes** β€” `configs/base.yaml` Β§`grounding_training` |
854
+ | change head | local | **yes** β€” `configs/base.yaml` Β§`change` |
855
+ | optical_sar fusion head | local, seed sweep | **yes** |
856
+ | change_vqa head | **external GPU (Kaggle)** | **partly** β€” the promotion gate, evaluation and serving wiring are reproducible; there is no one-command retrain |
857
+ | vlm LoRA adapter | **external GPU** | **partly** β€” same |
858
+
859
+ (`release/repo/README.md` Β§What "reproduce" means.)
860
+
861
+ ## 20. Local training (router, grounding, change, optical-SAR)
862
+
863
+ ### 20.1 Router β€” CPU-only, and fast
864
+
865
+ ```bash
866
+ python scripts/train_router.py --smoke # fast, stub encoder
867
+ python scripts/train_router.py --stub # fast, no model download
868
+ python scripts/train_router.py # full run, real MiniLM
869
+ ```
870
+
871
+ (`scripts/train_router.py:1-10`)
872
+
873
+ The script's own measured note: *"Runs entirely on CPU. Measured cost with the real encoder: ~40 s to load
874
+ MiniLM on a cold cache, ~0.1 s to embed the corpus, ~1 s to train 60 epochs. There is no reason to spend
875
+ Kaggle GPU quota on this."* The encoder is frozen, so embeddings are cached and the adapter trains on
876
+ cached vectors (`configs/base.yaml:67-69`).
877
+
878
+ ### 20.2 Grounding β€” two separable stages
879
+
880
+ ```bash
881
+ # stage 1 only β€” measure the corpus before committing to training
882
+ python scripts/train_grounding.py --data-root <root> --extract-only
883
+
884
+ # both stages, real run
885
+ python scripts/train_grounding.py --data-root <root> --checkpoint <ckpt.pt>
886
+
887
+ # 2-batch forward+backward+checkpoint+reload, on CPU
888
+ python scripts/train_grounding.py --data-root <root> --debug
889
+ ```
890
+
891
+ (`scripts/train_grounding.py:6-22`)
892
+
893
+ Splitting extraction from training means a hyperparameter change does not re-encode 15,699 images, and a
894
+ crashed training run does not lose the cache. The stated bar: *"the Phase 7 zero-shot baseline scored
895
+ 0.0972 mean best IoU over 16,159 real eval records. This head must beat it by MIN_IMPROVEMENT_IOU to
896
+ justify existing. The comparison is printed whether or not it passes."*
897
+
898
+ ### 20.3 Change β€” a thin CLI over `training/change/train.py`
899
+
900
+ ```bash
901
+ # validate the data and the split; build nothing, train nothing
902
+ python scripts/train_change.py --data-root <LEVIR-CD root> --dry-run
903
+
904
+ # real run on CPU
905
+ python scripts/train_change.py --data-root <LEVIR-CD root> --device cpu
906
+
907
+ # a first real-data run that does not commit an hour
908
+ python scripts/train_change.py --data-root <root> --limit 256 --epochs 3
909
+ ```
910
+
911
+ (`scripts/train_change.py:8-16`)
912
+
913
+ `--dry-run` is *"the leakage check without the cost: it loads the items, performs the scene-disjoint
914
+ split, runs `assert_image_disjoint`, prints the accounting and exits. It does NOT build the detector, so
915
+ it cannot trigger a pretrained-weights download, and it writes no files."*
916
+
917
+ ### 20.4 Optical-SAR fusion β€” CPU-only, no result claimed
918
+
919
+ ```bash
920
+ # validate the caches and the split; train nothing, write nothing
921
+ python scripts/train_fusion.py --train-cache train.npz --val-cache val.npz --dry-run
922
+
923
+ # a real run on CPU (the fusion head is CPU-only)
924
+ python scripts/train_fusion.py --train-cache train.npz --val-cache val.npz --arm A --epochs 20
925
+ ```
926
+
927
+ (`scripts/train_fusion.py:8-15`)
928
+
929
+ > **No performance number is a result here.** The script states it: *"This loop has never seen the real
930
+ > paired BigEarthNet-S1+S2 corpus. Every run record it writes carries `result_status` and
931
+ > `pre_registered_metric_computed: false`; the pre-registered 11.5 metric is not computed."*
932
+ > (`scripts/train_fusion.py:17-20`.)
933
+
934
+ ## 21. External GPU training (change-VQA, VLM LoRA)
935
+
936
+ Both external runs happen on **Kaggle** with **GPU T4 Γ—2**. Neither is a one-command retrain from this
937
+ release.
938
+
939
+ ### 21.1 Change-VQA (R-02)
940
+
941
+ **Guide:** `docs/R02_KAGGLE_TRAINING_GUIDE.md` (37 KB). **Runbook:** `RUNBOOK_CHANGE_VQA_KAGGLE.md`.
942
+
943
+ | Property | Value | Source |
944
+ |---|---|---|
945
+ | notebook | `notebooks/kaggle_change_vqa_train.ipynb` | `docs/R02_KAGGLE_TRAINING_GUIDE.md` Β§4 |
946
+ | accelerator | **GPU T4 Γ—2** | Β§5 |
947
+ | internet | **On** (MiniLM downloads) | Β§5 |
948
+ | cells | **all 34, top to bottom** | Β§7 |
949
+ | hard stop | `HARD_STOP_SECONDS = 3 * 3600` (the plan's 3 h) | Β§7 |
950
+ | trainer defaults | `epochs=40, batch_size=128, seed=42, patience=6, time_limit=10800s` | Β§7 |
951
+
952
+ The path-discovery cells find the code root (four markers) and the CDVQA root (12 annotations + 4 image
953
+ dirs), and section 3c verifies the frozen STANet checkpoint by size and SHA256
954
+ (`RUNBOOK_CHANGE_VQA_KAGGLE.md` Β§5). **Test splits are not readable by the trainer:**
955
+ `training/change_vqa/train.py` loads only Train and Val, and *"there is no option in either file that
956
+ changes that"* (`scripts/train_change_vqa.py:11-15`).
957
+
958
+ ### 21.2 VLM LoRA (Phase 6)
959
+
960
+ **Runbook:** `RUNBOOK_PHASE6_VLM_KAGGLE.md`.
961
+
962
+ | Property | Value | Source |
963
+ |---|---|---|
964
+ | notebook | `notebooks/kaggle_phase6_vlm_lora.ipynb` | `RUNBOOK_PHASE6_VLM_KAGGLE.md` Β§1 |
965
+ | accelerator | **GPU T4 Γ—2** | Β§4 |
966
+ | internet | **off** (base model attached as input) | Β§4 |
967
+ | precision | **`fp16`, not the plan's `bf16`** β€” T4 is SM 7.5 | Β§4 |
968
+ | base model | `HuggingFaceTB/SmolVLM-500M-Instruct`, ~1.02 GB `model.safetensors` | Β§4 |
969
+
970
+ The trainer raises a `ConfigError` on `bf16` rather than silently falling back, and the deviation is
971
+ recorded in the manifest under `plan_deviations` (`RUNBOOK_PHASE6_VLM_KAGGLE.md` Β§4).
972
+
973
+ > **The VLM adapter's status is `ACCEPTANCE-REJECTED`.** Its metrics are *usable* (`exact_match 0.963`)
974
+ > but it was not promoted. **USABLE β‰  ACCEPTED** (`DOCS_STYLE_GUIDE.md` Β§3). Do not describe the VLM path
975
+ > as accepted.
976
+
977
+ ## 22. Calibration
978
+
979
+ ```bash
980
+ python scripts/fit_calibration.py --dry-run # check inputs, exit
981
+ python scripts/fit_calibration.py # fit, write artifact
982
+ ```
983
+
984
+ (`scripts/fit_calibration.py:19-21`)
985
+
986
+ It fits on **validation** data, and both the fitter entry point and the artifact writer **refuse a
987
+ held-out split** β€” *"Fitting on Test or Test2 would make the reported confidence a function of the answers
988
+ it is used to score β€” a leak, not a calibration."* (`scripts/fit_calibration.py:7-13`.)
989
+
990
+ > **Calibration made ECE worse** β€” `0.013755 β†’ 0.014929` β€” and is **retained only because it is in the
991
+ > frozen config**. Never present it as an improvement (`DOCS_STYLE_GUIDE.md` Β§3).
992
+
993
+ ---
994
+
995
+ # Part VIII β€” Coding conventions
996
+
997
+ ## 23. What the code actually does
998
+
999
+ These conventions are visible across the files read. They are not aspirational style rules; each is
1000
+ observable in the code.
1001
+
1002
+ ### 23.1 Every module has a substantial docstring stating *why*
1003
+
1004
+ The files read are densely commented at the module and function level, and the comments explain decisions
1005
+ and failure modes rather than restating the code. Examples: `core/registry.py:1-76` (a 76-line module
1006
+ docstring on the capability-key trap), `deploy/codespace/launch.sh:1-24` (why the launcher is defensive),
1007
+ `gateway/app.py:53-106` (the annotation-scope trap). **A change that removes the reasoning from a
1008
+ comment removes the reason a future reader will not re-introduce the bug.**
1009
+
1010
+ ### 23.2 `from __future__ import annotations` is standard
1011
+
1012
+ It appears at the top of nearly every module (`core/config.py:11`, `core/registry.py:78`,
1013
+ `specialists/base.py:19`, `gateway/app.py:48`, `deploy/render/main.py:40`, `deploy/render/codespaces.py:16`,
1014
+ `deploy/codespace/warm_cache.py:28`, `scripts/*.py`). It is convenient β€” and it is the direct cause of
1015
+ Trap 6 (Β§30).
1016
+
1017
+ ### 23.3 Findings are recorded as short codes, inline
1018
+
1019
+ The code refers to findings by code (`C-1`, `C-6`, `C-8`, `F5-2`, `P7-1`, `F-15`, `F-16`, `F-17`) and
1020
+ states the failure each guard prevents. This is a documentation convention enforced by comments, e.g.
1021
+ `core/config.py:121-142` (finding F5-2, the 17Γ— overrun), `core/registry.py:227-235` (F-17, a config
1022
+ surface removed rather than left as a silent no-op).
1023
+
1024
+ ### 23.4 Pydantic models set `extra="forbid"` and carry validators
1025
+
1026
+ Every schema model forbids extra fields and uses `@field_validator` / `@model_validator` for cross-field
1027
+ rules (Β§11). New models should follow the same pattern.
1028
+
1029
+ ### 23.5 Loggers are module-level and named `satquery.<area>`
1030
+
1031
+ `logging.getLogger("satquery.orchestrator")` (`deploy/render/main.py:74`),
1032
+ `logging.getLogger("satquery.registry")` (`core/registry.py:95`), `logging.getLogger(__name__)`
1033
+ (`gateway/app.py:123`). Log calls use `%s`-style free text β€” there are no structured/JSON logs
1034
+ (`docs/architecture/10-observability-and-ops.md` Β§4.4).
1035
+
1036
+ ### 23.6 Client-visible strings are scrubbed
1037
+
1038
+ Server-side diagnostics that name absolute paths must not reach a client. `core/errors.scrub_paths` reduces
1039
+ a path to a basename before it is published (`core/registry.py:297-312,560-598`;
1040
+ `core/errors.py:39`). Any new client-visible field that could carry a path or an exception string must go
1041
+ through the same scrub.
1042
+
1043
+ ### 23.7 `noqa` comments state the reason
1044
+
1045
+ Where a lint suppression is used, the reason is written, e.g. `# noqa: E402 (see comment above)` in
1046
+ `gateway/app.py:85-106`, and `# noqa: BLE001 - report, never crash the warm step` in
1047
+ `deploy/codespace/warm_cache.py:74`.
1048
+
1049
+ ### 23.8 Tests assert the *cause*, not the symptom
1050
+
1051
+ The regression test for the annotation-scope trap asserts that the route has no query parameters and that
1052
+ `gateway.app.Request is starlette.requests.Request` β€” *"rather than the symptom, because asserting the
1053
+ symptom would be brittle"* (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13). Follow this when writing a
1054
+ regression test.
1055
+
1056
+ ### 23.9 The 88-column soft limit
1057
+
1058
+ The files read wrap around 88 characters. There is no committed linter config in the files read, so this
1059
+ is a convention, not an enforced rule β€” **the exact formatter/linter configuration is `UNKNOWN β€” not
1060
+ established from the available evidence`.**
1061
+
1062
+ ---
1063
+
1064
+ # Part IX β€” Known development traps
1065
+
1066
+ ## 24. The traps, in one table
1067
+
1068
+ | # | Trap | One-line consequence |
1069
+ |---|---|---|
1070
+ | 1 | the monorepo `deploy/` is stale and untracked | it is **not** the deployed source |
1071
+ | 2 | the sandbox proxy is dead | outbound calls need `--noproxy '*'` / `ProxyHandler({})` |
1072
+ | 3 | pytest exists only in the venv | the system `pytest` resolves to the wrong interpreter |
1073
+ | 4 | Chrome drops synthetic CDP key events without OS focus | a browser-driven run silently answers the default query |
1074
+ | 5 | Cloudflare `_headers` rules concatenate | a later rule cannot "fix" an earlier one |
1075
+ | 6 | `from __future__ import annotations` + FastAPI | an unresolvable `Request` annotation becomes a **required query parameter** |
1076
+ | 7 | the full-suite run trips the bulk-delete guard | 4 spurious `test_safe_delete_shim` failures |
1077
+ | 8 | a stale serve process keeps answering | it reports the **previous revision's** capabilities |
1078
+
1079
+ ## 25. Trap 1 β€” the stale untracked `deploy/`
1080
+
1081
+ **Symptom.** You edit `deploy/render/main.py`, deploy your change, and nothing changes in production β€” or
1082
+ you read `deploy/render/main.py` and cannot find the tunnel code the live service runs.
1083
+
1084
+ **Root cause.** `deploy/` inside the monorepo is **stale and untracked**. `git status` reports
1085
+ `?? deploy/` (verified in the working copy). It is **not** the deployed source
1086
+ (`release/repo/docs/DEPLOYMENT.md` Β§1; `release/CURRENT_RELEASE_STATE.md` Β§6).
1087
+
1088
+ **Evidence of divergence.** The monorepo `deploy/render/main.py` (532 lines) exposes `/api/health` with a
1089
+ `config` block that has **no** `tunnel` field and no `transport_mode` / `tunnel_timeout_s` /
1090
+ `wake_timeout_s` keys (`deploy/render/main.py:444-466`), whereas the **live** payload carries all of them
1091
+ (`release/repo/docs/DEPLOYMENT.md` Β§5). The monorepo copy also lacks `tunnel_agent.py` and `doctor.sh`,
1092
+ both of which `launch.sh` references (`deploy/codespace/launch.sh:83,89,153,160-168,179`).
1093
+
1094
+ **Fix.** Fetch the deployed file from the private repository and diff before editing. Treat the monorepo
1095
+ `deploy/` as documentation of intent, not as source.
1096
+
1097
+ ## 26. Trap 2 β€” the dead sandbox proxy needs `--noproxy '*'`
1098
+
1099
+ **Symptom.** Outbound HTTP calls fail or hang in the authoring sandbox.
1100
+
1101
+ **Root cause.** The sandbox proxy is dead; requests are routed to it and never reach the target.
1102
+
1103
+ **Fix.** Disable proxies for the call (`release/repo/docs/REPRODUCIBILITY.md` Β§10.1):
1104
+
1105
+ ```bash
1106
+ curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health
1107
+ ```
1108
+
1109
+ ```python
1110
+ opener = urllib.request.build_opener(urllib.request.ProxyHandler({}))
1111
+ ```
1112
+
1113
+ `release/tools/hf_verify.py:37-38` does exactly this (`opener_no_proxy()`). In a normal environment the
1114
+ flag is harmless; in the sandbox it is mandatory.
1115
+
1116
+ ## 27. Trap 3 β€” pytest only in the venv
1117
+
1118
+ **Symptom.** Invoking the system `pytest` fails or resolves to a different interpreter.
1119
+
1120
+ **Root cause.** pytest is installed only in `.venv`.
1121
+
1122
+ **Fix.** Always invoke the venv interpreter explicitly (`release/repo/docs/REPRODUCIBILITY.md` Β§10.2):
1123
+
1124
+ ```bash
1125
+ .venv/Scripts/python.exe -m pytest tests/unit/test_frontend_live_wiring.py -q
1126
+ ```
1127
+
1128
+ ## 28. Trap 4 β€” Chrome drops synthetic CDP key events
1129
+
1130
+ **Symptom.** A browser-driven run silently answers the *default* query; the query box looks untouched;
1131
+ `mock_nodes` is non-zero.
1132
+
1133
+ **Root cause.** Chrome **drops synthesized key events when the browser window does not hold OS focus**.
1134
+ `press_key` / `fill_input` (real CDP key events) are focus-gated; `Input.insertText` (`type_text`) is not.
1135
+
1136
+ **Measured.** With Chrome backgrounded, `press_key("Z")` left `#qtext.value` unchanged, while
1137
+ `type_text("Q")` inserted fine (`release/repo/docs/REPRODUCIBILITY.md` Β§6.4).
1138
+
1139
+ **Fix.** Use `type_text` (not `fill_input`), and **assert the input state before dispatch** β€” `q_ok`
1140
+ (the query box really held the query), `obs_ok` (`#obsTail == 'ready'`), `t0_ok` (both frames present
1141
+ where required). This is the single most dangerous trap because it produces a **silent false pass** β€” the
1142
+ pipeline "works", it just answered a different question
1143
+ (`release/repo/docs/REPRODUCIBILITY.md` Β§10.5).
1144
+
1145
+ ## 29. Trap 5 β€” Cloudflare `_headers` concatenate
1146
+
1147
+ **Symptom.** A specific cache-control rule does not take effect; the browser caches a file you expected it
1148
+ to revalidate.
1149
+
1150
+ **Root cause.** Cloudflare `_headers` rules **concatenate, they do not override.** Two matching rules are
1151
+ merged: a specific rule nested under a broad `/assets/img/*` rule yields
1152
+ `max-age=604800, …, max-age=0, must-revalidate` β€” and **Chromium takes the FIRST `max-age`**. The file's
1153
+ own "later rules override" comment is **false** (`release/repo/docs/DEPLOYMENT.md` Β§10).
1154
+
1155
+ **Fix.** Order rules so the *broadest* rule appears last, and never rely on a later rule overriding an
1156
+ earlier one. Related: Cloudflare **308-redirects `X.html` β†’ `/X`**, so reference the extensionless path
1157
+ (`release/repo/docs/REPRODUCIBILITY.md` Β§10.4).
1158
+
1159
+ ## 30. Trap 6 β€” the annotation-scope trap
1160
+
1161
+ **Symptom.** Every `POST` route returns `422` with FastAPI's own shape, without ever entering the handler:
1162
+
1163
+ ```
1164
+ POST /v1/analyze β†’ 422
1165
+ {"detail":[{"type":"missing","loc":["query","request"],"msg":"Field required"}]}
1166
+ ```
1167
+
1168
+ (`docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13.)
1169
+
1170
+ **Root cause.** The module uses `from __future__ import annotations`, so every annotation is a **string**
1171
+ at runtime. FastAPI resolves those strings via `get_typed_signature`, which calls
1172
+ `eval(annotation, func.__globals__)`. If `Request` is imported **inside** `create_app`, the route closures
1173
+ capture the name as a *local* of `create_app`; it never appears in `gateway.app.__globals__`, so
1174
+ resolution fails and FastAPI is left holding a bare `ForwardRef('Request')`. **FastAPI does not raise** β€”
1175
+ it silently falls back to treating the parameter as a **required query parameter named `request`**
1176
+ (`gateway/app.py:53-73`).
1177
+
1178
+ Three things were wrong at once: the request body was never read, the gateway's own validation never ran,
1179
+ and the error shape was FastAPI's `{"detail": ...}` rather than the contract's `{"error": {...}}`. **Every**
1180
+ POST route was affected, including `/v1/assets`.
1181
+
1182
+ **The asymmetry to internalise.** An unresolvable **parameter** annotation is *silently reinterpreted*
1183
+ (the handler never runs), while an unresolvable **return** annotation *raises*
1184
+ (`pydantic.errors.PydanticUndefinedAnnotation: name 'JSONResponse' is not defined`, which made
1185
+ `create_app()` unbuildable). Same root cause, opposite diagnosability
1186
+ (`gateway/app.py:88-106`).
1187
+
1188
+ **Fix.** Bind the annotation subjects at **module scope** (`gateway/app.py:79-106`):
1189
+
1190
+ ```python
1191
+ if TYPE_CHECKING: # pragma: no cover
1192
+ from starlette.requests import Request
1193
+ from starlette.responses import Response
1194
+
1195
+ #: Runtime bindings used as annotation subjects in this module. Deliberately
1196
+ #: module-level so `eval()` can find them. See the comment above.
1197
+ from starlette.requests import Request # noqa: E402 (see comment above)
1198
+ from starlette.responses import Response # noqa: E402 (see comment above)
1199
+ from fastapi.responses import JSONResponse # noqa: E402 (see comment above)
1200
+ ```
1201
+
1202
+ The regression test asserts the **cause**: the route has no query parameters, and
1203
+ `gateway.app.Request is starlette.requests.Request`.
1204
+
1205
+ > **The general rule.** Any name used as a FastAPI route annotation in a module with
1206
+ > `from __future__ import annotations` **must** be importable from that module's globals. Import it at
1207
+ > module scope, not inside the factory.
1208
+
1209
+ ## 31. Trap 7 β€” the full-suite bulk-delete guard
1210
+
1211
+ **Symptom.** Running the whole `tests/unit` tree trips 4Γ— `test_safe_delete_shim` failures.
1212
+
1213
+ **Root cause.** Windows **verbatim-path** defects in the sandbox's bulk-delete guard. The precise
1214
+ condition under which the shim intermittently triggers on Windows is `UNKNOWN β€” not established from the
1215
+ available evidence` (`release/repo/docs/REPRODUCIBILITY.md` Β§10.3).
1216
+
1217
+ **Fix / workaround.** Run the targeted suites (106 and 183 pass cleanly); treat the 4 shim failures as
1218
+ environmental, not regressions (Β§17).
1219
+
1220
+ ## 32. Trap 8 β€” the stale serve process
1221
+
1222
+ **Symptom.** The Codespace answers `/v1/health` and `/v1/capabilities`, but reports the **previous
1223
+ revision's** capabilities.
1224
+
1225
+ **Root cause.** A serve process started before a code or environment change keeps serving from old code.
1226
+ The serve process reads its environment exactly once, at startup (`deploy/codespace/launch.sh:100-102`).
1227
+
1228
+ **Fix.** `launch.sh` already handles it: it records a **stamp** of the revision and the asset
1229
+ configuration and restarts the server when the stamp disagrees
1230
+ (`deploy/codespace/launch.sh:100-147`). The rule for a developer: **a restart, not a reload, is required
1231
+ after any env or revision change.**
1232
+
1233
+ > *"A stale serve process is worse than no process: it answers /v1/health and /v1/capabilities from OLD
1234
+ > code, so the deployment looks alive while reporting the previous revision's capabilities."*
1235
+ > (`deploy/codespace/launch.sh:111-113`)
1236
+
1237
+ ### Related platform traps worth knowing
1238
+
1239
+ | Trap | Detail |
1240
+ |---|---|
1241
+ | a forwarded Codespace port returns `302` for a private repo | this is *why* the outbound tunnel exists (`release/repo/docs/DEPLOYMENT.md` Β§10) |
1242
+ | the tunnel agent must be started by the devcontainer `postStartCommand` | a restarted Codespace comes up with `agent_connected: false` otherwise |
1243
+ | never retry `POST /api/infer` at the gateway | a retry consumes inference twice |
1244
+ | `containerEnv` applies only at container **creation** | an env change needs a restart, which is why `launch.sh` re-exports on every start (`deploy/codespace/launch.sh:49-52`) |
1245
+ | the Codespace filesystem is **ephemeral** | uploaded assets and logs vanish with the Codespace (`deploy/codespace/launch.sh:44-47`) |
1246
+
1247
+ ---
1248
+
1249
+ # Part X β€” Status and evidence
1250
+
1251
+ ## 33. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for development
1252
+
1253
+ | # | Item | Status |
1254
+ |---|---|---|
1255
+ | 1 | B-07 β€” tunnel gaps; patch prepared, **not deployed** | **`OPEN`** |
1256
+ | 2 | B-02 β€” `/api/health` `codespace_name` trailing `\n` | **`OPEN` (cosmetic)** |
1257
+ | 3 | A `LICENSE` file | **`OPEN`** β€” **no LICENSE file exists**; README says to add one before public release |
1258
+ | 4 | The monorepo README | **materially stale** β€” it calls the frontend "hermetic", describes a 4-endpoint `/v1/*` contract, omits the tunnel, and points at the stale `deploy/` (`release/CURRENT_RELEASE_STATE.md` Β§6) |
1259
+ | 5 | `hf/SETUP.md` and `hf/README.md` | **stale** β€” they assert the project ships no weights, which is now false (`release/CURRENT_RELEASE_STATE.md` Β§6) |
1260
+ | 6 | The exact formatter / linter configuration | **`UNKNOWN`** β€” no committed config in the files read (Β§23.9) |
1261
+ | 7 | The precise full-suite collected test count | **`UNKNOWN`** β€” per-suite counts are known, the total is not |
1262
+ | 8 | The exact Windows trigger for the delete-shim flake | **`UNKNOWN`** |
1263
+ | 9 | Whether `doctor.sh` / `tunnel_agent.py` exist in the deployed inference repo | **`UNKNOWN`** β€” the monorepo copy lacks them |
1264
+ | 10 | A system-level end-to-end benchmark | **`NOT RUN`** β€” none exists |
1265
+ | 11 | The router **test**-split number | **`NOT RUN`** β€” only validation (`n = 86`, ungated) exists |
1266
+ | 12 | Captioning benchmark | **`NOT RUN`** β€” implemented, not benchmarked |
1267
+ | 13 | A one-command retrain for the two external artifacts | **not implemented** β€” deliberate boundary (Β§21) |
1268
+ | 14 | `torch.compile`, CUDA, quantisation, thread capping | **not implemented** β€” CPU-first by design (Β§1) |
1269
+ | 15 | Enforcement of a one-model cache (`cache_max_models: 1`) | **UNVERIFIED** β€” the value is reported, not proven enforced (`docs/DEPLOYMENT_DECISION.md` Β§5) |
1270
+
1271
+ > **The stale-README trap is worth its own line.** `README.md` in the monorepo is *"materially stale"*: it
1272
+ > calls the frontend *"hermetic οΏ½οΏ½οΏ½ no backend calls"* (it calls `/api/*` on Render), puts Render/Codespace
1273
+ > as *"in progress"* (both deployed), describes a 4-endpoint `/v1/*` contract (the live contract is
1274
+ > `/api/*`), omits the tunnel, and points at the stale untracked `deploy/` as the deployment source
1275
+ > (`release/CURRENT_RELEASE_STATE.md` Β§6). The public release documentation is authoritative; the
1276
+ > monorepo README is not.
1277
+
1278
+ ## 34. Where the evidence lives
1279
+
1280
+ | What | Where |
1281
+ |---|---|
1282
+ | the frozen config, invariants and hash | `configs/base.yaml`; `core/config.py:76-80,94-222` |
1283
+ | the binding schemas | `core/schemas.py` |
1284
+ | the specialist interface | `specialists/base.py` |
1285
+ | the capability table and lazy construction | `core/registry.py` |
1286
+ | the planner mapping | `core/planner.py:121-142` |
1287
+ | the dependency profiles and contract notes | `requirements.txt` |
1288
+ | the test markers | `pytest.ini` |
1289
+ | the launcher and warm-up | `deploy/codespace/launch.sh`, `deploy/codespace/warm_cache.py` |
1290
+ | the annotation-scope trap | `gateway/app.py:53-106`; `docs/STEP7_BACKEND_CHAIN_REPORT.md` Β§13 |
1291
+ | the deploy manifest is inert | `docs/PHASE18_DEPLOYMENT_PACKAGING.md`; `scripts/validate_deploy_config.py` |
1292
+ | the change-VQA Kaggle guide | `docs/R02_KAGGLE_TRAINING_GUIDE.md`; `RUNBOOK_CHANGE_VQA_KAGGLE.md` |
1293
+ | the VLM LoRA Kaggle runbook | `RUNBOOK_PHASE6_VLM_KAGGLE.md` |
1294
+ | the test suites and their expected results | `release/repo/docs/REPRODUCIBILITY.md` Β§5 |
1295
+ | the environment traps | `release/repo/docs/REPRODUCIBILITY.md` Β§10 |
1296
+ | the platform traps | `release/repo/docs/DEPLOYMENT.md` Β§10 |
1297
+ | the live topology, env vars, cold start | `release/repo/docs/DEPLOYMENT.md` |
1298
+ | the operations manual | [../OPERATIONS.md](OPERATIONS.md) |
1299
+ | the factual inventory | `release/CURRENT_RELEASE_STATE.md` |
1300
+
1301
+ ### Cross-references
1302
+
1303
+ | For | See |
1304
+ |---|---|
1305
+ | the frozen config and `Config.hash == 78f1e3700da15aa1` | [07 β€” Configuration and Freeze](architecture/07-configuration-freeze.md) |
1306
+ | the specialist contract in depth | [05 β€” Specialists](architecture/05-specialists.md) |
1307
+ | the router's five heads and the label space | [04 β€” Router](architecture/04-router.md) |
1308
+ | the evidence and confidence stages | [06 β€” Evidence and Confidence](architecture/06-evidence-and-confidence.md) |
1309
+ | the request lifecycle and the nine-state spine | [03 β€” Request Lifecycle](architecture/03-request-lifecycle.md) |
1310
+ | the four endpoints and error codes | [08 β€” The API Contract](architecture/08-api-contract.md) |
1311
+ | how to operate the live stack | [../OPERATIONS.md](OPERATIONS.md) |
1312
+ | per-artifact hyperparameters | [../TRAINING.md](TRAINING.md) |
1313
+ | what a third party can and cannot reproduce | [../REPRODUCIBILITY.md](REPRODUCIBILITY.md) |
1314
+
1315
+ ---
1316
+
1317
+ > **Chapter summary.** Python 3.11+ and a CPU are sufficient; there is no CUDA requirement, and all
1318
+ > placement is `.to(device)`, never `.cuda()`. The single registry is `configs/base.yaml`, its hash
1319
+ > **`78f1e3700da15aa1`** is frozen, and **editing the config moves the hash and invalidates every artifact
1320
+ > keyed to it** β€” the one non-reversible action in the repository. `core/schemas.py` is the binding
1321
+ > contract and every specialist returns exactly `SpecialistResult`; the specialist interface is four
1322
+ > methods in `specialists/base.py`, and a new specialist is added by implementing it and adding a
1323
+ > `SpecialistSpec` row keyed on **capability**, not name. Tests run from the venv, targeted, because the
1324
+ > full suite trips the sandbox's bulk-delete guard. Four artifacts train locally; two β€” change-VQA and the
1325
+ > VLM LoRA β€” train on an external GPU and are not one-command reproducible. Eight development traps are
1326
+ > recorded, of which the most dangerous are the stale untracked `deploy/`, the annotation-scope trap that
1327
+ > turns a `Request` parameter into a required query parameter, and the Chrome focus trap that produces a
1328
+ > silent false pass. `NOT RUN`/`OPEN` items include B-07 and B-02 (both `OPEN`), the absent `LICENSE`, the
1329
+ > stale monorepo README and `hf/` docs, and several `UNKNOWN β€” not established from the available
1330
+ > evidence` gaps.