thundercode commited on
Commit
496299b
·
verified ·
1 Parent(s): e1c7447

release: add docs/SERVING.md

Browse files
Files changed (1) hide show
  1. docs/SERVING.md +1301 -0
docs/SERVING.md ADDED
@@ -0,0 +1,1301 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # SatQuery AI — Serving
2
+
3
+ **Chapter scope.** This chapter documents the SatQuery AI inference service end to end: how the
4
+ service is composed, the four-endpoint contract it exposes and the `/api/*` mirror of that contract,
5
+ the entrypoint requirements any host must satisfy, the lazy model-loading model, the annotation-scope
6
+ defect that once made every upload fail with a `422`, the ephemeral asset store, the error taxonomy
7
+ and its machine codes, the tunnel transport, and — stated plainly — what the service does *not* do.
8
+
9
+ **Grounding.** Every claim below comes from a file that was read for this chapter, cited inline, e.g.
10
+ `(app/space_app.py)`, `(docs/API_CONTRACT.md §2.4)`. No endpoint, field, environment variable, status
11
+ code, or number is invented. Where the evidence does not exist, the text says exactly:
12
+ `UNKNOWN — not established from the available evidence`.
13
+
14
+ **Status vocabulary** follows `release/DOCS_STYLE_GUIDE.md` §2: `IMPLEMENTED` · `VERIFIED` ·
15
+ `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` · `OPEN` · `RESOLVED` ·
16
+ `CLOSED`.
17
+
18
+ **Nothing in this chapter is a system-level accuracy claim.** Per `release/DOCS_STYLE_GUIDE.md` §3
19
+ there is **no end-to-end benchmark** for SatQuery AI. This chapter describes a *service*; it does not
20
+ score one.
21
+
22
+ ---
23
+
24
+ ## 1. What the serving tier is
25
+
26
+ SatQuery AI's serving tier is a **Python HTTP service** that exposes the project's analysis capability
27
+ over four endpoints. It is built on FastAPI/Starlette, it is served by `uvicorn`, and it is composed
28
+ by three modules:
29
+
30
+ | Module | Role |
31
+ |---|---|
32
+ | `app/serving.py` | The **composition root**: builds a deployable registry and controller, wiring trained artifacts through the registry's `builders=` seam. |
33
+ | `app/space_app.py` | The **HTTP application**: builds the FastAPI app (`build_space_app()`), owns the four routes, the asset store, and the error handlers. |
34
+ | `app/deployment.py` | The **capability adapter**: turns internal registry state into the public capability vocabulary and produces the health and capabilities payloads. |
35
+
36
+ Around those three sit:
37
+
38
+ - `core/controller.py` — `AnalysisController`, the control tier that runs the pipeline.
39
+ - `core/registry.py` — `SpecialistRegistry`, which discovers specialists from a spec table and builds
40
+ them lazily.
41
+ - `core/planner.py` — `PolicyPlanner`, the deterministic policy planner.
42
+ - `core/errors.py` — the error taxonomy (23 codes) and the path scrubber.
43
+ - `gateway/app.py` + `gateway/policy.py` — the gateway (an optional front tier; see §3.3).
44
+ - `deploy/codespace/serve.py` — the 26-line process entrypoint that calls `build_space_app()`.
45
+ - `deploy/codespace/launch.sh` — the launcher that starts the service and its tunnel agent.
46
+ - `deploy/render/main.py` — the Render orchestrator that exposes the `/api/*` mirror.
47
+
48
+ The service's job is narrow and worth stating: **accept an analysis request, run the pipeline, return a
49
+ `ResultEnvelope`.** It does not render a UI, it does not stream, and it does not persist results. §11
50
+ lists what it does not do in full.
51
+
52
+ ---
53
+
54
+ ## 2. The composition root: `app/serving.py`
55
+
56
+ `app/serving.py` is 275 lines. Its module docstring calls itself "the public serving entry point" and
57
+ states that it is "the thin, public composition root that wires a *deployable* controller".
58
+
59
+ ### 2.1 The three wired artifacts
60
+
61
+ The module declares three module-level `Path` constants. Each is a *repo-local artifact identity*, not
62
+ a config key:
63
+
64
+ | Constant | Path (relative to `REPO_ROOT`) |
65
+ |---|---|
66
+ | `CHANGE_CHECKPOINT` | `artifacts/change/levir_change_v001/head.pt` |
67
+ | `CHANGE_VQA_HEAD` | `artifacts/change_vqa/run/head.pt` |
68
+ | `FUSION_HEAD` | `artifacts/optical_sar/fusion_head_production_v001/head.pt` |
69
+
70
+ Their documented identities, as stated in the module's comments:
71
+
72
+ - **`CHANGE_CHECKPOINT`** — "The trained, benchmarked change head (test pooled IoU 0.8122)." It is
73
+ "the single source of truth for where serving looks for it; tests monkeypatch this to simulate an
74
+ absent artifact." It is also the detector whose features `scripts/prepare_change_vqa.py` builds, "so
75
+ training and serving share it." A test
76
+ (`tests/unit/test_app_serving.py::test_serving_and_preparation_share_one_stanet_checkpoint`) keeps
77
+ the two literals equal and fails if either side drifts.
78
+ - **`CHANGE_VQA_HEAD`** — "The R-02 change-VQA reasoning head." Written by
79
+ `scripts/train_change_vqa.py` (`--output-dir`, default `artifacts/change_vqa/run`) and read by
80
+ `scripts/evaluate_change_vqa.py` (`DEFAULT_CHECKPOINT`, the same path). It "does NOT exist in a fresh
81
+ checkout: it is produced by the external Kaggle run and returned to the maintainer". Absent ⇒ the
82
+ specialist constructs, reports itself unavailable, and answers nothing.
83
+ - **`FUSION_HEAD`** — "The verified production optical-SAR fusion head (Phase 12). Its identity is
84
+ pre-registered, not inferred: sha256
85
+ `785815729a3a39fc34dc41894efaf00d8739365d970a3f830a326e68ae888dab`, 14,427,457 bytes, 1,201,711
86
+ parameters, and the checkpoint self-identifies with the embedded `config_hash`
87
+ `78f1e3700da15aa1` and `arm='A'`."
88
+
89
+ Note the last one carefully: the fusion head's *self-identification* carries the same frozen config
90
+ hash `78f1e3700da15aa1` that `release/DOCS_STYLE_GUIDE.md` §3 records as the project's frozen config
91
+ hash. The artifact and the config agree by construction.
92
+
93
+ ### 2.2 Why artifacts are wired through `builders=`, not through config
94
+
95
+ This is the single most important design decision in the serving tier, and `app/serving.py` documents
96
+ it at length. The mechanism:
97
+
98
+ `Config.hash` (`core/config.py:79-80`) is a **sha256 over the WHOLE registry**. Adding one key moves
99
+ the hash. The shipped change head records `78f1e3700da15aa1` in
100
+ `artifacts/change/levir_change_v001/model_metadata.json`, and `scripts/eval_change.py` **refuses to
101
+ score on a hash drift (exit 3)**.
102
+
103
+ Therefore: editing `configs/base.yaml` to point at a trained head would **invalidate the project's own
104
+ benchmark number**. The supported wiring path is instead the registry's `builders=` override
105
+ (`core/registry.py:420-433`), keyed by spec name — `"change"` (`core/registry.py:204-213`) — and it is
106
+ a **call-site argument, not config**, so the hash is untouched.
107
+
108
+ The registry calls a builder as `builder(self.config, **kwargs)` and only passes config keys that
109
+ resolve (`_builder_kwargs`, `core/registry.py:435-453`). Since `change.checkpoint_path` is unset,
110
+ nothing arrives and the real builder would degrade; the override supplies the missing argument.
111
+
112
+ Stated as a general rule: **in this project, a serving-side artifact path may not be added to
113
+ `configs/base.yaml`, because doing so would move the frozen config hash and invalidate every benchmark
114
+ number keyed to it.** The `builders=` seam is the hash-exempt channel for such paths.
115
+
116
+ ### 2.3 Degrade, do not crash
117
+
118
+ The module's docstring states the governing principle: "A serving path must run even when the artifact
119
+ is absent."
120
+
121
+ The distinction it enforces is precise:
122
+
123
+ - **Absent** is a *deployment case*. When the checkpoint does not exist the module applies NO override,
124
+ so the registry resolves the real builder with no `checkpoint_path` and the documented `DEGRADED`
125
+ contract applies (`specialists/change/specialist.py`).
126
+ - **Corrupt** is a *defect*. The real builder still surfaces it as `ModelLoadError` — "the two are
127
+ deliberately not conflated."
128
+
129
+ The `change_vqa` override is applied **unconditionally**, because its builder's contract is
130
+ finer-grained: a missing head and a missing detector are both *named* refusals
131
+ (`ChangeVQASpecialist.has_head`, `unavailable_reason`), so wiring it can never turn "absent" into a
132
+ crash. Both paths are passed as `None` when the file does not exist, "which the builder reads as
133
+ 'artifact genuinely absent' rather than 'path I was told about is broken'."
134
+
135
+ ### 2.4 The three builders
136
+
137
+ **`_wired_change_builder(config, **kwargs)`** — imports `build_change_specialist` lazily ("so importing
138
+ `app.serving` stays cheap and does not pull the model stack (torch) into a process that never serves a
139
+ change query"), sets `kwargs["checkpoint_path"] = str(CHANGE_CHECKPOINT)`, and delegates.
140
+
141
+ **`_wired_change_vqa_builder(config, **kwargs)`** — exists to close a **train/serve skew** finding
142
+ (named "finding F2" in the code). The skew: `scripts/prepare_change_vqa.py` builds its change features
143
+ from the *trained* STANet (`DEFAULT_CHANGE_CHECKPOINT`), but serving had no equivalent wiring —
144
+ `change.checkpoint_path` is unset in `configs/base.yaml`, so the registry passed no `checkpoint_path`
145
+ and the specialist "would construct an UNTRAINED STANet and answer from a representation the head was
146
+ never fitted on." The class's `feature_spec_mismatch()` already refused to answer on that skew, "so
147
+ the failure was loud rather than silent — but a deployment that can never answer is still not a
148
+ deployment." The override supplies the SAME checkpoint `_wired_change_builder` uses, so the detector
149
+ backing a change answer and the detector behind the head's training features are "one artifact by
150
+ construction; the spec check stays armed as the second line of defence, not the only one."
151
+
152
+ **`_wired_optical_sar_builder(config, **kwargs)`** — exists to close a different structural defect:
153
+ "the encoder was unreachable by default." The mechanism, quoted from the module:
154
+ `specialists/optical_sar/specialist.py:946` builds CROMA only when handed a `checkpoint_path` that
155
+ exists:
156
+
157
+ ```python
158
+ if checkpoint_path is not None and Path(checkpoint_path).exists():
159
+ ```
160
+
161
+ `croma.checkpoint_path` is not in `configs/base.yaml` — and must not be, or `Config.hash` moves — so
162
+ the registry passed no `checkpoint_path`, the gate was `False`, and the default serving composition
163
+ ran with `encoder=None`. The registry then correctly reported `DEGRADED` ("no encoder; running on
164
+ fallback"): "a deployment that could never answer an optical/SAR question. The checkpoint was present
165
+ on disk the whole time; nothing asked for it."
166
+
167
+ The fix resolves the checkpoint from the PINNED identity through the hash-exempt channel
168
+ (`croma.resolve_checkpoint_path`: **env → config → pinned Hub cache, offline first**), then hands it to
169
+ the real builder. The resolution is recorded as `source` and **logged**, not attached to the returned
170
+ specialist — with an explicit reason given in the code: "An unread attribute on a production object is
171
+ how a contract quietly grows a second, undocumented shape; and it must not be published either,
172
+ because the trace reaches the client and v1 has no auth."
173
+
174
+ That last sentence is a design principle worth extracting: **anything the trace carries is public,
175
+ because v1 has no auth.** So composition-time facts that must not leak are logged rather than attached.
176
+
177
+ The fusion head is wired the same way for the same reason: `has_head`
178
+ (`specialists/optical_sar/specialist.py:175-187`) is `False` without it, so the capability would stay
179
+ `DEGRADED` even with the encoder loaded. Absent head ⇒ `None` ⇒ degrade.
180
+
181
+ ### 2.5 `build_serving_registry()`
182
+
183
+ ```python
184
+ def build_serving_registry(config=None, *, device=None) -> SpecialistRegistry
185
+ ```
186
+
187
+ - `config` — the central `core.config.Config`. "Loaded unchanged when omitted; never mutated."
188
+ - `device` — a torch device string; defaults to `config.device_preference`.
189
+ - Returns "a `SpecialistRegistry` that constructs nothing yet (`discover()` reads a spec table)."
190
+
191
+ Its body builds a `builders` dict:
192
+
193
+ ```python
194
+ builders = {
195
+ "change_vqa": _wired_change_vqa_builder,
196
+ "optical_sar": _wired_optical_sar_builder,
197
+ }
198
+ if CHANGE_CHECKPOINT.exists():
199
+ builders["change"] = _wired_change_builder
200
+ return SpecialistRegistry.discover(cfg, device=device, builders=builders)
201
+ ```
202
+
203
+ Note the asymmetry and the reason for it, which the code comments on: `change_vqa` and `optical_sar`
204
+ are registered **UNCONDITIONALLY**, unlike `change`. The comment explains: "The resolver decides at
205
+ build time whether an artifact exists, so gating registration on a path this module does not know yet
206
+ would be circular. Wiring it cannot turn 'absent' into a crash: the resolver returns `None`, the
207
+ builder degrades, and a construction failure is retained as an `UNAVAILABLE` entry by
208
+ `SpecialistRegistry.build` rather than escaping."
209
+
210
+ So: **`change` is gated on the checkpoint existing; `change_vqa` and `optical_sar` are not, because
211
+ their builders accept `None` as "absent".**
212
+
213
+ ### 2.6 `build_serving_controller()`
214
+
215
+ ```python
216
+ def build_serving_controller(config=None, *, device=None) -> AnalysisController
217
+ ```
218
+
219
+ It is "Constructed with `registry=`, `planner=` and `config=` only."
220
+
221
+ ```python
222
+ registry = build_serving_registry(cfg, device=device)
223
+ return AnalysisController(
224
+ registry=registry,
225
+ planner=PolicyPlanner(registry),
226
+ config=cfg,
227
+ )
228
+ ```
229
+
230
+ The critical documented consequence: **"No router is attached, so a caller drives it with
231
+ `AnalysisRequest(..., force_task=...)`; a natural-language router can be supplied by the caller's own
232
+ composition if the router weights are available."**
233
+
234
+ This is the single most important behavioural fact about the serving tier's request handling: **the
235
+ deployed service is driven by an explicit `force_task`, not by natural-language routing.** It explains
236
+ why the frontend's `interpret()` (see the `FRONTEND.md` chapter) does the lexical routing in the
237
+ browser and then sends a `force_task`: the browser-side interpretation is what fills the gap left by
238
+ the deliberately router-less serving composition.
239
+
240
+ `__all__` exports `CHANGE_CHECKPOINT`, `CHANGE_VQA_HEAD`, `FUSION_HEAD`, `build_serving_controller`,
241
+ and `build_serving_registry`.
242
+
243
+ ---
244
+
245
+ ## 3. The HTTP application: `app/space_app.py`
246
+
247
+ `app/space_app.py` is 736 lines and owns the HTTP surface.
248
+
249
+ ### 3.1 The four routes
250
+
251
+ `build_space_app()` assembles a FastAPI application with four routes:
252
+
253
+ | Method | Path | Kind | Notes |
254
+ |---|---|---|---|
255
+ | `GET` | `/v1/health` | cheap | Health block; includes `device` and `gpu_available`. |
256
+ | `GET` | `/v1/capabilities` | cheap | Capability block; per-task availability and reasons. |
257
+ | `POST` | `/v1/analyze` | **COSTLY** | Runs the pipeline; returns a `ResultEnvelope`. |
258
+ | `POST` | `/v1/assets` | **COSTLY** | Uploads an asset; returns an opaque `asset_id`. |
259
+
260
+ The "cheap vs COSTLY" distinction is not decoration: `gateway/app.py` declares
261
+
262
+ ```python
263
+ COSTLY_ROUTES = ("/v1/analyze", "/v1/assets")
264
+ ```
265
+
266
+ and the gateway's policy (`gateway/policy.py`) applies its body-size caps, file-size caps, rate limit,
267
+ and upstream timeout with those routes in mind. A cheap route can be polled; a COSTLY route cannot.
268
+ (§3.3 covers the gateway.)
269
+
270
+ `build_space_app()` also installs two error handlers:
271
+
272
+ - a `StarletteHTTPException` handler, and
273
+ - a generic `Exception` handler (recorded in the deployment docs as **F-12b**).
274
+
275
+ The generic handler matters: without it, an unhandled exception would return a framework-default body
276
+ that leaks internals. With it, the service returns a translated error. See §8.
277
+
278
+ ### 3.2 The ZeroGPU duration map
279
+
280
+ The module declares a per-task duration budget used when the service is hosted on a ZeroGPU-style
281
+ platform that requires an advance duration declaration:
282
+
283
+ | Task | Duration |
284
+ |---|---|
285
+ | `vqa` | 20 |
286
+ | `caption` | 20 |
287
+ | `grounding` | 45 |
288
+ | `change` | 30 |
289
+ | `optical_sar` | 45 |
290
+ | `change_vqa` | 30 |
291
+
292
+ The helper `decorate_gpu()` applies the declaration, and `_spaces_module()` resolves the platform
293
+ module. The numbers are the declared *budgets*, not measured latencies; the captured grounding run
294
+ records a measured `step_001` of 209.873 ms (see the `FRONTEND.md` chapter §7.4), which is a single
295
+ step's timing, not a task duration, and the two are not comparable.
296
+
297
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §3.4 documents this same map as the "ZeroGPU duration map". On the
298
+ **active** topology the service runs on a CPU Codespace (`SATQUERY_DEVICE=cpu`, per
299
+ `deploy/codespace/launch.sh` and `docs/DEPLOYMENT_TOPOLOGY.md` §5), where the GPU decoration is inert.
300
+
301
+ ### 3.3 The gateway and the `/api/*` mirror
302
+
303
+ There are two front-facing surfaces, and it is important not to conflate them.
304
+
305
+ **(a) The gateway (`gateway/app.py`).** A thin front tier that proxies a **4-route allowlist** to the
306
+ inference service. Its declarations:
307
+
308
+ | Symbol | Value | Meaning |
309
+ |---|---|---|
310
+ | `PROXIED_ROUTES` | 4 routes | The allowlist. |
311
+ | `BLOCKED_ROUTES` | empty | Nothing is explicitly blocked. |
312
+ | `COSTLY_ROUTES` | `("/v1/analyze", "/v1/assets")` | The routes that cost real work. |
313
+
314
+ It exposes `/v1/gateway/health` (its own health, distinct from `/v1/health`), installs a
315
+ `StarletteHTTPException` handler (F-3), and proxies the four routes. Notable mechanisms inside
316
+ `_proxy()`:
317
+
318
+ - **F-2** — it strips client CORS headers and *asserts* that none remain (`_is_cors_header()`,
319
+ `_CORS_HEADER_PREFIX`). This prevents a client from injecting an `Access-Control-*` header that the
320
+ gateway would then pass upstream.
321
+ - **F-6** — it applies a **streaming cap** on the response body rather than buffering unbounded.
322
+ - It deliberately **does not retry** (there is an explicit no-retry comment): a retry of a COSTLY route
323
+ would double the work.
324
+ - `_read_body_bounded()` (F-9) is a thin adapter that bounds the request body it reads.
325
+ - `_client_ip()` derives the client IP (used by the rate limiter), and `_env()` reads configuration.
326
+
327
+ The gateway is an **optional** front tier. Its module docstring notes it is unimportable in a
328
+ sandbox — i.e. it is written to be deployed, not imported by test runners — and the module-level `app`
329
+ is created inside a `try/except` for that reason.
330
+
331
+ **(b) The Render orchestrator (`deploy/render/main.py`, 532 lines).** The orchestrator exposes the
332
+ `/api/*` mirror of the four endpoints:
333
+
334
+ | Orchestrator route | Mirrors |
335
+ |---|---|
336
+ | `/api/health` | `/v1/health` |
337
+ | `/api/infer` | `/v1/analyze` |
338
+ | `/api/capabilities` | `/v1/capabilities` |
339
+ | `/api/assets` | `/v1/assets` |
340
+
341
+ This is the surface the frontend actually calls: `SQ.ENDPOINTS` is
342
+ `{assets:'/assets', infer:'/infer', capabilities:'/capabilities', health:'/health'}`
343
+ (`frontend/assets/js/live.js`) and the default base is `/api`, so the frontend's `/api/infer` maps to
344
+ the orchestrator's `/api/infer`, which maps to the service's `/v1/analyze`. Note the name change:
345
+ **the frontend says "infer"; the service says "analyze"; they are the same endpoint.**
346
+
347
+ The orchestrator's internals:
348
+
349
+ | Symbol | Behaviour |
350
+ |---|---|
351
+ | `_github_token()` | Reads the GitHub token used to wake the Codespace. |
352
+ | `_codespace_name()` | Reads and **strips** the Codespace name — the strip is the fix for the B-02 trailing-`\n` defect (see §12). |
353
+ | `_codespace_port()` | Defaults to `8000`. |
354
+ | `_wake_timeout_s()` | Defaults to `120`. |
355
+ | `_upstream_timeout_s()` | Defaults to `90`. |
356
+ | `_DEV_ORIGINS`, `_PRODUCTION_ORIGINS` | `_PRODUCTION_ORIGINS = ("https://satquery.pages.dev",)`; `_allowed_origins()` composes the CORS allowlist. |
357
+ | `OrchestratorError`, `WakeTimeout`, `OrchestratorConfigError`, `OrchestratorUpstreamError` | The orchestrator's own error types. |
358
+ | `_envelope()` | Wraps a response/error into the orchestrator's envelope shape. |
359
+ | `ensure_codespace_up()` | Wakes the Codespace if it is asleep (the wake sequence). |
360
+ | `_proxy()` | Forwards the request upstream. |
361
+ | `create_app()` | Builds the app with the four routes. |
362
+ | `_handle_orchestrator_error()` | Translates an orchestrator error into a response. |
363
+
364
+ **Documented drift, recorded not hidden.** `deploy/render/main.py`'s own docstring notes that it is
365
+ **superseded by the tunnel design** per the delivery documents, while remaining the source present in
366
+ this working copy. The deployed backend is the `SatQuery-Backend` repository (`main.py`, 768 lines,
367
+ with a tunnel), whose deployed HEAD is `89d80eaddec5` (`release/DOCS_STYLE_GUIDE.md` §3). The local
368
+ `deploy/render/main.py` therefore does **not** carry the tunnel implementation. See §9 and §12.
369
+
370
+ `render.yaml` declares the orchestrator service concretely:
371
+
372
+ ```yaml
373
+ startCommand: uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT
374
+ healthCheckPath: /api/health
375
+ ```
376
+
377
+ with environment variables `PORT`, `SATQUERY_ALLOWED_ORIGINS`, `GITHUB_TOKEN`, `CODESPACE_NAME`,
378
+ `CODESPACE_PORT` (`"8000"`), `SATQUERY_DEVICE` (`"cpu"`), `SATQUERY_WAKE_TIMEOUT_S` (`"120"`), and
379
+ `SATQUERY_UPSTREAM_TIMEOUT_S` (`"90"`). The plan is free, and **all secret values are declared
380
+ `sync: false`** — i.e. they are injected by the platform, not committed. (No value is reproduced in
381
+ this chapter; per the release rules, this documentation contains no credentials.)
382
+
383
+ ### 3.4 The `/v1` vs `/api` naming table
384
+
385
+ Because two naming schemes coexist, here is the mapping in one place:
386
+
387
+ | Concept | Service (`/v1`) | Orchestrator mirror (`/api`) |
388
+ |---|---|---|
389
+ | Health | `GET /v1/health` | `GET /api/health` |
390
+ | Capabilities | `GET /v1/capabilities` | `GET /api/capabilities` |
391
+ | Analysis | `POST /v1/analyze` | `POST /api/infer` |
392
+ | Asset upload | `POST /v1/assets` | `POST /api/assets` |
393
+ | Gateway's own health | `GET /v1/gateway/health` | — |
394
+
395
+ The `/v1/` prefix is the service's versioned contract (`docs/API_CONTRACT.md` §1). The `/api/` prefix
396
+ is the orchestrator's mirror. A client that speaks `/api/infer` is speaking to the mirror, not to the
397
+ service.
398
+
399
+ ---
400
+
401
+ ## 4. The four-endpoint contract in detail
402
+
403
+ `docs/API_CONTRACT.md` is the frozen, frontend-facing contract (917 lines). This section summarises
404
+ what it pins, because the service must satisfy it exactly.
405
+
406
+ ### 4.1 Conventions, and the one schema exception
407
+
408
+ `docs/API_CONTRACT.md` §1.1: unknown fields are **rejected**. The schemas use Pydantic
409
+ `extra="forbid"`, with exactly **one** exception: `GeoMetadata` is `extra="allow"`. The reason is that
410
+ geospatial metadata is an open set — a raster may carry CRS, transform, resolution, and arbitrary
411
+ derived fields — so forbidding extras there would reject legitimate metadata rather than protect the
412
+ contract.
413
+
414
+ The consequence for a client: sending an unexpected field on any *other* model is a validation error,
415
+ not a silently-ignored field. This is a deliberate strictness choice, and it is why the contract is
416
+ worth reading before writing a client.
417
+
418
+ ### 4.2 `GET /v1/health` (§2.1)
419
+
420
+ Returns the health block. Two fields are worth pinning:
421
+
422
+ - **`device`** is a **closed set**, validated by the F-8 rule in `app/deployment.py`: the legal values
423
+ are `_LEGAL_DEVICES = {cpu, cuda, mps}`. `_effective_device()` returns `None` for an unrecognised
424
+ device rather than echoing it back. So a client can rely on `device` being one of three values or
425
+ absent.
426
+ - **`gpu_available: false` is normal on ZeroGPU.** The contract records the *measured* degraded output,
427
+ and states that a `false` here is not a fault on that platform.
428
+
429
+ The reason to state this in the docs at all: a naive client would treat `gpu_available: false` as an
430
+ error. The contract says otherwise.
431
+
432
+ ### 4.3 `GET /v1/capabilities` (§2.2, §2.3, §2.3.1)
433
+
434
+ Returns the capability block: per-task availability plus a `reason` when a task is unavailable.
435
+
436
+ - `reason` is **required when `available: false`**. A capability block that said "unavailable" without
437
+ saying why would be less useful than one that names the missing artifact.
438
+ - **`modalities` appears only on `optical_sar`.** `app/deployment.py` declares `_MODALITIES` with only
439
+ `optical_sar` in it, so no other task carries a `modalities` field.
440
+
441
+ **§2.3 / §2.3.1 — the five-word vocabulary, and why `loaded`/`degraded` are never emitted.** The
442
+ public capability vocabulary has five states, and `app/deployment.py` translates internal registry
443
+ states into them via `CONTRACT_STATES` (5) and `REGISTRY_TO_CONTRACT`. The internal registry states are
444
+ `AVAILABLE` / `DEGRADED` / `UNAVAILABLE` (`core/registry.py`), and the registry has
445
+ `PLANABLE_STATES` marking which of those the planner may plan against.
446
+
447
+ The important negative fact: the public contract **never emits the words `loaded` or `degraded`**. The
448
+ internal vocabulary and the public vocabulary are deliberately different, and the translation is the
449
+ adapter's job. A client that wrote `if status == 'degraded'` would be reading a word the contract does
450
+ not use.
451
+
452
+ ### 4.4 `POST /v1/analyze` (§2.4)
453
+
454
+ Accepts an `AnalysisRequest` and returns a `ResultEnvelope`.
455
+
456
+ **Multipart is NOT implemented.** This is stated in the contract and it constrains every client: an
457
+ asset is uploaded separately to `/v1/assets`, and the analysis request references it by `asset_id`.
458
+ A client that tried to send the image inline as a multipart part would be rejected. This is why the
459
+ frontend's upload is a raw-bytes POST and why the analysis request is JSON
460
+ (`frontend/assets/js/live.js`; `FRONTEND.md` §5.5, §6.8.1).
461
+
462
+ The request carries the task and, in the serving composition, a `force_task` (see §2.6 — the deployed
463
+ controller has no router attached).
464
+
465
+ The response's fields that the frontend must read are enumerated in the contract: the answer, the
466
+ evidence, the regions, the confidence (raw and calibrated), the timings, the provenance (run id,
467
+ policy, protocol, schema), the geospatial block, and the warnings. The captured envelope on
468
+ `frontend/assets/data/anatomy-run.js` is a real instance of this shape (`FRONTEND.md` §7.4).
469
+
470
+ **Artifact refs are `null` in v1.** The contract records an explicit ruling (F-16) that artifact
471
+ references are `null` — the service does not return a URL or a handle to a produced artifact in v1.
472
+ This is a capability limit, not an oversight, and a client must not depend on an artifact ref being
473
+ present.
474
+
475
+ ### 4.5 `POST /v1/assets` (§2.5)
476
+
477
+ Uploads an asset and returns an opaque `asset_id`. The contract records the design as "Option A" and
478
+ pins:
479
+
480
+ | Property | Value |
481
+ |---|---|
482
+ | `asset_id` opacity | The client must treat the handle as opaque. |
483
+ | Size cap | Enforced (F-6 / F-7). |
484
+ | Content-type allowlist | Five types. |
485
+ | Retries | Documented. |
486
+ | Lifetime | The handle is **ephemeral** with a TTL. |
487
+
488
+ The service-side implementation of all five is in `app/space_app.py` (§7).
489
+
490
+ ### 4.6 Enums (§3)
491
+
492
+ | Enum | Cardinality | Values |
493
+ |---|---|---|
494
+ | `Task` | **7** | The task vocabulary. |
495
+ | `CoordinateSystem` | **3** | The coordinate-system vocabulary. |
496
+ | `Modality` | **4** | The modality vocabulary. |
497
+
498
+ Seven tasks is worth noting because `core/registry.py`'s `default_specs()` declares **six** specialists
499
+ (`vqa`, `caption`, `grounding`, `change`, `change_vqa`, `optical_sar`). The `Task` enum having seven
500
+ values while six specialists exist means the enum is the *request* vocabulary and the spec table is the
501
+ *implementation* vocabulary; the difference is a task the request enum names but that no specialist
502
+ serves directly. Which specific value accounts for the difference:
503
+ `UNKNOWN — not established from the available evidence` (the enum's member list was not read
504
+ verbatim for this chapter; only its cardinality is recorded here).
505
+
506
+ ### 4.7 The confidence contract (§4)
507
+
508
+ The contract documents:
509
+
510
+ - **The measured ECE caveat.** Calibration's ECE went **0.013755 → 0.014929 — worse**. The transform is
511
+ retained only because it is in the frozen config (`release/DOCS_STYLE_GUIDE.md` §3).
512
+ - **`T = 0.9772731820958189`** — the temperature.
513
+ - **16,441 Val rows** — the calibration sample count. This is the same figure the captured envelope
514
+ records as `calibration_samples: 16441.0` (`frontend/assets/data/anatomy-run.js`; `FRONTEND.md`
515
+ §7.4). The public page and the contract agree.
516
+
517
+ The honest reading of this section: **the service returns a calibrated confidence, and the calibration
518
+ is documented to have made ECE slightly worse.** A client must not present the calibrated confidence as
519
+ an accuracy. Per the style guide, there is no end-to-end benchmark, so a per-run confidence is a
520
+ per-run confidence.
521
+
522
+ ### 4.8 The error contract (§5)
523
+
524
+ See §8 for the full treatment. The contract's §5.1 gives the status map, §5.2 the full 23-code
525
+ taxonomy, and §5.3 the gateway-origin `rate_limited` code. §5.1 also records the **trailing-slash 307
526
+ footgun** (Starlette `redirect_slashes`), which is why a client should compose exact URLs.
527
+
528
+ ### 4.9 Latency, quotas, auth, CORS (§6, §7, §7.1)
529
+
530
+ - **§6 — latency and quotas.** The contract records the latency expectations and any quotas.
531
+ - **§7 — auth: none.** v1 has **no authentication**. This is a first-class design fact with
532
+ consequences that appear all over the codebase: it is why `core/errors.py` scrubs paths (F-15), why
533
+ `app/serving.py` logs rather than attaches the CORS/checkpoint `source`, and why the trace must not
534
+ carry anything sensitive.
535
+ - **§7.1 — CORS.** CORS is configured on the orchestrator, whose `_PRODUCTION_ORIGINS` includes the
536
+ Pages origin `https://satquery.pages.dev` (`deploy/render/main.py`). The gateway additionally strips
537
+ client-supplied CORS headers (F-2, `gateway/app.py`).
538
+
539
+ **§9 — the minimal integration checklist.** The contract closes with a checklist for a new client,
540
+ which is the shortest path for anyone writing against this service.
541
+
542
+ ---
543
+
544
+ ## 5. Entrypoint requirements
545
+
546
+ Any host that runs this service must satisfy five requirements. `docs/DEPLOYMENT_ARCHITECTURE.md`
547
+ §3.3 enumerates them, and §3.3.1 adds a sixth consideration (a single capability authority). The
548
+ requirements are:
549
+
550
+ 1. **A Python process with the project's dependencies.** `deploy/codespace/launch.sh` performs a
551
+ preflight dependency check for `yaml`, `pydantic`, `fastapi`, `uvicorn`, and `httpx` before it
552
+ starts anything. A host that does not have these cannot start the service.
553
+ 2. **A callable application object.** `deploy/codespace/serve.py` is the reference implementation:
554
+
555
+ ```python
556
+ app = build_space_app()
557
+ uvicorn.run(app, host="0.0.0.0", port=port)
558
+ ```
559
+
560
+ with `port = int(os.environ.get("PORT", "8000"))`. The entrypoint therefore must (a) build the app
561
+ via `build_space_app()` and (b) bind a port from the environment with a default.
562
+ 3. **A port binding on `0.0.0.0`.** The reference binds `0.0.0.0`, not `127.0.0.1`, so the service is
563
+ reachable from outside the process's own namespace.
564
+ 4. **An environment that can reach the artifacts** (or degrade cleanly without them). Because
565
+ `app/serving.py` wires artifacts through the `builders=` seam and degrades when they are absent, a
566
+ host without the artifacts still *starts* — it just reports the affected capabilities as
567
+ unavailable. This is what makes "degrade, do not crash" a deployment property rather than a slogan.
568
+ 5. **A health-checkable endpoint.** The orchestrator's `render.yaml` sets
569
+ `healthCheckPath: /api/health`, so the platform probes that path. A host that cannot answer a health
570
+ probe will be considered unhealthy and restarted or removed from rotation.
571
+
572
+ **§3.3.1 — a single capability authority.** The architecture doc adds that there must be exactly one
573
+ authority for capability state: `app/deployment.py`. The registry knows internal state
574
+ (`AVAILABLE`/`DEGRADED`/`UNAVAILABLE`); the deployment adapter translates it into the public five-word
575
+ vocabulary. A second place that decided capability state would create two answers to "is this task
576
+ available?", which is exactly the kind of drift the project's discipline forbids.
577
+
578
+ ### 5.1 The launcher: `deploy/codespace/launch.sh`
579
+
580
+ `deploy/codespace/launch.sh` is 194 lines and is the reference launcher. Its steps, as read:
581
+
582
+ 1. **Preflight dependency checks** for `yaml`, `pydantic`, `fastapi`, `uvicorn`, `httpx`.
583
+ 2. **Port and stamp guards** — so two launchers do not fight over the same port and a stale stamp does
584
+ not mislead.
585
+ 3. **`_restart_serve()`** — starts the service with
586
+ `setsid nohup python deploy/codespace/serve.py`, i.e. detached from the launcher's terminal so the
587
+ service survives the shell.
588
+ 4. **The supervised tunnel-agent loop** — starts the tunnel agent with
589
+ `setsid nohup bash -c '… python deploy/codespace/tunnel_agent.py …'` and supervises it, restarting
590
+ it if it exits. See §9.
591
+ 5. **Environment** — exports `SATQUERY_DEVICE=cpu`, `SATQUERY_ASSET_ENABLED=1`,
592
+ `SATQUERY_ASSET_DIR=/tmp/satquery-assets`, and
593
+ `SATQUERY_HUB_URL=https://satquery-backend-m4yv.onrender.com`.
594
+ 6. **Verification** — step 3 verifies the agent "announced to hub", so the launcher does not report
595
+ success merely because the process started.
596
+
597
+ Note that `SATQUERY_DEVICE=cpu` in the launcher matches `SATQUERY_DEVICE: "cpu"` in `render.yaml` and
598
+ the CPU-first reconciliation in `docs/DEPLOYMENT_TOPOLOGY.md` §5.
599
+
600
+ > **Honesty note.** `deploy/codespace/launch.sh` references `deploy/codespace/tunnel_agent.py`, and
601
+ > `docs/DEPLOYMENT_TOPOLOGY.md` §2 and `docs/FINAL_DELIVERY_TODO.md` §1.3 both name that file as part
602
+ > of the `SatQuery-Inference` deployment. **That file does not exist in this working copy.** The local
603
+ > `deploy/` directory is stale/untracked (`docs/FINAL_DELIVERY_REPORT.md` §6 records "local `deploy/`
604
+ > stale"; `docs/FINAL_DELIVERY_TODO.md` §5 records the corresponding blocker). What the tunnel agent
605
+ > does is therefore described in §9 from the *evidence that does exist* (the launcher's invocation, the
606
+ > topology doc's description, and the transport value in the captured envelope), and the agent's
607
+ > internals are marked `UNKNOWN — not established from the available evidence`.
608
+
609
+ ---
610
+
611
+ ## 6. Lazy model loading and `cache_max_models: 1`
612
+
613
+ ### 6.1 The lazy-loading contract
614
+
615
+ The serving tier does **not** load models at import time. Two mechanisms enforce this:
616
+
617
+ **(a) `build_serving_registry()` constructs nothing.** Its own docstring says it returns "a
618
+ `SpecialistRegistry` that constructs nothing yet (`discover()` reads a spec table)." `core/registry.py`
619
+ confirms the shape: `default_specs()` returns six spec rows, and `discover()` reads that table. The
620
+ spec table is data; no model is instantiated by reading it.
621
+
622
+ **(b) Builders import lazily.** `_wired_change_builder` imports `build_change_specialist` *inside* the
623
+ function, with the stated reason: "so importing `app.serving` stays cheap and does not pull the model
624
+ stack (torch) into a process that never serves a change query." The same pattern appears in the other
625
+ two builders. The consequence is that `import app.serving` does not import torch at all.
626
+
627
+ This matters because `core/registry.py`'s `SpecialistRegistry.__init__` has a torch-import path
628
+ (`self.device = device or config.device_preference`). Keeping the import inside builders means a
629
+ process that only serves, say, capabilities never pays for torch.
630
+
631
+ ### 6.2 The spec table and lazy construction
632
+
633
+ `core/registry.py`:
634
+
635
+ | Symbol | Role |
636
+ |---|---|
637
+ | `RegistryState` | `AVAILABLE` / `DEGRADED` / `UNAVAILABLE`. |
638
+ | `PLANABLE_STATES` | Which states the planner may plan against. |
639
+ | `SpecialistSpec` | One row of the spec table. |
640
+ | `default_specs()` | **Six** rows: `vqa`, `caption`, `grounding`, `change`, `change_vqa`, `optical_sar`. |
641
+ | `RegistryEntry` | The registry's record for one spec; `to_trace()` **scrubs `detail`**. |
642
+ | `SpecialistRegistry.discover()` | Reads the spec table; constructs nothing. |
643
+ | `SpecialistRegistry.available()` | Returns `tuple(sorted(self._specs))` — a sorted tuple, so the order is stable. |
644
+ | `SpecialistRegistry.specs()` | The spec table. |
645
+ | `SpecialistRegistry.entry(name)` | One entry. |
646
+
647
+ The `default_specs()` rows carry their asset requirements: `requires_assets` is `1`, `2`, or `None`
648
+ depending on the task (a single-image task needs 1; a paired task needs 2; a task that needs no asset
649
+ has `None`). They also carry `optional_config_keys`. These are the same requirements the frontend's
650
+ `PAIRED_TASKS = {change, change_vqa, optical_sar}` reflects on the client side (`FRONTEND.md` §5.3) —
651
+ and it is worth noting the two lists agree: the three paired tasks are exactly the three whose
652
+ `requires_assets` is 2.
653
+
654
+ `RegistryEntry.to_trace()` scrubbing `detail` is a privacy mechanism: the trace reaches the client, and
655
+ v1 has no auth, so the entry's raw detail does not travel.
656
+
657
+ ### 6.3 `cache_max_models: 1`
658
+
659
+ The serving configuration caps the model cache at **one** model. The consequence is the important part:
660
+ with a cache of one, serving a task evicts the previously loaded model. A sequence of requests across
661
+ two tasks therefore loads and evicts repeatedly rather than holding both.
662
+
663
+ Why this is the right default for this deployment: the active host is a CPU Codespace
664
+ (`SATQUERY_DEVICE=cpu`) with limited memory, and the project's posture is CPU-first
665
+ (`docs/DEPLOYMENT_TOPOLOGY.md` §5). Holding several models resident would risk memory exhaustion, and
666
+ an `OutOfMemoryError` is a defined failure in the taxonomy (`core/errors.py`, `out_of_memory`,
667
+ `recoverable=True`) precisely because memory pressure is an expected condition.
668
+
669
+ The honest cost of `cache_max_models: 1`: **a multi-task workload pays repeated model-load cost.** This
670
+ is a latency property, not a correctness one. It is stated here rather than omitted because it is a
671
+ real consequence a reader should know before benchmarking latency.
672
+
673
+ Where `cache_max_models` is declared in config: `UNKNOWN — not established from the available
674
+ evidence` for the exact key location (the value's *effect* — a cap of one — is what is documented
675
+ here; the config file line was not read for this chapter).
676
+
677
+ ---
678
+
679
+ ## 7. The asset store
680
+
681
+ ### 7.1 Purpose and shape
682
+
683
+ `POST /v1/assets` exists because multipart is not implemented (§4.4). An asset is uploaded once,
684
+ receives an opaque handle, and the handle is referenced by the analysis request.
685
+
686
+ `app/space_app.py` implements the store with a module-level cache (`_ASSET_STORE`) and an accessor
687
+ `get_asset_store()`. Its configuration comes from environment variables:
688
+
689
+ | Helper | Default | Meaning |
690
+ |---|---|---|
691
+ | `_asset_max_files()` | **32** | Maximum number of files held. |
692
+ | `_asset_ttl_seconds()` | **900.0** | Handle lifetime, in seconds (15 minutes). |
693
+ | `_asset_root()` | system tempdir fallback | Where asset bytes are written. |
694
+ | `_asset_max_file_bytes()` | — | Per-file byte cap; **refuses a non-positive or non-integer value** (F-7). |
695
+
696
+ `_ALLOWED_ASSET_CONTENT_TYPES` declares the **five** accepted content types, matching
697
+ `docs/API_CONTRACT.md` §2.5 and the client's `SQ.CONTENT_TYPES` (`frontend/assets/js/live.js`:
698
+ `tif`, `tiff`, `png`, `jpg`, `jpeg`).
699
+
700
+ ### 7.2 Fail-closed availability
701
+
702
+ The store's availability gate is `_asset_store_available()`, which requires **BOTH**:
703
+
704
+ - `SATQUERY_ASSET_ENABLED`, and
705
+ - `SATQUERY_ASSET_DIR`.
706
+
707
+ If either is missing, the store is unavailable and `POST /v1/assets` returns **503**. This is
708
+ **fail-closed**: the service refuses uploads rather than accepting them into a store it cannot
709
+ guarantee. That is the correct posture for an ephemeral store — a handle issued by a store that cannot
710
+ serve it back is worse than no handle.
711
+
712
+ The launcher (`deploy/codespace/launch.sh`) sets both:
713
+
714
+ ```
715
+ SATQUERY_ASSET_ENABLED=1
716
+ SATQUERY_ASSET_DIR=/tmp/satquery-assets
717
+ ```
718
+
719
+ so the deployed Codespace has the store enabled with a temp-dir root. On a host where the variables are
720
+ absent, the 503 is the expected behaviour and the frontend surfaces it via `translateError()`
721
+ (`FRONTEND.md` §6.7).
722
+
723
+ ### 7.3 Handle opacity and lifetime
724
+
725
+ The handle is `asset_<32 hex>` — 32 hex characters, which is `secrets.token_hex(16)`. Two properties
726
+ follow:
727
+
728
+ 1. **It is unguessable.** 16 random bytes (128 bits) means a client cannot enumerate handles.
729
+ 2. **It is opaque.** Nothing about the underlying file is encoded in it. The client must not parse it,
730
+ and the frontend's `uploadAsset()` explicitly *asserts* the handle exists and passes it back
731
+ unexamined (`frontend/assets/js/live.js`; `FRONTEND.md` §6.8.1).
732
+
733
+ The **TTL** (default 900.0 s) and the **file cap** (default 32) together mean the store is a short-lived
734
+ staging area, not a database. The practical consequences for a client:
735
+
736
+ - An upload and its analysis must happen **within the TTL**.
737
+ - A workload that uploads more than 32 files concurrently will hit the cap.
738
+ - Nothing survives a service restart: the store is in-memory plus a temp directory.
739
+
740
+ ### 7.4 Path scrubbing on the way out
741
+
742
+ `core/errors.py` implements **F-15** path scrubbing (`scrub_paths()`), which is directly relevant to the
743
+ asset store because asset errors are client-visible. The mechanism:
744
+
745
+ - `_WINDOWS_DRIVE_PATH`, `_UNC_PATH`, and `_POSIX_PATH` match **absolute** paths.
746
+ - The replacement keeps only the **final component** ("basename reduction"), so
747
+ `"cannot read C:\\a\\b\\weights.pt"` becomes `"cannot read weights.pt"` — "still diagnostic, no
748
+ longer a location disclosure."
749
+ - **Relative paths are deliberately not matched**, and the reason is documented: "A rule broad enough to
750
+ catch `artifacts/change/head.pt` also catches `and/or` and the path segments of a URL, and a
751
+ scrubber that mangles ordinary prose is a worse defect than the disclosure it fixes. The measured
752
+ leaks are all absolute."
753
+ - **URLs are left intact on purpose**: `https://github.com/antofuller/CROMA` appears inside one of the
754
+ very messages this scrubs, and mangling it "would be a worse defect than the one being repaired."
755
+ The `_POSIX_PATH` lookbehind refuses to start a match immediately after `:` or `/`, which is the
756
+ mechanism that keeps the URL intact.
757
+
758
+ The module also records the *history* of the fix, which is instructive: a blunt replacement of the
759
+ whole message with a generic string was tried first, "but it discarded path-free diagnostics the client
760
+ can legitimately act on (`... has no builder 'build_x'`, `no GPU in this dimension`), and three
761
+ existing tests that pin exactly those diagnostics failed. **A fix that forces legitimate tests to be
762
+ weakened is aimed at the wrong granularity.**"
763
+
764
+ F-15's owner ruling (2026-09-23) is quoted in the file: *"sanitize all client-facing exception
765
+ messages; retain full exception details only in server-side diagnostics."* The reason it was needed:
766
+ exception messages in this repo routinely embed an absolute path (e.g. `specialists/optical_sar/croma.py`
767
+ raises a message naming a vendored directory; `specialists/change/stanet.py` raises one naming an
768
+ encoder-weights path), and those strings reach client-visible fields — and v1 has no auth.
769
+
770
+ ---
771
+
772
+ ## 8. Error translation and machine codes
773
+
774
+ ### 8.1 The taxonomy: 23 codes
775
+
776
+ `core/errors.py` (316 lines) defines the taxonomy. Every failure the system can produce is one of these
777
+ codes, and the module's docstring states the rule plainly: "Never raise a bare Exception from specialist
778
+ or controller code."
779
+
780
+ The base class is `SatQueryError`, whose attributes are documented in the file:
781
+
782
+ | Attribute | Meaning |
783
+ |---|---|
784
+ | `code` | Stable machine-readable identifier, used in traces. |
785
+ | `user_message` | Text safe to show the operator. |
786
+ | `detail` | Technical detail for the execution trace (**never chain-of-thought**). |
787
+ | `recoverable` | Whether the controller may continue with a fallback. |
788
+
789
+ It carries a `to_trace()` method returning `{code, detail, recoverable, context}`.
790
+
791
+ The taxonomy, grouped as the file groups it:
792
+
793
+ **Input / raster.**
794
+
795
+ | Code | Class | `recoverable` |
796
+ |---|---|---|
797
+ | `input_error` | `InputError` | default |
798
+ | `raster_read_error` | `RasterReadError` | default |
799
+ | `missing_crs` | `MissingCRSError` | **True** — "Degraded, not fatal: non-geospatial analysis may still be possible." |
800
+ | `unsupported_bands` | `UnsupportedBandsError` | default |
801
+ | `oversized_image` | `OversizedImageError` | **True** — recoverable via downscale. |
802
+
803
+ **Pairing.**
804
+
805
+ | Code | Class | Note |
806
+ |---|---|---|
807
+ | `pair_incompatible` | `PairCompatibilityError` | — |
808
+ | `pair_misaligned` | `PairMisalignmentError` | Subclass of the above. |
809
+ | `temporal_pair_invalid` | `TemporalPairError` | Subclass of the above. |
810
+
811
+ **Routing / planning.**
812
+
813
+ | Code | Class |
814
+ |---|---|
815
+ | `routing_error` | `RoutingError` |
816
+ | `unsupported_query` | `UnsupportedQueryError` |
817
+ | `invalid_request` | `InvalidRequestError` |
818
+ | `workflow_plan_error` | `WorkflowPlanError` |
819
+
820
+ **Specialists.**
821
+
822
+ | Code | Class | `recoverable` |
823
+ |---|---|---|
824
+ | `specialist_error` | `SpecialistError` | default |
825
+ | `model_load_error` | `ModelLoadError` | default |
826
+ | `model_unavailable` | `ModelUnavailableError` | **True** — "the controller degrades the workflow." |
827
+ | `out_of_memory` | `OutOfMemoryError` | **True** — retry at lower resolution. |
828
+ | `specialist_timeout` | `SpecialistTimeoutError` | **True** |
829
+
830
+ **Output integrity.**
831
+
832
+ | Code | Class |
833
+ |---|---|
834
+ | `schema_validation_error` | `SchemaValidationError` |
835
+ | `coordinate_error` | `CoordinateError` |
836
+ | `confidence_range_error` | `ConfidenceRangeError` |
837
+
838
+ **Leakage / evaluation.**
839
+
840
+ | Code | Class |
841
+ |---|---|
842
+ | `leakage_violation` | `LeakageError` |
843
+ | `benchmark_freeze_error` | `BenchmarkFreezeError` |
844
+
845
+ That is **23 codes**, matching `__all__`'s 23 entries and the "23-code taxonomy" recorded in
846
+ `docs/API_CONTRACT.md` §5.2 and `gateway/policy.py`'s `_CODE_STATUS`.
847
+
848
+ ### 8.2 The `specialist_timeout` recoverability correction
849
+
850
+ One entry deserves its own treatment because the file documents a *defect* it corrected.
851
+ `SpecialistTimeoutError` was inheriting `recoverable=False` from `SatQueryError`, and the file explains
852
+ why that was wrong, with two independent reasons:
853
+
854
+ 1. `docs/API_CONTRACT.md` is the frozen frontend-facing contract, and §5.1 **maps 504 with
855
+ `recoverable: true`**. A frontend that reads `recoverable: false` "will not offer a retry for the one
856
+ failure the contract explicitly tells it to retry."
857
+ 2. The plan's Failure Matrix (§57) lists Timeout with the recovery "abort specialist" and the fallback
858
+ "partial result" — i.e. the controller continues rather than failing the request. A terminal
859
+ `recoverable=False` contradicts that.
860
+
861
+ The file also records *why the defect was invisible from the inside*: "the controller currently only
862
+ reuses `.code` for its budget-skip trace entry (`core/controller.py:464`), so nothing in the pipeline
863
+ constructed this class and the wrong default was never observable from the inside — only from a
864
+ client." This is a good example of the project's practice of documenting *how* a bug could hide.
865
+
866
+ ### 8.3 The status map and the gateway-origin code
867
+
868
+ `gateway/policy.py` declares `_CODE_STATUS`, the map from each of the 23 codes to an HTTP status, and:
869
+
870
+ ```python
871
+ GATEWAY_ORIGIN_CODES = {"rate_limited"}
872
+ _CODE_STATUS["rate_limited"] = 429
873
+ ```
874
+
875
+ So `rate_limited` is a **gateway-origin** code: it is not one of the 23 taxonomy codes produced by the
876
+ service, it is produced by the gateway's own rate limiter, and it maps to **429**. `docs/API_CONTRACT.md`
877
+ §5.3 records it separately for exactly this reason — a client should understand that a 429 came from the
878
+ gateway, not from the analysis pipeline.
879
+
880
+ `DEFECT_CODES` (5) names the codes that indicate a *defect* rather than a normal failure. The
881
+ distinction matters: a defect code means the system did something wrong, whereas most codes describe a
882
+ legitimate condition (a missing CRS, a bad upload, a timeout).
883
+
884
+ ### 8.4 `translate_error()`
885
+
886
+ `translate_error()` maps an error to its client-facing form. Its role in the architecture is stated in
887
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §2.3: **the code is passed unchanged.** The gateway translates the
888
+ *shape* (into its envelope, with a request id) but does not rewrite the code — so a client sees the
889
+ service's own code, not a gateway-invented one.
890
+
891
+ Supporting symbols: `_REQUEST_ID_RE` (validates a request id's shape) and `new_request_id()` (mints
892
+ one). A request id is what makes a client-side report correlatable with a server-side log.
893
+
894
+ ### 8.5 `GatewayConfig` and its validators
895
+
896
+ `gateway/policy.py` declares `GatewayConfig` with these defaults:
897
+
898
+ | Field | Default |
899
+ |---|---|
900
+ | `max_body_bytes` | 8 MiB |
901
+ | `max_file_bytes` | 4 MiB |
902
+ | `rate_limit_per_ip` | 10 |
903
+ | `rate_limit_window_s` | 60.0 |
904
+ | `upstream_timeout_s` | 90.0 |
905
+ | `allowed_content_types` | 5 |
906
+
907
+ Its `__post_init__` validators reject a misconfiguration rather than letting it fail later:
908
+
909
+ - an origin with a **trailing slash** is rejected,
910
+ - an empty value is rejected,
911
+ - a `*` wildcard is rejected,
912
+ - and a timeout that is **not greater than 45** is rejected.
913
+
914
+ The last one is interesting: the 45-second floor is tied to the GPU duration map's longest budget
915
+ (`grounding` and `optical_sar` are both **45** in `app/space_app.py`'s `GPU_DURATIONS`). An upstream
916
+ timeout below the longest task budget would cut off a legitimate run, so the validator forbids it.
917
+
918
+ Note the relationship between the two size caps: the gateway's `max_file_bytes` (4 MiB) is *smaller*
919
+ than its `max_body_bytes` (8 MiB), which is coherent — a file cap inside a body cap.
920
+
921
+ ### 8.6 The F-12b generic handler
922
+
923
+ Back in `app/space_app.py`, the generic `Exception` handler (F-12b) is what makes the taxonomy
924
+ *airtight at the edge*: an exception that escaped the pipeline's own handling is still translated into a
925
+ response rather than surfacing as a framework default. `docs/DEPLOYMENT_ARCHITECTURE.md` §5 lists F-12
926
+ and F-12b among the failure modes, alongside F-11, F-13, F-14, F-15, F-15b, F-15c, F-16, F-16c, F-17,
927
+ F-18, and F-19. (F-15c is the gateway's transport-failure detail, `_TRANSPORT_FAILURES` /
928
+ `_transport_failure_detail()` in `gateway/app.py`.)
929
+
930
+ ---
931
+
932
+ ## 9. The tunnel agent and the transport
933
+
934
+ ### 9.1 Why a tunnel exists
935
+
936
+ The service runs on a host (a GitHub Codespace) that is not directly reachable at a stable public
937
+ address in the way a normal web service is. The orchestrator on Render is the public face. Something
938
+ must carry a request from the orchestrator to the service. That "something" is the transport, and the
939
+ captured envelope records the transport it used:
940
+
941
+ ```
942
+ transport: "tunnel"
943
+ ```
944
+
945
+ (`frontend/assets/data/anatomy-run.js`; `FRONTEND.md` §7.4). The frontend's live client also reads a
946
+ transport response header, `x-satquery-transport` (`frontend/assets/js/live.js`), which is how a client
947
+ can see which transport carried its response.
948
+
949
+ ### 9.2 The two transports
950
+
951
+ `docs/DEPLOYMENT_TOPOLOGY.md` and the delivery documents describe two transport designs:
952
+
953
+ 1. **Forwarded-port transport.** The orchestrator reaches the Codespace through a forwarded port. In
954
+ this design a private repository yields a **302** (a redirect), which is why a 302 is a documented
955
+ behaviour rather than an error.
956
+ 2. **Outbound tunnel transport.** The service-side agent **long-polls** `POST /tunnel/agent` to the
957
+ hub, so the connection is *outbound* from the Codespace. An outbound tunnel avoids requiring the
958
+ Codespace to be reachable inbound, which is the property that makes it robust on a platform that
959
+ does not expose inbound ports.
960
+
961
+ The tunnel design supersedes the forwarded-port design: `deploy/render/main.py`'s docstring says it is
962
+ superseded by the tunnel design per the delivery documents, and the deployed backend repository is the
963
+ one that carries the tunnel.
964
+
965
+ ### 9.3 The agent's role, and what is known about it
966
+
967
+ The agent's role, assembled from the evidence that exists:
968
+
969
+ - **`deploy/codespace/launch.sh` starts and supervises it** with
970
+ `setsid nohup bash -c '… python deploy/codespace/tunnel_agent.py …'`, detached from the launcher's
971
+ terminal and restarted if it exits. So the agent is a long-running process, not a one-shot.
972
+ - **It announces to the hub.** The launcher's step 3 verifies that the agent "announced to hub", so
973
+ announcing is part of the agent's contract and the launcher treats a failed announcement as a failed
974
+ launch.
975
+ - **`SATQUERY_HUB_URL` names the hub.** The launcher sets it to
976
+ `https://satquery-backend-m4yv.onrender.com`, which is the same host the frontend's
977
+ `<meta name="satquery-api-base">` names (`frontend/mission.html`). So the hub, the orchestrator, and
978
+ the API base are one host.
979
+ - **It is supervised, and it is started after the service.** The launcher starts the service
980
+ (`_restart_serve()`) and *then* starts the agent, which is the correct order: an agent that
981
+ announced before the service was listening would advertise a dead endpoint.
982
+
983
+ **What the agent does internally** — its poll loop, its request framing, its reconnection strategy, its
984
+ handling of a hub restart — is `UNKNOWN — not established from the available evidence`, because
985
+ `deploy/codespace/tunnel_agent.py` does not exist in this working copy (§5.1's honesty note). The
986
+ deployed backend repository (HEAD `89d80eaddec5`) is where the tunnel implementation lives, and it was
987
+ not read for this chapter.
988
+
989
+ ### 9.4 B-07: tunnel gaps, patch prepared but not deployed
990
+
991
+ Per `release/DOCS_STYLE_GUIDE.md` §3 and `docs/FINAL_DELIVERY_TODO.md` §5: **B-07 is OPEN. It is tunnel
992
+ gaps, and the patch is prepared but NOT deployed.** This status must not be upgraded. The correct
993
+ statement is:
994
+
995
+ > B-07 — tunnel gaps. Patch prepared, not deployed. **OPEN.**
996
+
997
+ The consequence for a reader: the tunnel transport works well enough to have carried the runs recorded
998
+ in the delivery documents (including the captured `run_d124d8b9adea`, whose `transport` is `"tunnel"`),
999
+ and it also has known gaps whose fix is written but not live. Both halves are true at once.
1000
+
1001
+ ---
1002
+
1003
+ ## 10. The deployment topology
1004
+
1005
+ ### 10.1 The active topology
1006
+
1007
+ `docs/DEPLOYMENT_TOPOLOGY.md` is the **active** topology document. Its components:
1008
+
1009
+ | Component | Host | Role |
1010
+ |---|---|---|
1011
+ | Static tier | Cloudflare Pages | The eleven pages (see `FRONTEND.md`). |
1012
+ | Public backend | Render (`satquery-orchestrator`) | The `/api/*` mirror; wake + proxy; CORS. |
1013
+ | Inference | GitHub Codespace | Runs the service (`build_space_app()`), CPU-first, plus the tunnel agent. |
1014
+ | Model artifacts | Hugging Face | Artifact hosting; also the public release surface. |
1015
+
1016
+ The document contains a Mermaid topology diagram and a **wake sequence**, plus §3's per-component
1017
+ responsibilities and environment variables, §4's five old blockers, §5's reconciliation (CPU-first),
1018
+ and §6's preconditions.
1019
+
1020
+ ### 10.2 Deployed HEADs
1021
+
1022
+ Per `release/DOCS_STYLE_GUIDE.md` §3:
1023
+
1024
+ | Component | Deployed HEAD |
1025
+ |---|---|
1026
+ | Frontend | `2d7ae53b482d` |
1027
+ | Backend | `89d80eaddec5` |
1028
+ | Inference | `5a0936ace491` |
1029
+
1030
+ ### 10.3 The measured live environment
1031
+
1032
+ `docs/DEPLOYMENT_TOPOLOGY.md` §3.2 records the **measured live env-var set**. Two entries in that
1033
+ section are worth flagging because the section also notes that some names listed historically are
1034
+ **not** in the live config: `SATQUERY_UPSTREAM_URL` and `HF_TOKEN` are named in the section's own prose
1035
+ while the section's measured note says they are not present. This is documentation drift inside the
1036
+ topology document, recorded here rather than propagated.
1037
+
1038
+ `docs/DEPLOYMENT_ARCHITECTURE.md` carries a superseded-topology banner and still names Railway /
1039
+ HF-Space hosts in its body while the active hosts are Render / Codespace. Both documents are kept, with
1040
+ the banner making the supersession explicit — which is the project's stated practice (mirroring
1041
+ `P10-T02`).
1042
+
1043
+ ### 10.4 `docs/DEPLOYMENT_ARCHITECTURE.md` §2 — gateway responsibilities
1044
+
1045
+ The architecture document's §2 enumerates the gateway's responsibilities and the 4-route allowlist, and
1046
+ §2.3 pins the error-translation rule (code passed unchanged). §3.1 assigns entrypoint ownership, §3.2
1047
+ lists constraints, §3.3 lists the five entrypoint requirements, §3.3.1 the single capability authority,
1048
+ §3.4 the ZeroGPU duration map, §4 the env-var vocabulary (a long table with F-6/F-7/F-8/F-9 notes), §5
1049
+ the failure-mode table (F-11…F-19), §6 what is excluded, §7 implementation status, and §8 deployment
1050
+ preconditions.
1051
+
1052
+ ### 10.5 The pipeline the service runs
1053
+
1054
+ The service's work is done by `core/controller.py`'s `AnalysisController.run()`, whose stages are:
1055
+
1056
+ ```
1057
+ RECEIVE → PARSE → VALIDATE → PLAN → EXECUTE → AGGREGATE → VERIFY → RESPOND
1058
+ ```
1059
+
1060
+ The captured grounding envelope's eight steps are `RECEIVE` → `RESPOND`, i.e. the same eight-stage
1061
+ pipeline (`frontend/assets/data/anatomy-run.js`). Notable details from `core/controller.py`:
1062
+
1063
+ - **`_asset_label`** — a basename reduction applied to asset labels, the same idiom as F-13/F-14 and
1064
+ the same idiom `core/errors.py::scrub_paths` uses for F-15. "One rule, one implementation, applied at
1065
+ every client-facing write site."
1066
+ - **F-19** — the registry is **re-snapshotted after execute**:
1067
+ `trace.parameters["registry"] = self.registry.describe()` is written *after* the EXECUTE stage, so the
1068
+ trace records the registry state that actually ran rather than the state at request entry.
1069
+ - **`_execute()`** — applies a **budget between steps**; and per F-15, sets
1070
+ `trace.errors[].message = user_message` (the sanitized message, not the raw detail).
1071
+ - **`_execute_one()`** — implements **F-20**, a producer-side repair for unhandled exceptions, so a
1072
+ specialist that raises something unexpected is still recorded as a result rather than escaping.
1073
+ - **`health()`** — **deprecated**: it "Constructs everything", and it was retired as the public path.
1074
+ This is why `app/deployment.py` owns the health payload instead: the public health path must be
1075
+ cheap, and a health check that constructs every model is not cheap.
1076
+ - `_route()`, `_resolve_assets()`, `_modalities()` — the routing, asset-resolution, and modality
1077
+ helpers.
1078
+
1079
+ ---
1080
+
1081
+ ## 11. What the service does NOT do
1082
+
1083
+ Stated explicitly, because the depth of §2–§10 could otherwise imply more capability than exists.
1084
+
1085
+ - **No Gradio GUI.** The service is an HTTP API. There is no Gradio interface in this serving tier; the
1086
+ user interface is the static frontend (`FRONTEND.md`), which talks to the service over HTTP. Whether
1087
+ a Gradio surface exists anywhere else in the project: `UNKNOWN — not established from the available
1088
+ evidence` for this chapter (the serving modules read contain no Gradio application).
1089
+ - **No streaming.** There is no server-sent-events or websocket channel. A request is answered with a
1090
+ single response. The frontend's eight-event display is driven *client-side* from that one response
1091
+ plus two headers (`X-SatQuery-State`, `x-satquery-transport`), not pushed from the server
1092
+ (`FRONTEND.md` §14).
1093
+ - **No batching.** A request is one analysis. There is no batch endpoint, and `POST /v1/analyze` takes
1094
+ one `AnalysisRequest`.
1095
+ - **No queue.** There is no job queue and no async job model: a COSTLY route does its work within the
1096
+ request, bounded by the upstream timeout (`upstream_timeout_s` default 90.0) and the gateway's
1097
+ timeout floor (> 45). This is why the gateway deliberately does **not** retry (`gateway/app.py`): a
1098
+ retry of a COSTLY route would duplicate work rather than dequeue it.
1099
+ - **No authentication.** v1 has no auth (`docs/API_CONTRACT.md` §7). This has downstream consequences
1100
+ throughout: path scrubbing (F-15), trace scrubbing (`RegistryEntry.to_trace()` scrubs `detail`), and
1101
+ logging-instead-of-attaching composition facts (`app/serving.py`).
1102
+ - **No multipart upload.** Assets are uploaded separately (`docs/API_CONTRACT.md` §2.4).
1103
+ - **No artifact refs.** Artifact references are `null` in v1 (the F-16 ruling).
1104
+ - **No persistence.** The asset store is ephemeral (TTL 900.0 s, cap 32 files) and there is no run
1105
+ store. A restart loses everything.
1106
+ - **No natural-language routing in the serving composition.** `build_serving_controller()` attaches no
1107
+ router, so a caller drives it with `force_task` (`app/serving.py`; §2.6).
1108
+ - **No model preloading.** Models load lazily and the cache holds one (`cache_max_models: 1`; §6).
1109
+ - **No end-to-end benchmark.** Per `release/DOCS_STYLE_GUIDE.md` §3 this does not exist, and no
1110
+ system-level accuracy is claimed anywhere in this chapter.
1111
+
1112
+ ---
1113
+
1114
+ ## 12. Status summary and blockers
1115
+
1116
+ ### 12.1 Status by subsystem
1117
+
1118
+ | Subsystem | Status |
1119
+ |---|---|
1120
+ | `app/serving.py` composition root (`build_serving_registry`, `build_serving_controller`) | `IMPLEMENTED` |
1121
+ | Artifact wiring via the `builders=` seam (change / change_vqa / optical_sar) | `IMPLEMENTED` |
1122
+ | `app/space_app.py` (`build_space_app()`, four routes, two error handlers) | `IMPLEMENTED` |
1123
+ | `app/deployment.py` capability adapter (two vocabularies, five contract states) | `IMPLEMENTED` |
1124
+ | Four-endpoint contract (`/v1/health`, `/v1/capabilities`, `/v1/analyze`, `/v1/assets`) | `IMPLEMENTED` |
1125
+ | `/api/*` orchestrator mirror | `IMPLEMENTED`; deployed backend HEAD `89d80eaddec5` |
1126
+ | Gateway (4-route allowlist, `COSTLY_ROUTES`, F-2/F-3/F-6/F-9) | `IMPLEMENTED` |
1127
+ | Lazy model loading; `cache_max_models: 1` | `IMPLEMENTED` |
1128
+ | Asset store (opaque handles, TTL, cap, allowlist, fail-closed 503) | `IMPLEMENTED` |
1129
+ | Error taxonomy (23 codes) + `_CODE_STATUS` + gateway-origin `rate_limited` | `IMPLEMENTED` |
1130
+ | Path scrubbing (F-15) | `IMPLEMENTED` |
1131
+ | Tunnel transport | `IMPLEMENTED`; carried `run_d124d8b9adea` (`transport: "tunnel"`) |
1132
+ | B-02 `codespace_name` trailing `\n` | Fixed in `deploy/render/main.py` via a strip; recorded as cosmetic, **OPEN** |
1133
+ | B-07 tunnel gaps | Patch prepared, **NOT deployed** — **OPEN** |
1134
+
1135
+ ### 12.2 The blockers, stated exactly
1136
+
1137
+ | ID | Statement | Status |
1138
+ |---|---|---|
1139
+ | **B-07** | Tunnel gaps. Patch prepared, not deployed. | **OPEN** — never to be upgraded. |
1140
+ | **B-02** | `codespace_name` trailing `\n`. Cosmetic. The orchestrator's `_codespace_name()` strips it. | **OPEN** (cosmetic) |
1141
+ | Local `deploy/` | The local `deploy/` directory is stale/untracked; `deploy/codespace/tunnel_agent.py` is absent; `deploy/render/main.py` is superseded by the deployed backend. | **KNOWN** (`docs/FINAL_DELIVERY_TODO.md` §5 B-03; `docs/FINAL_DELIVERY_REPORT.md` §6) |
1142
+ | Change capability | Recorded as degraded in the delivery documents at the time of writing. | `KNOWN` — per `docs/FINAL_DELIVERY_REPORT.md` §6 |
1143
+ | P2-T03 | Cosmetic. | **OPEN** (cosmetic) |
1144
+
1145
+ `docs/FINAL_DELIVERY_TODO.md` §5 records the full blocker register: B-01 **CLOSED**, B-02
1146
+ **DOWNGRADED**, B-03 **KNOWN**, B-04 **ACCEPTED**, B-05 **ACCEPTED**, B-06 **KNOWN**, B-07 **OPEN**,
1147
+ B-08 **CLOSED**. Note that B-01 (which `docs/FINAL_DELIVERY_REPORT.md` §6 records as HF BLOCKED at the
1148
+ time of that report) is **CLOSED** in the later TODO register — so the correct current statement is
1149
+ that B-01 is CLOSED, with the earlier report's BLOCKED status being superseded.
1150
+
1151
+ ### 12.3 The G-1 annotation-scope defect
1152
+
1153
+ This is the most instructive serving defect in the project and deserves its own treatment.
1154
+
1155
+ **The mechanism.** `app/space_app.py` uses `from __future__ import annotations`. Under that import,
1156
+ annotations are **strings**, resolved lazily by FastAPI via `eval` against a namespace. If a parameter's
1157
+ annotation names a type (`Request`) that is **bound in a narrower scope** than the function that FastAPI
1158
+ introspects, then FastAPI's `eval` resolves that name against the **wrong globals**. The name fails to
1159
+ resolve as a type, and FastAPI **silently reinterprets the parameter as a REQUIRED QUERY PARAMETER named
1160
+ `request`**.
1161
+
1162
+ **The symptom.** Every upload gets:
1163
+
1164
+ ```json
1165
+ 422 {"detail":[{"loc":["query","request"]}]}
1166
+ ```
1167
+
1168
+ This is the worst kind of bug: a **server-side** defect that presents as a **client-side** validation
1169
+ error. A client developer reads "missing required query parameter `request`" and concludes they
1170
+ mis-called the API. They did not.
1171
+
1172
+ **Why it is silent.** There is no exception at import time. The app builds. The route registers. Only
1173
+ the *interpretation* of the parameter changed, and it changed in a way that produces a plausible-looking
1174
+ error.
1175
+
1176
+ **The twin, and the asymmetry.** The related case is a **return annotation** naming `JSONResponse`. In
1177
+ that case the resolution failure does **not** degrade silently — it raises **`PydanticUndefinedAnnotation`**,
1178
+ and it raises **at import/definition time**, so `build_space_app()` is **never called at all**. The app
1179
+ therefore does not exist.
1180
+
1181
+ So the defect has two halves with **opposite** failure modes:
1182
+
1183
+ | Annotation position | Failure mode |
1184
+ |---|---|
1185
+ | **Parameter** annotation | **Silent.** The parameter is reinterpreted as a required query parameter. The app runs and every upload 422s. |
1186
+ | **Return** annotation | **Loud.** `PydanticUndefinedAnnotation` is raised before `build_space_app()` can be called; the app never starts. |
1187
+
1188
+ The asymmetry is why the defect is worth documenting: the *loud* half is easy to find (the app will not
1189
+ start), and the *silent* half is the dangerous one (the app starts and lies about why it is failing).
1190
+
1191
+ **The repair pattern.** `app/space_app.py` lines 55–91 carry module-scope comment blocks binding
1192
+ `Request`, `Response`, and `JSONResponse` at **module scope**, so that FastAPI's `eval` resolves the
1193
+ names against the module's globals. The gateway has the **twin** of this: `gateway/app.py` also binds
1194
+ `Request`, `Response`, and `JSONResponse` at module level for the same reason. The rule extracted:
1195
+
1196
+ > **Under `from __future__ import annotations`, every type used in a FastAPI route signature must be
1197
+ > bound at the module scope where the route function is defined — because FastAPI resolves annotations
1198
+ > by `eval` against that module's globals, and a narrower-scope binding resolves to nothing.**
1199
+
1200
+ The correct status for G-1: the **repair is IMPLEMENTED** (the module-scope bindings are present in both
1201
+ `app/space_app.py` and `gateway/app.py`). The **defect is RESOLVED** in the code read. Whether an
1202
+ earlier deployment ever served the silent-422 behaviour is a historical question: the recorded live
1203
+ validation ran 24 runs with 8/8 per pass (`release/DOCS_STYLE_GUIDE.md` §3), which is consistent with a
1204
+ working upload path in the deployed build — but the exact deployment at which the fix landed is
1205
+ `UNKNOWN — not established from the available evidence`.
1206
+
1207
+ ### 12.4 Other failure modes recorded in the architecture doc
1208
+
1209
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §5 lists the failure-mode table. The ones most relevant to serving:
1210
+
1211
+ | ID | Subject |
1212
+ |---|---|
1213
+ | F-6 | Streaming size cap (also `gateway/app.py` `_proxy()`). |
1214
+ | F-7 | `_asset_max_file_bytes()` refuses a non-positive or non-integer value. |
1215
+ | F-8 | `device` validation → `_effective_device()` returns `None` for an unrecognised value; `_LEGAL_DEVICES = {cpu, cuda, mps}`. |
1216
+ | F-9 | `_read_body_bounded()` in the gateway. |
1217
+ | F-11 | (per §5) |
1218
+ | F-12 / F-12b | The generic exception handler in `build_space_app()`. |
1219
+ | F-13 / F-14 | `_asset_label` basename reduction. |
1220
+ | F-15 / F-15b | Path scrubbing; the F-15b variant. |
1221
+ | F-15c | Gateway transport-failure detail (`_TRANSPORT_FAILURES`, `_transport_failure_detail()`). |
1222
+ | F-16 / F-16c | The artifact-refs-`null` ruling; the F-16c variant. |
1223
+ | F-17 / F-18 / F-19 | F-19 is the post-execute registry re-snapshot in `core/controller.py`. |
1224
+ | F-20 | Producer-side repair for unhandled exceptions in `_execute_one()`. |
1225
+
1226
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §4's env-var vocabulary table carries the F-6/F-7/F-8/F-9 notes
1227
+ inline, and §6 states what is excluded from the deployment, §7 its implementation status, and §8 the
1228
+ deployment preconditions.
1229
+
1230
+ ---
1231
+
1232
+ ## 13. NOT RUN / OPEN / BLOCKED (serving)
1233
+
1234
+ Per `release/DOCS_STYLE_GUIDE.md` §4, every doc ends with this list.
1235
+
1236
+ **NOT RUN**
1237
+ - No end-to-end benchmark of the service (project-wide fact per `release/DOCS_STYLE_GUIDE.md` §3; the
1238
+ service is not exempt, and no system-level accuracy is claimed).
1239
+ - No load/latency benchmark of the four endpoints under `cache_max_models: 1`.
1240
+ - No test of the tunnel under a hub restart.
1241
+ - No verification of the gateway's rate limiter under sustained load.
1242
+ - No verification of the asset store's cap (32) and TTL (900.0 s) boundaries end to end.
1243
+
1244
+ **OPEN**
1245
+ - **B-07 — tunnel gaps. Patch prepared, NOT deployed.** OPEN. (Never to be upgraded.)
1246
+ - **B-02 — `codespace_name` trailing `\n`. Cosmetic.** OPEN. (The strip is present in
1247
+ `deploy/render/main.py`.)
1248
+ - **P2-T03 — cosmetic.** OPEN.
1249
+ - **F-15 path scrubbing** — the *measured* leaks are all absolute paths; relative-path leaks were
1250
+ deliberately not covered. The scoping is documented as intentional; whether any relative-path leak
1251
+ exists is `UNKNOWN — not established from the available evidence`.
1252
+ - **Documentation drift inside the topology docs** — `docs/DEPLOYMENT_TOPOLOGY.md` §3.2 names
1253
+ `SATQUERY_UPSTREAM_URL` and `HF_TOKEN` while its own measured note says they are not in the live
1254
+ config; `docs/DEPLOYMENT_ARCHITECTURE.md` names Railway / HF-Space hosts under a superseded-topology
1255
+ banner. Recorded; OPEN as documentation debt.
1256
+ - **`deploy/codespace/tunnel_agent.py`** — referenced by `launch.sh` and two delivery docs, absent from
1257
+ this working copy. The agent's internals are
1258
+ `UNKNOWN — not established from the available evidence`.
1259
+ - **`Task` enum's seventh value** — the enum has seven values while six specialists are declared; which
1260
+ value accounts for the difference is `UNKNOWN — not established from the available evidence`.
1261
+ - **`cache_max_models` config key location** — the value's effect (a cap of one) is documented; the
1262
+ exact key location is `UNKNOWN — not established from the available evidence`.
1263
+ - **G-1's fix deployment point** — the repair is IMPLEMENTED in the code read; the deployment at which
1264
+ it landed is `UNKNOWN — not established from the available evidence`.
1265
+ - **B-01** — `CLOSED` per `docs/FINAL_DELIVERY_TODO.md` §5 (superseding the earlier report's BLOCKED
1266
+ status). Recorded here so it is not re-opened.
1267
+ - **No LICENSE file exists** — project-wide, OPEN (`release/DOCS_STYLE_GUIDE.md` §3).
1268
+
1269
+ **BLOCKED**
1270
+ - Nothing in the serving *code* read for this chapter is blocked.
1271
+ - **Deployment-level:** the local `deploy/` tree is stale/untracked, so the tunnel implementation
1272
+ cannot be read from this working copy — the corresponding investigation is BLOCKED on that tree being
1273
+ refreshed (or on the deployed backend repository being read instead).
1274
+ - **B-01 at the time of `docs/FINAL_DELIVERY_REPORT.md`** was BLOCKED (HF); it is CLOSED per the later
1275
+ TODO register. The earlier status is superseded, not deleted.
1276
+
1277
+ ---
1278
+
1279
+ ## 14. Where the evidence lives
1280
+
1281
+ | Claim area | Evidence file(s) |
1282
+ |---|---|
1283
+ | Composition root; the three artifact constants and their identities; the `builders=` seam and why config must not be edited; degrade-don't-crash; the three builders and the defects they close; `build_serving_registry()`; `build_serving_controller()` (no router → `force_task`) | `app/serving.py` |
1284
+ | HTTP application; `build_space_app()`; the four routes; the two error handlers (incl. F-12b); `GPU_DURATIONS`; `decorate_gpu()`; `_spaces_module()`; `get_controller()`; `describe_deployment()`; the asset-store helpers (`_asset_max_files()` 32, `_asset_ttl_seconds()` 900.0, `_asset_root()`, `_asset_max_file_bytes()` F-7, `_ALLOWED_ASSET_CONTENT_TYPES` 5, `_asset_store_available()` requiring both env vars); `main()` | `app/space_app.py` |
1285
+ | Capability adapter: `CONTRACT_STATES` (5), `REGISTRY_TO_CONTRACT`, `_REQUIREMENTS`, `_MISSING_REASONS`, `_HUB_REASONS`, `_optical_sar_artifacts()`, `_resolve_croma_checkpoint()`, `_requirement_artifacts()`, `_missing_shipped()`, `_hub_unconfigured()`, `_configured_path()`, `_HUB_BACKED`, `CapabilityReport`, `_MODALITIES`, `DeploymentReport`, `_schema_version()`, `_registry_capabilities()`, `_asset_count()`, `_artifact_evidence()`, `_report_for()`, `deployment_report()`, `_effective_device()` (F-8), `_LEGAL_DEVICES`, `_cuda_detected()`, `health_payload()`, `capabilities_payload()` | `app/deployment.py` |
1286
+ | Error taxonomy (23 codes), `SatQueryError` + `to_trace()`, the `specialist_timeout` recoverability correction, F-15 path scrubbing (`_WINDOWS_DRIVE_PATH`, `_UNC_PATH`, `_POSIX_PATH`, `scrub_paths()`) | `core/errors.py` |
1287
+ | `_CODE_STATUS` (23 codes), `DEFECT_CODES` (5), `GATEWAY_ORIGIN_CODES`, `rate_limited` → 429, `translate_error()`, `_REQUEST_ID_RE`, `new_request_id()`, `GatewayConfig` + validators | `gateway/policy.py` |
1288
+ | Gateway: `PROXIED_ROUTES` (4), `BLOCKED_ROUTES`, `COSTLY_ROUTES`, `/v1/gateway/health`, F-3 handler, `_read_body_bounded()` (F-9), `_proxy()` (F-2 CORS strip + assertion, F-6 streaming cap, no-retry), `_is_cors_header()`, `_CORS_HEADER_PREFIX`, `_client_ip()`, `_env()`, module-level `Request`/`Response`/`JSONResponse` bindings (the G-1 twin) | `gateway/app.py` |
1289
+ | Registry: `RegistryState`, `PLANABLE_STATES`, `SpecialistSpec`, `default_specs()` (6 rows, `requires_assets`), `RegistryEntry.to_trace()` scrubs `detail`, `discover()`, `available()`, `specs()`, `entry()`; the `builders=` override site (lines 420-433) and the spec-name key (lines 204-213); `_builder_kwargs` (lines 435-453) | `core/registry.py` |
1290
+ | Controller: the eight-stage pipeline; `_asset_label` (F-13/F-14); the F-19 post-execute registry re-snapshot; `health()` deprecated ("Constructs everything"); `_route()`, `_resolve_assets()`, `_modalities()`, `_execute()` (budget; F-15 `user_message`), `_execute_one()` (F-20) | `core/controller.py` |
1291
+ | The frozen contract: conventions + `extra="forbid"` / `GeoMetadata extra="allow"` (§1.1); health (§2.1, device closed set, `gpu_available: false` normal on ZeroGPU); capabilities (§2.2, §2.3, §2.3.1 five-word vocabulary, `modalities` only on optical_sar); analyze (§2.4, multipart NOT implemented, artifact refs `null` per F-16); assets (§2.5, opacity, caps, allowlist, lifetime); enums (§3, Task 7 / CoordinateSystem 3 / Modality 4); confidence (§4, ECE 0.013755→0.014929, T = 0.9772731820958189, 16,441 Val rows); errors (§5, §5.1 status map + 307 footgun, §5.2 23 codes, §5.3 `rate_limited`); latency/quotas (§6); auth (§7) + CORS (§7.1); status (§8); integration checklist (§9) | `docs/API_CONTRACT.md` |
1292
+ | Five entrypoint requirements; §3.3.1 single capability authority; gateway responsibilities + 4-route allowlist + COSTLY; §2.3 error translation (code unchanged); §3.4 ZeroGPU duration map; §4 env-var vocabulary; §5 failure-mode table (F-11…F-19); §6 exclusions; §7 status; §8 preconditions; superseded-topology banner | `docs/DEPLOYMENT_ARCHITECTURE.md` |
1293
+ | Active topology; components; Mermaid topology + wake sequence; §3 per-component responsibilities/env vars; §4 five old blockers; §5 CPU-first reconciliation; §6 preconditions | `docs/DEPLOYMENT_TOPOLOGY.md` |
1294
+ | Orchestrator: `_github_token()`, `_codespace_name()` (strip = B-02), `_codespace_port()` 8000, `_wake_timeout_s()` 120, `_upstream_timeout_s()` 90, `_DEV_ORIGINS`, `_PRODUCTION_ORIGINS`, `_allowed_origins()`, error classes, `_envelope()`, `ensure_codespace_up()`, `_proxy()`, `create_app()` (four `/api/*` routes), `_handle_orchestrator_error()`; the superseded-by-tunnel docstring | `deploy/render/main.py` |
1295
+ | Orchestrator service declaration: start command, `healthCheckPath: /api/health`, env-var names, `sync: false` on secrets | `render.yaml` |
1296
+ | Service entrypoint: `app = build_space_app()`; `uvicorn.run(host="0.0.0.0", port=...)`; `PORT` default 8000 | `deploy/codespace/serve.py` |
1297
+ | Launcher: preflight deps; port/stamp guards; `_restart_serve()`; the supervised tunnel-agent loop; the env vars (`SATQUERY_DEVICE=cpu`, `SATQUERY_ASSET_ENABLED=1`, `SATQUERY_ASSET_DIR`, `SATQUERY_HUB_URL`); the "announced to hub" verification | `deploy/codespace/launch.sh` |
1298
+ | Captured run: `run_id`, `transport: "tunnel"`, `config_hash`, confidence + `temperature` + `calibration_samples`, warnings, steps | `frontend/assets/data/anatomy-run.js` |
1299
+ | Deployed HEADs (`2d7ae53b482d`, `89d80eaddec5`, `5a0936ace491`); B-07 OPEN patch prepared not deployed; B-02 cosmetic OPEN; no E2E benchmark; live validation 24 runs / 0 mock nodes / 94.4444 % | `release/DOCS_STYLE_GUIDE.md` |
1300
+ | Commits; live topology; E2E run-id table; metrics; blockers; test results (94 + 183 passed); truthfulness statement | `docs/FINAL_DELIVERY_REPORT.md` |
1301
+ | Status board; artifact inventory; real measured metrics; nine known blockers (incl. item 9 Cloudflare concatenation); blocker register B-01…B-08; evidence register E-01…E-14; final verification checklist | `docs/FINAL_DELIVERY_TODO.md` |