thundercode commited on
Commit
495e039
·
verified ·
1 Parent(s): 50cf649

release: add docs/architecture/02-deployment-topology.md

Browse files
docs/architecture/02-deployment-topology.md ADDED
@@ -0,0 +1,1633 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # 02 — Deployment Topology
2
+
3
+ **Parent:** [Architecture hub](../ARCHITECTURE.md) · **Sibling:** [01 System overview](01-system-overview.md) ·
4
+ **Next:** [03 Request lifecycle](03-request-lifecycle.md)
5
+
6
+ **Status tags used in this document:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` ·
7
+ `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` · `OPEN` · `RESOLVED` · `CLOSED`.
8
+
9
+ > **One-paragraph summary.** SatQuery AI is deployed as **four tiers**: a static browser client on
10
+ > **Cloudflare Pages** (`satquery.pages.dev`), a thin stateless **gateway/orchestrator on Render**
11
+ > (`satquery-orchestrator` at `satquery-backend-m4yv.onrender.com`), a **FastAPI inference service inside
12
+ > a GitHub Codespace** (FastAPI on port `8000`), and **Hugging Face** as the project/model-presence tier.
13
+ > The gateway does not dial into the Codespace. Instead the Codespace **dials out** to the gateway over a
14
+ > long-poll **tunnel** (`POST /tunnel/agent`), because a *forwarded* Codespace port returns **HTTP 302**
15
+ > for a private repository. That inversion is the single most consequential decision in the topology, and
16
+ > it is why the deployment works at all with private repositories.
17
+
18
+ ---
19
+
20
+ ## 1. The four tiers
21
+
22
+ ### 1.1 Tier map
23
+
24
+ ```
25
+ USER / BROWSER
26
+ │ HTTPS
27
+ ▼
28
+ Cloudflare Pages — static frontend satquery.pages.dev
29
+ │ HTTPS, JSON, /api/*
30
+ ▼
31
+ Render — orchestrator / API gateway satquery-backend-m4yv.onrender.com
32
+ │ service: satquery-orchestrator
33
+ │ outbound long-poll (POST /tunnel/agent) ← direction is INVERTED
34
+ ▼
35
+ GitHub Codespace — FastAPI inference potential-space-trout-r4ppw969w45j2pvvw :8000
36
+ │ build_space_app(): /v1/health · /v1/capabilities · /v1/analyze · /v1/assets
37
+ │ specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
38
+ ▼
39
+ Hugging Face — project card + pinned model references
40
+ ```
41
+
42
+ Sources: `docs/DEPLOYMENT_TOPOLOGY.md` §1; `docs/FINAL_DELIVERY_TODO.md` §1.2 (the same ASCII topology
43
+ reproduced in the delivery single-source-of-truth); `docs/FINAL_DELIVERY_REPORT.md` §2.
44
+
45
+ ```mermaid
46
+ flowchart LR
47
+ U["Browser<br/>satquery.pages.dev"] -->|"HTTPS"| CF["Cloudflare Pages<br/>static frontend"]
48
+ CF -->|"HTTPS JSON /api/*"| R["Render<br/>satquery-orchestrator"]
49
+ R -->|"POST /tunnel/agent<br/>long-poll (outbound)"| A["Codespace tunnel agent"]
50
+ A -->|"http://127.0.0.1:8000"| I["FastAPI<br/>build_space_app()"]
51
+ I --> S[("SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet")]
52
+ I -.->|"model refs"| HF["Hugging Face<br/>project + pinned models"]
53
+ R -.->|"HF proxy path<br/>(not used in live config)"| HF
54
+ ```
55
+
56
+ ### 1.2 Tier responsibilities, at a glance
57
+
58
+ | Tier | Host / identity | Runs | Holds secrets? | Holds state? |
59
+ |---|---|---|---|---|
60
+ | Browser | the user's machine | `frontend/` static JS | **No** — never | no |
61
+ | Static | Cloudflare Pages, `satquery.pages.dev` | HTML/CSS/JS only | **No** | no |
62
+ | Gateway | Render, `satquery-orchestrator`, `satquery-backend-m4yv.onrender.com` | `deploy/render/main.py` (`SatQuery-Backend` in production) | **Yes** — `GITHUB_TOKEN` (and `HF_TOKEN` if the proxy path is used) | no |
63
+ | Inference | GitHub Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port `8000` | `app/space_app.py::build_space_app()` via `deploy/codespace/serve.py` | no gateway secrets | ephemeral asset store only |
64
+ | Presence | Hugging Face (`hf/`) | project README / model cards | no | no |
65
+
66
+ Sources: `docs/DEPLOYMENT_TOPOLOGY.md` §3.1–§3.4; `render.yaml`; `deploy/render/main.py`;
67
+ `app/space_app.py`; `.devcontainer/devcontainer.json`.
68
+
69
+ ### 1.3 Why exactly four tiers and not three
70
+
71
+ The plan forbids "unnecessary microservices" (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1, §6, quoting plan
72
+ §73/§74). A gateway is nevertheless present, and `docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 gives three
73
+ concrete reasons rather than an architectural preference:
74
+
75
+ 1. **The inference host cannot hold the security boundary.** It is a public ASGI app on third-party
76
+ infrastructure. Rate limiting, size caps, CORS and secret custody belong outside it
77
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 1).
78
+ 2. **A request that can be rejected on shape must never reach inference.** In the original design the
79
+ scarce resource was the ZeroGPU `5 GPU-minutes/day` budget; in the active design it is inference wall
80
+ time on a CPU Codespace. Either way the gateway is where a malformed request dies cheaply
81
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 2).
82
+ 3. **The plan's §74 boundary excludes auth, multi-tenancy and queues.** So the gateway is a *proxy with
83
+ validation*, and must not grow into a platform (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 3, §2.2).
84
+
85
+ The conclusion recorded in the source document is **"two services, not three"** — a static client, a
86
+ gateway, and one inference service (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1). Hugging Face is a presence
87
+ tier, not a runtime tier, in the active design.
88
+
89
+ ---
90
+
91
+ ## 2. Why a gateway exists
92
+
93
+ This section is the load-bearing one. A reader who understands only one part of the deployment should
94
+ understand this: **the gateway is not there to compute anything. It is there to be the boundary.**
95
+
96
+ ### 2.1 The responsibility table (authoritative)
97
+
98
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 is the authoritative statement. Reproduced with the active
99
+ host name substituted (`Railway` → `Render`):
100
+
101
+ | Responsibility | Detail | Why it must be here |
102
+ |---|---|---|
103
+ | **Schema validation** | reject malformed bodies with the §5 error envelope | avoids spending inference on a request that will fail |
104
+ | **Size limits** | per-request body cap *and* per-file cap | the inference host cannot refuse a body it has already received |
105
+ | **Rate limiting** | per-IP count + window | back-pressure against accidental loops; **fairness, not security** — see §2.4 |
106
+ | **CORS** | explicit allowlist of the frontend origin | **never `*`** |
107
+ | **Request IDs** | generate, inject, echo `X-Request-Id` | correlation across two services |
108
+ | **Timeouts** | upstream timeout **shorter** than the inference host's own budget | prevents a hung proxy holding a connection |
109
+ | **Secret custody** | `GITHUB_TOKEN` (and `HF_TOKEN` if used) live here only | the browser never sees them |
110
+ | **Error translation** | inference errors → the documented envelope | the error contract is a gateway product |
111
+ | **Body relaying for upload** | read and forward the raw body for `POST /v1/assets` | the upload path is not JSON-shaped, so JSON-oriented handling does not apply |
112
+
113
+ ### 2.2 The CORS allowlist is explicit, and never a wildcard
114
+
115
+ The orchestrator's CORS list is assembled by `_allowed_origins()` in `deploy/render/main.py`, in a
116
+ documented order:
117
+
118
+ 1. `SATQUERY_ALLOWED_ORIGINS` — the operator's comma-separated list. The authoritative source for any
119
+ additional deployment origin.
120
+ 2. `_PRODUCTION_ORIGINS` — `("https://satquery.pages.dev",)`, **always present**, so a deployment that
121
+ forgets the environment variable still serves the real frontend. `deploy/render/main.py` records the
122
+ reasoning: *"an empty allowlist would otherwise take the live site down, which is a worse failure than
123
+ the one this guards."*
124
+ 3. `_DEV_ORIGINS` — 20 enumerated `host:port` pairs (10 ports × `localhost`/`127.0.0.1`), added unless
125
+ `SATQUERY_ALLOW_DEV_ORIGINS` is set to `0`/`false`/`no`/`""`.
126
+
127
+ The dev-origin list is **enumerated, not a regex and not a suffix match** (`deploy/render/main.py`):
128
+
129
+ ```python
130
+ _DEV_ORIGINS: tuple[str, ...] = tuple(
131
+ f"http://{host}:{port}"
132
+ for host in ("localhost", "127.0.0.1")
133
+ for port in ("3000", "5500", "5173", "8000", "8080")
134
+ )
135
+ ```
136
+
137
+ A wildcard is refused in **two** places, deliberately:
138
+
139
+ * `_allowed_origins()` raises `ValueError` if `"*"` appears in the assembled list, and its docstring
140
+ records why the check exists there as well as in the config validator: *"this function cannot be the
141
+ way a `*` reaches `CORSMiddleware`, which does not run that validator."*
142
+ * `GatewayConfig.__post_init__` (`gateway/policy.py`) refuses a wildcard at construction, so a
143
+ misconfiguration fails at startup rather than on the first request.
144
+
145
+ `allow_credentials=False` is set explicitly in `create_app()` (`deploy/render/main.py`), matching the
146
+ contract's "no auth, no cookies" position (`docs/API_CONTRACT.md` §7; plan §74).
147
+
148
+ The gateway also **strips CORS headers coming back from upstream**, so the CORS answer is the gateway's
149
+ alone. `gateway/app.py::_proxy` asserts this rather than trusting it:
150
+
151
+ ```python
152
+ assert not any(_is_cors_header(k) for k in out_headers) or decision.headers, (
153
+ "a CORS header reached the response without a policy decision; the "
154
+ "upstream's headers are no longer filtered (see F-2)"
155
+ )
156
+ ```
157
+
158
+ ### 2.3 Size limits: two caps, both enforced twice, on purpose
159
+
160
+ Two independent caps exist, and they are different numbers with different jobs
161
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §4):
162
+
163
+ | Cap | Default | Scope | Where read |
164
+ |---|---|---|---|
165
+ | `SATQUERY_MAX_FILE_BYTES` | `4 * 1024 * 1024` = **4,194,304 bytes** | one uploaded file | **both** layers, from one variable |
166
+ | `SATQUERY_MAX_BODY_BYTES` | `8 * 1024 * 1024` | the whole request body | gateway |
167
+
168
+ `gateway/policy.py:221` declares `max_file_bytes: int = 4 * 1024 * 1024`;
169
+ `app/space_app.py::_asset_max_file_bytes()` returns `4 * 1024 * 1024` when the variable is unset. The
170
+ per-file cap is deliberately shared so the two layers cannot disagree about what "too large" means
171
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §4).
172
+
173
+ `SATQUERY_MAX_BODY_BYTES` is **enforced twice** — from the `Content-Length` header *and* while reading
174
+ the bytes — because the header check is *declarative*: it measures what the client claims. The measured
175
+ consequence is in `docs/DEPLOYMENT_ARCHITECTURE.md` §4 (F-6), with the cap at 8 MiB and a 12 MiB body:
176
+
177
+ | Client behaviour | Result | Peak allocation | Bytes read |
178
+ |---|---|---|---|
179
+ | `Content-Length` declared, 12 MiB | `413 oversized_image` | **0.2 MiB** | **0** |
180
+ | `Content-Length` omitted, 12 MiB | `502 model_unavailable` | **13.9 MiB** | **12 MiB** |
181
+
182
+ and allocation tracked body size exactly with no ceiling: `1/8/16/32/64 MiB in → 3.0/8.1/16.0/32.0/64.0 MiB
183
+ allocated`. The remedy was to make the cap unconditional by enforcing it **while reading**, in the single
184
+ shared reader `gateway/assets.py::read_body_bounded`, called by both layers
185
+ (`gateway/app.py::_read_body_bounded` is now a thin adapter over it; `app/space_app.py`'s `/v1/assets`
186
+ handler calls the same function — that is F-9, which found the Space calling `await request.body()` and
187
+ holding 64 MiB in → 128 MiB peak).
188
+
189
+ > **The honest framing, quoted from the source:** *"Operators should not treat the header check as the
190
+ > protection — it protects the gateway's memory against honest clients, not against hostile ones."*
191
+ > (`docs/DEPLOYMENT_ARCHITECTURE.md` §4, F-6 note.)
192
+
193
+ ### 2.4 Rate limiting is FAIRNESS, not security
194
+
195
+ This is a ruling, not an implementation detail. `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2 carries the
196
+ owner ruling of 2026-09-23:
197
+
198
+ > *"✅ RULED 2026-09-23 (owner ruling): the limiter is RETAINED as a fairness / rate-control mechanism
199
+ > only, and it is explicitly NOT a security or abuse-prevention boundary."*
200
+
201
+ The measurement that forced the ruling is reproduced here because it is the whole argument. Limit set to
202
+ **3 requests / 60 s**, **8 requests** sent in-process:
203
+
204
+ | Case | Statuses | Throttled |
205
+ |---|---|---|
206
+ | One client, no `X-Forwarded-For` | `502 502 502 429 429 429 429 429` | **5 / 8** |
207
+ | A fresh spoofed `X-Forwarded-For` per request | `502 502 502 502 502 502 502 502` | **0 / 8** |
208
+
209
+ The mechanism is `gateway/app.py::_client_ip`, which derives the rate-limit key from the **first hop of
210
+ `X-Forwarded-For`** — a client-supplied header. Its own docstring already said the value is
211
+ attacker-controlled and is *"a rate-limit key, not an identity"*; what the measurement added is that the
212
+ limiter **does not hold at all** against a caller willing to vary one header.
213
+
214
+ Consequences a deployment must honour (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2):
215
+
216
+ * **Do not size abuse protection on this limiter.** It is not that control.
217
+ * **A `429` is a fairness signal, not a security signal**, and its **absence is not evidence** that no
218
+ abuse occurred.
219
+ * The gateway remains the request-side boundary for **shape, size and content type** — the things it can
220
+ actually enforce. Rate is not one of them.
221
+
222
+ The limit itself is two variables because the limit *is* the pair (`docs/DEPLOYMENT_ARCHITECTURE.md` §4):
223
+ `SATQUERY_RATE_LIMIT_PER_IP` and `SATQUERY_RATE_LIMIT_WINDOW_S`; `10` and `60.0` mean "ten per minute".
224
+
225
+ > **Why there is no code fix.** Correctly trusting `X-Forwarded-For` requires knowing how many proxy hops
226
+ > the platform inserts — a deployment fact not verifiable from the build host. Hard-coding an assumption
227
+ > would replace a *documented* weakness with an *undocumented* one
228
+ > (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2).
229
+
230
+ ### 2.5 Request IDs
231
+
232
+ The gateway generates a request id, injects it on the upstream leg, and echoes it to the client
233
+ (`gateway/app.py::_proxy`):
234
+
235
+ ```python
236
+ headers = policy.upstream_headers(dict(request.headers), token=token)
237
+ headers["X-Request-Id"] = decision.request_id
238
+ ```
239
+
240
+ Every non-2xx envelope the gateway owns carries the same id, including ones raised by the framework's own
241
+ 404/405 handler, which is registered explicitly (`gateway/app.py`):
242
+
243
+ ```python
244
+ @app.exception_handler(StarletteHTTPException)
245
+ async def _contract_envelope_for_transport_errors(request, exc):
246
+ code = "routing_error" if exc.status_code < 500 else "satquery_error"
247
+ status, body = translate_error(...)
248
+ ```
249
+
250
+ The comment above that handler records the measurement that motivated it: before the fix,
251
+ `GET /v1/whocares → 404 {"detail":"Not Found"}` and `GET /v1/assets → 405 {"detail":"Method Not
252
+ Allowed"}`, while every handler-owned path answered with the contract envelope. A client written to the
253
+ contract parses `error.code` and would get a `KeyError` **exactly when it is trying to explain a failure
254
+ to a user**. `app/space_app.py` carries the same handler for the same reason (F-12/F-12b) — it was
255
+ previously registered on the gateway only.
256
+
257
+ ### 2.6 Timeouts, and the no-retry rule
258
+
259
+ | Timeout | Default | Meaning |
260
+ |---|---|---|
261
+ | `SATQUERY_UPSTREAM_TIMEOUT_S` | `90.0` | gateway → inference request timeout; must sit inside the task budget |
262
+ | `SATQUERY_WAKE_TIMEOUT_S` | `120` | how long the gateway polls for readiness before giving up |
263
+ | `SATQUERY_TUNNEL_TIMEOUT_S` | `150` | how long a tunnel request parks before returning `tunnel_offline` |
264
+
265
+ Defaults are declared in `deploy/render/main.py`:
266
+
267
+ ```python
268
+ def _wake_timeout_s() -> float:
269
+ return float(os.environ.get("SATQUERY_WAKE_TIMEOUT_S", "120"))
270
+
271
+ def _upstream_timeout_s() -> float:
272
+ return float(os.environ.get("SATQUERY_UPSTREAM_TIMEOUT_S", "90"))
273
+ ```
274
+
275
+ **The gateway never retries `POST /v1/analyze`.** `gateway/app.py::_proxy` states it inline:
276
+
277
+ ```python
278
+ except Exception as exc: # network-level failure
279
+ # NO RETRY. A retry on /v1/analyze would spend GPU quota twice
280
+ # (docs/DEPLOYMENT_ARCHITECTURE.md section 2.2).
281
+ ```
282
+
283
+ and `docs/DEPLOYMENT_TOPOLOGY.md` §2 repeats it for the active design: *"Render must not retry
284
+ `POST /api/infer` on its own — a retry would consume inference a second time. The client decides on
285
+ retry."* The client-side consequence is a hard rule in the frontend contract: **never automatically retry
286
+ `POST /v1/analyze`** (`docs/FRONTEND_INTEGRATION.md` §6.1).
287
+
288
+ ### 2.7 Secret custody
289
+
290
+ | Secret | Lives | Never |
291
+ |---|---|---|
292
+ | `GITHUB_TOKEN` | Render environment only | in the browser, in the repo, in a client bundle |
293
+ | `HF_TOKEN` | Render environment only, *if* the HF proxy path is used | as above |
294
+
295
+ The live Render configuration was measured on 2026-09-25 and **has no `SATQUERY_UPSTREAM_URL` and no
296
+ `HF_TOKEN`** (`docs/DEPLOYMENT_TOPOLOGY.md` header note; `release/repo/docs/DEPLOYMENT.md` §3.1). The
297
+ token that *is* present is `GITHUB_TOKEN` — needed only by the GitHub-API wake path, and reported in the
298
+ health payload as a boolean, never a value:
299
+
300
+ ```python
301
+ "has_github_token": bool(os.environ.get("GITHUB_TOKEN")),
302
+ ```
303
+
304
+ `docs/FRONTEND_INTEGRATION.md` §7 states the frontend requirement plainly: **no secrets in the browser**,
305
+ talk only to the gateway, and never call the inference host directly — *"it is not the security boundary
306
+ and its CORS will not welcome you."*
307
+
308
+ > **This document contains no credential, token, key or password, and no path to a credential file.**
309
+ > Every secret is described by *where it lives*, never by its value.
310
+
311
+ ### 2.8 Error translation, and one rule about codes
312
+
313
+ The gateway translates upstream failures into the documented envelope but **passes the `code` through
314
+ unchanged** (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.3):
315
+
316
+ > *"The `code` is passed through **unchanged**. The gateway must not invent codes: the taxonomy in
317
+ > `core/errors.py` is the single source of truth, and a gateway that remapped it would make the
318
+ > frontend's error handling unpredictable."*
319
+
320
+ The envelope shape is fixed (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.3):
321
+
322
+ ```json
323
+ {
324
+ "error": {
325
+ "code": "pair_misaligned",
326
+ "message": "The images are not sufficiently co-registered for spatial analysis.",
327
+ "detail": "RMSE 4.21 px exceeds the 2.0 px budget",
328
+ "recoverable": false,
329
+ "request_id": "req_01H...",
330
+ "run_id": "9f2c1c0e-..."
331
+ }
332
+ }
333
+ ```
334
+
335
+ The orchestrator's own translation table is small and explicit (`deploy/render/main.py`):
336
+
337
+ | Orchestrator error class | `code` | HTTP | `recoverable` |
338
+ |---|---|---|---|
339
+ | `WakeTimeout` | `wake_timeout` | `504` | `true` |
340
+ | `OrchestratorConfigError` | `orchestrator_config_error` | `500` | `false` |
341
+ | `OrchestratorUpstreamError` | `upstream_unreachable` | `502` | `true` |
342
+ | connection/timeout to upstream (`_proxy`) | `upstream_unreachable` | `502` | `true` |
343
+ | other transport error (`_proxy`) | `upstream_error` | `502` | `true` |
344
+ | non-JSON upstream body (`_proxy`) | `schema_validation_error` | `502` | `true` |
345
+ | non-JSON request body (`/api/infer`) | `invalid_request` | `400` | `false` |
346
+
347
+ A **non-JSON upstream body is a defect**, not a pass-through. `gateway/app.py::_proxy` enforces this for
348
+ **every** status, not only 2xx, and the comment records why: the guard originally read
349
+ `upstream.status_code < 400`, so a non-JSON 4xx/5xx — a proxy error page, an HTML 502 from a load
350
+ balancer, a plain-text stack trace — was forwarded verbatim. A sandbox egress proxy returned a 502 whose
351
+ body disclosed `os error 10061`; that is how it was found. `/v1/health` is exempt because a liveness probe
352
+ may legitimately answer non-JSON.
353
+
354
+ **A transport failure's raw exception text is never published** (F-15c, owner ruling 2026-09-23).
355
+ `gateway/app.py::_transport_failure_detail` maps the exception's MRO class names to a path-free
356
+ classification:
357
+
358
+ ```python
359
+ _TRANSPORT_FAILURES: tuple[tuple[str, str], ...] = (
360
+ ("TimeoutException", "the upstream did not respond within the gateway timeout"),
361
+ ("ConnectError", "the upstream could not be reached"),
362
+ ("ProxyError", "the gateway's egress proxy refused the connection"),
363
+ )
364
+ ```
365
+
366
+ The full exception still reaches the operator through `_log.error(..., exc_info=exc)`. It is **moved, not
367
+ deleted**.
368
+
369
+ ### 2.9 What the gateway must NOT do
370
+
371
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 is a closed list:
372
+
373
+ * No persistence. No database, no Redis, no session store.
374
+ * No model inference.
375
+ * No auth system (plan §74).
376
+ * No request queue (plan §73 forbids Redis-cluster/queue infrastructure).
377
+ * No retries on `POST /v1/analyze`.
378
+ * **No second copy of the capability table.** The gateway proxies `/v1/capabilities` and nothing else
379
+ decides that question. The authoritative sources for asset counts are `core.planner.CAPABILITY_ASSETS`
380
+ and `SpecialistSpec.requires_assets`; per `_indices_for`'s docstring, *"duplicating that logic here
381
+ would give two places to disagree."*
382
+ * **No asset storage.** The gateway relays upload bytes; it does not retain them. The store lives with the
383
+ inference host, which is the only component that will read them back
384
+ (`app/space_app.py::get_asset_store` docstring).
385
+
386
+ `deploy/render/main.py`'s module docstring states the same three absences in one line: *"It holds no
387
+ model, no state, no database, and performs **no auth** (per plan §73/§74)."* And it repeats the
388
+ capability-table rule: *"There is deliberately **no second copy** of the capability table here; the
389
+ gateway proxies `/v1/capabilities` and nothing else decides that question."*
390
+
391
+ ### 2.10 The proxied route allowlist
392
+
393
+ The gateway forwards an **allowlist, not a passthrough** (`gateway/app.py`):
394
+
395
+ ```python
396
+ PROXIED_ROUTES: tuple[str, ...] = (
397
+ "/v1/health",
398
+ "/v1/capabilities",
399
+ "/v1/analyze",
400
+ "/v1/assets",
401
+ )
402
+
403
+ COSTLY_ROUTES: tuple[str, ...] = ("/v1/analyze", "/v1/assets")
404
+ ```
405
+
406
+ The two tuples answer different questions and are deliberately separate — *"may this reach the Space at
407
+ all?"* versus *"does it cost a metered resource?"* — because collapsing them would make the rate
408
+ limiter's coverage depend on the proxy allowlist (`gateway/app.py`).
409
+
410
+ `BLOCKED_ROUTES` is **empty**, and the comment says it should stay that way: the tuple exists so a route
411
+ the contract discusses but the server does not implement answers **501 with a reason** instead of a 404
412
+ a frontend developer would debug as a typo.
413
+
414
+ On the orchestrator side the four routes are `/api/health`, `/api/infer`, `/api/capabilities`,
415
+ `/api/assets`, each proxying to the matching `/v1/*` route (`docs/DEPLOYMENT_TOPOLOGY.md` §3.2;
416
+ `deploy/render/main.py`). `/api/health` is the exception: it **never** answers for the inference host.
417
+ Its docstring says so — *"Reports its own configuration; never answers for the Codespace (that is
418
+ `/api/capabilities`)."*
419
+
420
+ ---
421
+
422
+ ## 3. Why the transport is an outbound tunnel
423
+
424
+ ### 3.1 The forwarded-port failure
425
+
426
+ A GitHub Codespace exposes a forwarded port publicly, but **for a private repository that forwarded URL
427
+ returns HTTP 302** — a redirect to a sign-in page, not the service. `docs/DEPLOYMENT_TOPOLOGY.md` records
428
+ this in its measured note:
429
+
430
+ > *"Transport is an **outbound tunnel**, not a polled forwarded port: the Codespace runs
431
+ > `deploy/codespace/tunnel_agent.py`, which dials out to `POST /tunnel/agent` (long-poll) and executes
432
+ > against `http://127.0.0.1:8000` locally."*
433
+
434
+ `release/repo/docs/DEPLOYMENT.md` §7 lists it among the platform traps:
435
+
436
+ > *"A forwarded Codespace port returns `302` for a private repo — which is *why* the tunnel exists."*
437
+
438
+ and `docs/DEPLOYMENT_DECISION.md`'s correction banner records the historical position and its reversal:
439
+
440
+ > *"Codespaces were **not** dropped; the forwarded-port path is dead (HTTP 302 for a private repo) and an
441
+ > outbound tunnel is used instead."*
442
+
443
+ ### 3.2 What the inversion buys
444
+
445
+ `deploy/codespace/launch.sh` states the property in its header comment, and it is worth quoting because it
446
+ is the whole reason the design is robust to repository visibility:
447
+
448
+ > *"The tunnel is why this works with a PRIVATE repository: the agent makes only outbound HTTPS calls, so
449
+ > GitHub's port-forwarding relay, port visibility and the repository's visibility are all irrelevant. The
450
+ > orchestrator never dials into this Codespace."*
451
+
452
+ Consequences, each observable:
453
+
454
+ | Property | Value under the tunnel |
455
+ |---|---|
456
+ | Repository visibility | irrelevant — only outbound HTTPS is used |
457
+ | Port visibility setting | irrelevant |
458
+ | Inbound firewall / NAT | no inbound connection is required at all |
459
+ | Who initiates | the **Codespace**, to `SATQUERY_HUB_URL` |
460
+ | What the hub needs | a long-poll endpoint and a way to match a response to a pending request |
461
+
462
+ ### 3.3 Direction, restated as a diagram
463
+
464
+ ```mermaid
465
+ sequenceDiagram
466
+ autonumber
467
+ participant CF as "Cloudflare Pages"
468
+ participant R as "Render hub"
469
+ participant TA as "Codespace tunnel agent"
470
+ participant API as "FastAPI :8000"
471
+
472
+ Note over TA,R: startup — agent dials OUT
473
+ TA->>R: POST /tunnel/agent (announce, long-poll)
474
+ R-->>TA: (holds the poll open)
475
+
476
+ CF->>R: POST /api/infer
477
+ R->>TA: deliver request on the open poll
478
+ TA->>API: POST http://127.0.0.1:8000/v1/analyze
479
+ API-->>TA: ResultEnvelope
480
+ TA-->>R: response
481
+ R-->>CF: envelope + X-SatQuery-State
482
+ ```
483
+
484
+ > **Honest note on the agent's internals.** `deploy/codespace/launch.sh` invokes
485
+ > `python deploy/codespace/tunnel_agent.py` and greps its log for the string `announced to hub`. That
486
+ > file is **not present in the monorepo working tree** and is **not tracked by git** (see §9.4), so its
487
+ > function names, arguments and payload shapes are
488
+ > `UNKNOWN — not established from the available evidence`. What *is* established is: the agent exists in
489
+ > the production `SatQuery-Inference` repository (`docs/FINAL_DELIVERY_TODO.md` §1.3), it dials
490
+ > `SATQUERY_HUB_URL`, it executes against `http://127.0.0.1:8000`, and it is supervised by
491
+ > `deploy/codespace/launch.sh`.
492
+
493
+ ### 3.4 The observable proof of transport
494
+
495
+ The frontend treats a response header as the evidence that the hub forwarded to the Codespace rather than
496
+ answering locally. `frontend/assets/js/live.js` reads it, and the unit suite pins the read:
497
+
498
+ > *"`x-satquery-transport: tunnel` is the proof that Render forwarded to the Codespace rather than
499
+ > answering locally. It is only readable before the response object is discarded."*
500
+ > (`tests/unit/test_frontend_live_wiring.py`, `test_the_client_reads_the_transport_header_as_evidence`)
501
+
502
+ The measured live value is `x-satquery-transport: tunnel` on `POST /api/infer` (`docs/FINAL_DELIVERY_TODO.md`
503
+ §1.4, §6 E-03; `docs/FINAL_DELIVERY_REPORT.md` §3 P3).
504
+
505
+ ---
506
+
507
+ ## 4. Wake flow
508
+
509
+ ### 4.1 The flow
510
+
511
+ The inference Codespace is CPU-first and **may be stopped when idle**. Before a request can be served the
512
+ hub starts it (if stopped) and polls health until it answers. The frontend shows *"Waking inference
513
+ engine…"* while this happens (`docs/DEPLOYMENT_TOPOLOGY.md` §2).
514
+
515
+ ```mermaid
516
+ sequenceDiagram
517
+ participant CF as "Cloudflare Pages"
518
+ participant R as "Render hub"
519
+ participant C as "GitHub Codespace"
520
+ participant HF as "Hugging Face"
521
+
522
+ CF->>R: GET /api/health (or POST /api/infer)
523
+ R->>C: is the Codespace running?
524
+ alt stopped
525
+ R->>C: start Codespace
526
+ R->>C: poll GET /v1/health
527
+ C-->>R: 200 {status: ok|degraded}
528
+ R-->>CF: "Waking inference engine…"
529
+ end
530
+ CF->>R: POST /api/infer (query + assets)
531
+ R->>C: POST /v1/analyze
532
+ C->>HF: resolve pinned model references
533
+ C-->>R: ResultEnvelope
534
+ R-->>CF: result (envelope + error translation)
535
+ ```
536
+
537
+ Source: `docs/DEPLOYMENT_TOPOLOGY.md` §2 (verbatim structure).
538
+
539
+ ### 4.2 The wake path in code
540
+
541
+ `deploy/render/main.py::ensure_codespace_up()` is the wake implementation. Its contract is precise:
542
+
543
+ ```python
544
+ async def ensure_codespace_up() -> tuple[str, bool]:
545
+ """Ensure the Codespace is running; return ``(base_url, woke)``.
546
+
547
+ Steps:
548
+ 1. ``GET`` the Codespace via the GitHub API.
549
+ 2. If ``state != "available"``, ``POST .../start``.
550
+ 3. Poll ``GET {base}/v1/health`` until 200 or until
551
+ ``SATQUERY_WAKE_TIMEOUT_S`` elapses.
552
+ """
553
+ ```
554
+
555
+ Its polling knobs are module constants:
556
+
557
+ ```python
558
+ _WAKE_POLL_INTERVAL_S = 2.0
559
+ _WAKE_HEALTH_TIMEOUT_S = 10.0
560
+ ```
561
+
562
+ and the failure mapping is explicit: a GitHub auth/transport failure becomes
563
+ `OrchestratorUpstreamError` (`502`, recoverable), a missing Codespace name becomes
564
+ `OrchestratorConfigError` (`500`, not recoverable), and an exhausted deadline raises `WakeTimeout`
565
+ (`504`, recoverable) with the last probe error in the detail.
566
+
567
+ ### 4.3 The response header the client reads
568
+
569
+ `/api/infer` tags the proxied response so the frontend can tell whether the delay was a cold start
570
+ (`deploy/render/main.py`):
571
+
572
+ ```python
573
+ out = await _proxy("POST", f"{base}/v1/analyze", json=body)
574
+ out.headers["X-SatQuery-State"] = "waking" if woke else "ready"
575
+ return out
576
+ ```
577
+
578
+ The unit suite pins both headers on the client side: `assert "x-satquery-transport" in source` and
579
+ `assert "X-SatQuery-State" in source` (`tests/unit/test_frontend_live_wiring.py`).
580
+
581
+ ### 4.4 The wake path is a *fallback* in the tunnel design
582
+
583
+ The measured note in `docs/DEPLOYMENT_TOPOLOGY.md` §2 is explicit that the tunnel design does not depend
584
+ on the GitHub-API wake:
585
+
586
+ > *"The GitHub-API wake path (`POST /user/codespaces/{name}/start`) still exists but the tunnel design
587
+ > relies on the agent reconnecting on Codespace start via the devcontainer `postStartCommand`."*
588
+
589
+ So there are two mechanisms and they are not equivalent:
590
+
591
+ | Mechanism | Trigger | Effect when it works | Effect when it fails |
592
+ |---|---|---|---|
593
+ | Devcontainer `postStartCommand` → `launch.sh` → tunnel agent | every Codespace start | agent reconnects; `agent_connected: true` | `agent_connected: false`; `/api/infer` parks to `SATQUERY_TUNNEL_TIMEOUT_S` |
594
+ | GitHub-API wake (`ensure_codespace_up`) | any `/api/*` request | starts a stopped Codespace, polls `/v1/health` | `wake_timeout` (`504`, recoverable) |
595
+
596
+ ---
597
+
598
+ ## 5. Cold start — documented, not hidden
599
+
600
+ Render's free tier **sleeps when idle**, and the Codespace **may be stopped** (the live GitHub value
601
+ recorded is `idle_timeout_minutes=30`, `docs/FINAL_DELIVERY_TODO.md` §6 E-04). The measured statement is:
602
+
603
+ > *"Render's free tier also sleeps when idle. Cold start is therefore tens of seconds and is **documented,
604
+ > not hidden**."* (`docs/DEPLOYMENT_TOPOLOGY.md` §2)
605
+
606
+ The UI consequence is recorded in `docs/FRONTEND_INTEGRATION.md` §6:
607
+
608
+ | Constraint | Value | UI consequence |
609
+ |---|---|---|
610
+ | Cold start | tens of seconds | *"A determinate-looking progress bar would lie. Use an indeterminate state with a 'this can take up to a minute' hint."* |
611
+
612
+ and the operator consequence in `docs/FINAL_DELIVERY_REPORT.md` §8:
613
+
614
+ > *"**Warm the demo stack** ~10 min before presenting: open the Codespace and confirm `GET /api/health`
615
+ > shows `tunnel.agent_connected:true`. If the Codespace idle-stops, restart it (the tunnel agent
616
+ > reconnects via the devcontainer `postStartCommand`)."*
617
+
618
+ > **No latency characterisation exists.** `docs/FRONTEND_INTEGRATION.md` §9 states it plainly:
619
+ > *"Latency is not characterized. No cold-start or throughput measurement has been taken against a live
620
+ > Space."* The phrase "tens of seconds" is a documented expectation, not a measurement. A precise cold-start
621
+ > distribution is `UNKNOWN — not established from the available evidence`.
622
+
623
+ ---
624
+
625
+ ## 6. `transport_mode: auto`, the fallthrough, and B-07
626
+
627
+ ### 6.1 The live transport configuration
628
+
629
+ The live Render service reports its transport settings in the health payload. Measured
630
+ 2026-09-25:
631
+
632
+ | Setting | Live value |
633
+ |---|---|
634
+ | `transport_mode` | `auto` |
635
+ | `tunnel_timeout_s` | `150.0` |
636
+ | `wake_timeout_s` | `120.0` |
637
+ | `upstream_timeout_s` | `90.0` |
638
+
639
+ Source: `release/repo/docs/DEPLOYMENT.md` §2 (live payload) and `docs/DEPLOYMENT_TOPOLOGY.md` header note.
640
+
641
+ ### 6.2 The fallthrough, exactly
642
+
643
+ `docs/FINAL_DELIVERY_TODO.md` §5 (blocker register, row B-07) records the confirmed root shape:
644
+
645
+ > *"Root shape confirmed 2026-09-25: in `auto` transport mode a tunnel timeout **falls through** to the
646
+ > forward path (`SatQuery-Backend/main.py:546`), which then burns `wake_timeout_s=120` on a 302 → the
647
+ > observed 504."*
648
+
649
+ `DELIVERY_REPORT_2026-09-25.md` §4 gives the mechanism and the arithmetic:
650
+
651
+ > *"in `auto` transport mode a tunnel timeout **falls through** to the forward path (`main.py:546` returns
652
+ > early only when `mode == "tunnel"`); the forward path then burns `wake_timeout_s = 120` on a 302.
653
+ > Measured timing ≈ 249 s ≈ `tunnel_timeout_s=150` + `wake_timeout_s=120`."*
654
+
655
+ So the worst case is:
656
+
657
+ ```
658
+ tunnel park 150 s (SATQUERY_TUNNEL_TIMEOUT_S)
659
+ + wake poll 120 s (SATQUERY_WAKE_TIMEOUT_S)
660
+ -------------------------
661
+ ≈ 249 s → a 504 the client waited four minutes for
662
+ ```
663
+
664
+ ```mermaid
665
+ flowchart TD
666
+ A["POST /api/infer<br/>transport_mode = auto"] --> B{"tunnel agent<br/>connected?"}
667
+ B -- yes --> C["execute via tunnel<br/>x-satquery-transport: tunnel"]
668
+ B -- "no / timeout" --> D["tunnel park expires<br/>SATQUERY_TUNNEL_TIMEOUT_S = 150 s"]
669
+ D --> E{"mode == tunnel?"}
670
+ E -- yes --> F["return tunnel_offline<br/>503 recoverable"]
671
+ E -- "no (auto) → FALLS THROUGH" --> G["forward path:<br/>forwarded port answers 302"]
672
+ G --> H["burns wake_timeout_s = 120 s<br/>polling health"]
673
+ H --> I["wake_timeout<br/>504 recoverable"]
674
+ style I fill:#fde,stroke:#c33
675
+ style D fill:#ffe,stroke:#cc3
676
+ ```
677
+
678
+ ### 6.3 B-07 is OPEN
679
+
680
+ **`B-07` — Transient tunnel-agent gaps — is `OPEN`.** Stated three times in the sources so it cannot be
681
+ mistaken:
682
+
683
+ > *"`B-07` | **Transient tunnel-agent gaps** | OPEN | A request can hang or return 504 (`tunnel_offline`
684
+ > / wake timeout; the forwarded port returns 302). Observed once live. Mitigation: keep the Codespace
685
+ > warm before the demo; the client shows an actionable retry message."* (`docs/FINAL_DELIVERY_REPORT.md`
686
+ > §6)
687
+
688
+ > *"`B-07` | Transient tunnel-agent gaps (agent briefly absent) → a request can hang or return 504
689
+ > (`tunnel_offline` / wake timeout, forward path 302) | … | **OPEN — patch prepared, not deployed.**"*
690
+ > (`docs/FINAL_DELIVERY_TODO.md` §5)
691
+
692
+ > *"B-07 backend patch **prepared, NOT deployed**."* (`docs/FINAL_DELIVERY_TODO.md` sprint-status note)
693
+
694
+ ### 6.4 The prepared patch — prepared, NOT deployed
695
+
696
+ The patch is `fix-b07-forward-unavailable.patch`, in the session workspace at
697
+ `.workbuddy-ai/scratch/deployed-backend/fix-b07-forward-unavailable.patch`
698
+ (`DELIVERY_REPORT_2026-09-25.md` §8). Its content and verification:
699
+
700
+ | Item | Detail |
701
+ |---|---|
702
+ | Base | the **deployed** `SatQuery-Backend/main.py` @ `89d80eaddec5` (769 lines) |
703
+ | Size | 9 hunks plus a 340-line test |
704
+ | Change A | adds `forward_unavailable` (`503`, `recoverable: true`) for a **terminal** 302/401/403 on the forward path, instead of burning the wake timeout |
705
+ | Change B | adds `upstream_timeout` (`504`) for "tunnel healthy but slow" |
706
+ | Change C | fixes `/api/health` `codespace_name` trailing `\n` via `.strip()` |
707
+ | Independent verification | `git apply --check` clean, `git apply` clean, `py_compile` OK |
708
+ | Presence check | `forward_unavailable` @ `main.py:326`, `upstream_timeout` @ `:601`, `codespace_name` `.strip()` @ `:686` |
709
+ | Deployment status | **NOT deployed** |
710
+
711
+ Sources: `DELIVERY_REPORT_2026-09-25.md` §4; `docs/FINAL_DELIVERY_TODO.md` §6 E-12.
712
+
713
+ > **A retracted claim, recorded because the honesty matters.** The report records that an earlier claim
714
+ > that the patch *"would not have prevented"* the observed 504 *"was wrong and was retracted"*. The
715
+ > corrected position: *"Change A is genuinely **on the failing path** — it converts a 504-after-249 s into
716
+ > a 503-early with an actionable code."* (`DELIVERY_REPORT_2026-09-25.md` §4.)
717
+
718
+ **Why it is not deployed:** *"the patch is not needed for the demo and touches the live backend. The
719
+ residual is better mitigated operationally (keep the Codespace warm, raise the idle timeout)."*
720
+ (`DELIVERY_REPORT_2026-09-25.md` §4.)
721
+
722
+ ### 6.5 Operational trap recorded with the patch
723
+
724
+ > *"the local `C:/Users/anish/SatQuery-Backend` (680 lines) is **STALE**. Always fetch the deployed
725
+ > `main.py` before touching backend code."* (`DELIVERY_REPORT_2026-09-25.md` §4)
726
+
727
+ This is the same class of trap as §9.4 below: **the working copy is not the deployed source.**
728
+
729
+ ---
730
+
731
+ ## 7. The full live health payload
732
+
733
+ ### 7.1 The measured payload
734
+
735
+ Probed live on 2026-09-25 against `https://satquery-backend-m4yv.onrender.com/api/health`
736
+ (`release/repo/docs/DEPLOYMENT.md` §2):
737
+
738
+ ```json
739
+ {"status":"ok","service":"satquery-orchestrator",
740
+ "tunnel":{"agent_connected":true,"agent_id":"codespaces-fd1038","pending":0,"completed":97},
741
+ "config":{"codespace_name":"potential-space-trout-r4ppw969w45j2pvvw\n","codespace_port":8000,
742
+ "transport_mode":"auto","tunnel_timeout_s":150.0,"wake_timeout_s":120.0,
743
+ "upstream_timeout_s":90.0,"device":"cpu","has_github_token":true}}
744
+ ```
745
+
746
+ Exact command used elsewhere in the project's evidence register:
747
+ `curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health`
748
+ (`docs/FINAL_DELIVERY_REPORT.md` §4).
749
+
750
+ ### 7.2 Field-by-field
751
+
752
+ | Field | Type | Meaning | Live value |
753
+ |---|---|---|---|
754
+ | `status` | string | the hub's own liveness | `"ok"` |
755
+ | `service` | string | the service identity | `"satquery-orchestrator"` |
756
+ | `tunnel.agent_connected` | bool | is a tunnel agent currently polling? | `true` |
757
+ | `tunnel.agent_id` | string | which agent identity holds the poll | `"codespaces-fd1038"` |
758
+ | `tunnel.pending` | int | requests delivered but not yet answered | `0` |
759
+ | `tunnel.completed` | int | requests completed since the agent connected | `97` |
760
+ | `config.codespace_name` | string | the target Codespace | `"…pvvw\n"` — **carries a trailing `\n`** |
761
+ | `config.codespace_port` | int | the inference port | `8000` |
762
+ | `config.transport_mode` | string | transport selection | `"auto"` |
763
+ | `config.tunnel_timeout_s` | float | tunnel park budget | `150.0` |
764
+ | `config.wake_timeout_s` | float | wake poll budget | `120.0` |
765
+ | `config.upstream_timeout_s` | float | proxy request timeout | `90.0` |
766
+ | `config.device` | string | declared device | `"cpu"` |
767
+ | `config.has_github_token` | bool | is a GitHub token configured? | `true` |
768
+
769
+ ### 7.3 `completed` was observed at three different values — do not treat any as a constant
770
+
771
+ The tunnel counter is a **monotonic runtime counter**, not a fixed fact. Three measured readings exist,
772
+ each with its own provenance:
773
+
774
+ | Reading | Where recorded |
775
+ |---|---|
776
+ | `completed: 97` | `release/repo/docs/DEPLOYMENT.md` §2 (the live health probe) |
777
+ | `completed: 314` | `docs/FINAL_DELIVERY_TODO.md` §1.4 and §6 E-02; `docs/DEPLOYMENT_TOPOLOGY.md` §2 |
778
+ | `completed: 338` | `docs/FINAL_DELIVERY_REPORT.md` §3 P2 |
779
+
780
+ They are consistent with each other — the counter grows — and the honest statement is
781
+ **"`completed` was measured at 97, 314 and 338 at three different times on 2026-09-25."** Quoting any one
782
+ of them as *the* value would be wrong.
783
+
784
+ ### 7.4 `codespace_name` carries a trailing newline — B-02, OPEN (cosmetic)
785
+
786
+ `config.codespace_name` reports `…pvvw\n`. This is **B-02**, and its status is `OPEN` **but cosmetic**:
787
+
788
+ > *"`P2-T03` `/api/health` `codespace_name` trailing `\n` | DEFERRED (cosmetic) | Wake path is safe
789
+ > (`_codespace_name()` strips, `main.py:123,357`); only the health payload reports the raw value."*
790
+ > (`docs/FINAL_DELIVERY_REPORT.md` §6)
791
+
792
+ > *"`B-02` | `/api/health` reports `codespace_name` with a trailing `\n` | **Cosmetic** — reporting only;
793
+ > the wake path strips via `_codespace_name()` (`main.py:123,357`) | P2-T03 | none needed |
794
+ > DOWNGRADED"* (`docs/FINAL_DELIVERY_TODO.md` §5)
795
+
796
+ The fix is known and one line — *"change line 619 to `_codespace_name()`, then Render redeploys"*
797
+ (`docs/FINAL_DELIVERY_TODO.md` §4, P2-T03) — and the row's own reasoning for deferring is that *"a
798
+ live-backend redeploy before the demo is not worth the risk."*
799
+
800
+ > **Do not upgrade this.** `B-02` is `OPEN`. It is not `RESOLVED`, and it is not `CLOSED`.
801
+
802
+ ### 7.5 The orchestrator's *own* `/api/health` shape in the repository
803
+
804
+ `deploy/render/main.py` — the monorepo copy, which is **not** the deployed source (§9.4) — declares a
805
+ different, simpler health payload. Reproduced because it documents the *contract* of the route even where
806
+ the deployed implementation has grown:
807
+
808
+ ```python
809
+ @app.get("/api/health")
810
+ async def health() -> dict[str, Any]:
811
+ """Orchestrator liveness. Reports its own configuration; never answers
812
+ for the Codespace (that is /api/capabilities)."""
813
+ return {
814
+ "status": "ok",
815
+ "service": "satquery-orchestrator",
816
+ "config": {
817
+ "codespace_name": os.environ.get("CODESPACE_NAME", ""),
818
+ "codespace_port": _codespace_port(),
819
+ "has_github_token": bool(os.environ.get("GITHUB_TOKEN")),
820
+ "allowed_origins": _allowed_origins(),
821
+ "production_origins": list(_PRODUCTION_ORIGINS),
822
+ "dev_origins_enabled": _dev_origins_enabled(),
823
+ "wake_timeout_s": _wake_timeout_s(),
824
+ "upstream_timeout_s": _upstream_timeout_s(),
825
+ "device": os.environ.get("SATQUERY_DEVICE", ""),
826
+ },
827
+ }
828
+ ```
829
+
830
+ Note the design decision visible here: `allowed_origins` reports the **effective** list, so an operator can
831
+ confirm **from outside** what the service will actually accept — not just what they set. The comment says
832
+ the dev entries being visible *"is how a production deployment proves it turned them off."*
833
+
834
+ > **The discrepancy is real and is stated rather than smoothed over.** The deployed payload carries a
835
+ > `tunnel` block and `config.transport_mode` / `config.tunnel_timeout_s`, which the monorepo copy does
836
+ > not. The monorepo copy is a **532-line** file with no tunnel code at all; the deployed
837
+ > `SatQuery-Backend/main.py` is **768–769 lines** with it (`docs/FINAL_DELIVERY_TODO.md` §1.1;
838
+ > `DELIVERY_REPORT_2026-09-25.md` §4).
839
+
840
+ ---
841
+
842
+ ## 8. Environment variables
843
+
844
+ ### 8.1 Render (orchestrator) — measured live values
845
+
846
+ | Variable | Live value | Purpose |
847
+ |---|---|---|
848
+ | `CODESPACE_NAME` | `potential-space-trout-r4ppw969w45j2pvvw` | which Codespace to target |
849
+ | `CODESPACE_PORT` | `8000` | the inference port on that Codespace |
850
+ | `SATQUERY_ALLOWED_ORIGINS` | `https://satquery.pages.dev` | CORS allowlist (the Pages origin) |
851
+ | `SATQUERY_DEVICE` | `cpu` | declared device |
852
+ | `SATQUERY_TRANSPORT` | `auto` | transport selection |
853
+ | `SATQUERY_TUNNEL_TIMEOUT_S` | `150` | tunnel park budget |
854
+ | `SATQUERY_WAKE_TIMEOUT_S` | `120` | wake poll budget |
855
+ | `SATQUERY_UPSTREAM_TIMEOUT_S` | `90` | gateway → upstream request timeout |
856
+ | `GITHUB_TOKEN` | present | GitHub API wake path; never sent to the browser |
857
+
858
+ Source: `docs/DEPLOYMENT_TOPOLOGY.md` header note (measured against `GET /api/health`);
859
+ `release/repo/docs/DEPLOYMENT.md` §3.1.
860
+
861
+ > **Two absences are as important as the presences.** There is **no `SATQUERY_UPSTREAM_URL`** and **no
862
+ > `HF_TOKEN`** in the live configuration (`docs/DEPLOYMENT_TOPOLOGY.md`; `release/repo/docs/DEPLOYMENT.md`
863
+ > §3.1). `SATQUERY_UPSTREAM_URL` is absent because the transport is the outbound tunnel, not a forwarded
864
+ > port; `HF_TOKEN` is absent because the HF proxy path is not used live.
865
+
866
+ ### 8.2 Render — the blueprint's declared variables
867
+
868
+ `render.yaml` (the blueprint) declares the same vocabulary as a service definition:
869
+
870
+ ```yaml
871
+ services:
872
+ - type: web
873
+ name: satquery-orchestrator
874
+ runtime: python
875
+ plan: free
876
+ buildCommand: pip install -r deploy/render/requirements.txt
877
+ startCommand: uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT
878
+ healthCheckPath: /api/health
879
+ envVars:
880
+ - key: PORT
881
+ sync: false
882
+ - key: SATQUERY_ALLOWED_ORIGINS
883
+ sync: false
884
+ - key: GITHUB_TOKEN
885
+ sync: false
886
+ - key: CODESPACE_NAME
887
+ sync: false
888
+ - key: CODESPACE_PORT
889
+ value: "8000"
890
+ - key: SATQUERY_DEVICE
891
+ value: "cpu"
892
+ - key: SATQUERY_WAKE_TIMEOUT_S
893
+ value: "120"
894
+ - key: SATQUERY_UPSTREAM_TIMEOUT_S
895
+ value: "90"
896
+ ```
897
+
898
+ Three things this file establishes that are easy to miss:
899
+
900
+ 1. `plan: free` — the free tier, which is *why* Render sleeps when idle (§5).
901
+ 2. `healthCheckPath: /api/health` — the platform's own liveness probe points at the orchestrator's
902
+ self-report route, which never touches the inference host.
903
+ 3. `sync: false` on `SATQUERY_ALLOWED_ORIGINS`, `GITHUB_TOKEN`, `CODESPACE_NAME` and `PORT` means those are
904
+ **operator-supplied**, not blueprint-committed. No secret value appears in the repository.
905
+
906
+ > The blueprint does **not** declare `SATQUERY_TRANSPORT` or `SATQUERY_TUNNEL_TIMEOUT_S`, which the live
907
+ > service reports. The blueprint and the live service have diverged. Whether the live service sets them
908
+ > through the dashboard or through a newer blueprint is `UNKNOWN — not established from the available
909
+ > evidence`; what is established is the live value set in §8.1.
910
+
911
+ ### 8.3 Codespace (inference) — declared and effective
912
+
913
+ | Variable | Where set | Purpose |
914
+ |---|---|---|
915
+ | `PORT` | `containerEnv` = `"8000"`, re-exported by `launch.sh` | platform-assigned; **must be read** (historical blocker #2) |
916
+ | `SATQUERY_DEVICE` | `containerEnv` = `"cpu"`, re-exported by `launch.sh` | `cpu` \| `cuda` \| `mps` \| `null`; read **without importing torch** |
917
+ | `SATQUERY_ASSET_ENABLED` | `containerEnv` = `"1"`, re-exported by `launch.sh` | enables `POST /v1/assets`; **both** this and the dir are required |
918
+ | `SATQUERY_ASSET_DIR` | `containerEnv` = `"/tmp/satquery-assets"`, re-exported by `launch.sh` | where uploaded bytes are written |
919
+ | `SATQUERY_MAX_FILE_BYTES` | not set live (default applies) | per-file cap, shared with Render |
920
+ | `SATQUERY_ASSET_MAX_FILES` | not set live (default applies) | optional handle capacity, default `32` |
921
+ | `SATQUERY_ASSET_TTL_S` | not set live (default applies) | optional handle lifetime, default `900.0` |
922
+ | `SATQUERY_HUB_URL` | defaulted by `launch.sh` | the hub the agent dials |
923
+ | `PYTHONPATH` | set by `launch.sh` | repo root, so `import app` resolves |
924
+
925
+ Sources: `.devcontainer/devcontainer.json`; `deploy/codespace/launch.sh`; `app/space_app.py`.
926
+
927
+ `.devcontainer/devcontainer.json` in full:
928
+
929
+ ```json
930
+ {
931
+ "name": "SatQuery AI — Codespace Inference",
932
+ "image": "mcr.microsoft.com/devcontainers/python:3.12",
933
+ "forwardPorts": [8000],
934
+ "portsAttributes": {
935
+ "8000": { "label": "SatQuery inference", "visibility": "public" }
936
+ },
937
+ "containerEnv": {
938
+ "SATQUERY_DEVICE": "cpu",
939
+ "PORT": "8000",
940
+ "SATQUERY_ASSET_ENABLED": "1",
941
+ "SATQUERY_ASSET_DIR": "/tmp/satquery-assets"
942
+ },
943
+ "postCreateCommand": "bash deploy/codespace/post_create.sh",
944
+ "postStartCommand": "bash deploy/codespace/launch.sh",
945
+ "customizations": { "vscode": { "extensions": ["ms-python.python"] } }
946
+ }
947
+ ```
948
+
949
+ > **A trap worth recording, from `launch.sh`'s own comment:** *"`containerEnv` is only applied when the
950
+ > container is CREATED, so setting it there alone would leave an already-running Codespace unconfigured
951
+ > until a rebuild. This script runs on every start and is therefore the effective source of truth."* The
952
+ > variables are therefore set **twice** — in `containerEnv` and in `launch.sh` — and `launch.sh` is the
953
+ > one that governs a running container.
954
+
955
+ ### 8.4 The historical vocabulary — still the contract
956
+
957
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §4 remains authoritative for the env-var *vocabulary*; only host names
958
+ moved. Its full table, reproduced, with the active host substituted:
959
+
960
+ | Variable | Where it lives (historical → active) | Purpose |
961
+ |---|---|---|
962
+ | `HF_TOKEN` | Railway only → **Render only, if used** | upstream credential; never sent to the browser |
963
+ | `SATQUERY_SPACE_URL` | Railway → **`SATQUERY_UPSTREAM_URL`** | upstream URL |
964
+ | `SATQUERY_ALLOWED_ORIGINS` | Railway → **Render** | CORS allowlist |
965
+ | `PORT` | Railway → **Render** | supplied by the platform |
966
+ | `SATQUERY_DEVICE` | Space → **Codespace** | `cpu` \| `cuda` \| `mps` \| `null`; read **without importing torch** |
967
+ | `SATQUERY_ASSET_ENABLED` | Space → **Codespace** | enables `POST /v1/assets`; fails closed |
968
+ | `SATQUERY_ASSET_DIR` | Space → **Codespace** | where uploaded bytes are written |
969
+ | `SATQUERY_MAX_FILE_BYTES` | **both** | per-file size cap, read by both layers from one variable |
970
+ | `SATQUERY_MAX_BODY_BYTES` | Railway → **Render** | whole-request body cap, above the per-file cap |
971
+ | `SATQUERY_UPSTREAM_TIMEOUT_S` | Railway → **Render** | gateway → upstream timeout; default `90.0` |
972
+ | `SATQUERY_RATE_LIMIT_PER_IP` / `_WINDOW_S` | Railway → **Render** | per-IP count + window |
973
+ | `SATQUERY_ASSET_MAX_FILES` / `SATQUERY_ASSET_TTL_S` | Space → **Codespace** | **optional** handle capacity / lifetime |
974
+
975
+ Three notes from that section are worth carrying forward because they explain *why* the vocabulary has
976
+ this shape:
977
+
978
+ 1. **`SATQUERY_MAX_FILE_BYTES` is applied while reading at both layers, not after** (F-9, F-6). Both layers
979
+ call the single reader `gateway/assets.py::read_body_bounded`, so the two enforcement points cannot
980
+ drift.
981
+ 2. **Both layers refuse an unparsable or non-positive value and name the variable** (F-7). Reading one
982
+ variable is not the same as agreeing on its value: the two parsers previously diverged in **opposite
983
+ directions** — `'abc'` raised at the gateway but silently defaulted to 4 MiB on the inference host;
984
+ `'0'` was accepted at the gateway but rejected on the inference host. A malformed cap now fails startup
985
+ at both layers rather than running on a limit nobody chose.
986
+ 3. **None of the asset variables is a config key**, and that is deliberate: adding a key to
987
+ `configs/base.yaml` moves `Config.hash` off `78f1e3700da15aa1` and invalidates the frozen Phase-9
988
+ benchmark. Asset storage is deployment state, so it is read from the environment.
989
+
990
+ The F-8 note on `SATQUERY_DEVICE` is also load-bearing and is reproduced in §8.5.
991
+
992
+ ### 8.5 `SATQUERY_DEVICE`: four read sites, and the case bug
993
+
994
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §4 records that the served value is validated against the contract's
995
+ closed set. Measured before the fix, `SATQUERY_DEVICE` had **four read sites** and only three normalised:
996
+
997
+ | Site | Behaviour before the fix |
998
+ |---|---|
999
+ | `core/config.py:88` (`Config.device_preference`) | raw — no strip, no lower |
1000
+ | `app/deployment.py:570` (`gpu_available`) | `.strip().lower()` |
1001
+ | `app/deployment.py:946` (the served device resolver) | `.strip()`, **no lower** — the odd one |
1002
+ | `app/deployment.py:970` (`_cuda_detected`) | `.strip().lower()` |
1003
+
1004
+ Two defects followed, both measured:
1005
+
1006
+ * **Case changed the answer.** `'cuda'` → `'cpu'` but `'CUDA'` → `'CUDA'`, so one payload could announce
1007
+ `gpu_available: true` alongside `device: "CUDA"` — a GPU is claimed and the device name is not a device.
1008
+ * **An unparsable value was echoed.** `'garbage'` → `device: "garbage"`, against a field the contract
1009
+ publishes as a closed set.
1010
+
1011
+ The served resolver now normalises and **validates**, returning `None` for anything outside
1012
+ `{"cpu", "cuda", "mps"}`. `None` is chosen over raising or over a silent `"cpu"`, because it is already a
1013
+ legal value for the field, it is honest, and defaulting to `"cpu"` *"would mean a typo silently changes
1014
+ which device the process is believed to use, which is the `_asset_max_file_bytes` mistake from F-7 in a
1015
+ different variable."* The guard is kept as a literal, not derived from the implementation, so it encodes
1016
+ the **contract's** set and cannot drift with the code:
1017
+
1018
+ ```python
1019
+ _LEGAL_DEVICES: frozenset[str] = frozenset({"cpu", "cuda", "mps"})
1020
+ ```
1021
+
1022
+ > `Config.device_preference` still returns the raw override, deliberately: it is a general-purpose property
1023
+ > whose other callers may legitimately want the operator's literal text, and narrowing it would be a wider
1024
+ > change than the defect warrants. **The served path is the one the contract constrains, so it is the one
1025
+ > that validates.**
1026
+
1027
+ ---
1028
+
1029
+ ## 9. The tunnel agent and the Codespace launcher
1030
+
1031
+ ### 9.1 `deploy/codespace/serve.py` — the entrypoint
1032
+
1033
+ The file is 25 lines and its whole job is to bind `build_space_app()` to `$PORT`:
1034
+
1035
+ ```python
1036
+ import os
1037
+
1038
+ from app.space_app import build_space_app
1039
+ import uvicorn
1040
+
1041
+ app = build_space_app()
1042
+
1043
+ if __name__ == "__main__":
1044
+ port = int(os.environ.get("PORT", "8000"))
1045
+ uvicorn.run(app, host="0.0.0.0", port=port)
1046
+ ```
1047
+
1048
+ Its docstring records the properties that make it import-safe on a CPU host with no GPU and no weights:
1049
+
1050
+ > *"`build_space_app()` is cheap to import: FastAPI is imported inside it and no model is loaded at module
1051
+ > scope, so this file stays import-safe on a CPU host with no GPU and no weights present."*
1052
+
1053
+ and it names the device resolution path: *"The serving controller (via
1054
+ `app.serving.build_serving_controller`) resolves `device` from the `SATQUERY_DEVICE` env var; set it to
1055
+ `cpu` for the CPU-first adaptation."*
1056
+
1057
+ ### 9.2 `deploy/codespace/launch.sh` — what runs on every Codespace start
1058
+
1059
+ The script is the devcontainer's `postStartCommand` target. It runs three stages plus a preflight.
1060
+
1061
+ **Stage 0 — preflight, refusing to start half-configured.** The header comment states why this exists:
1062
+
1063
+ > *"A silently-broken environment is the single worst failure mode here: the server dies, nothing listens
1064
+ > on the port, and the only external symptom is a bare 401/302 from GitHub's relay — which looks like a
1065
+ > visibility problem."*
1066
+
1067
+ It therefore checks the Python dependencies **including `httpx` explicitly**, and the comment records the
1068
+ incident:
1069
+
1070
+ > *"NOTE: httpx is checked explicitly. `tunnel_agent.py` imports it directly, and it was previously absent
1071
+ > from `requirements.txt` — so the agent died instantly and the supervised restart loop hid the error in a
1072
+ > log file."*
1073
+
1074
+ ```bash
1075
+ if ! python -c "import yaml, pydantic, fastapi, uvicorn, httpx" 2>/dev/null; then
1076
+ echo "ERROR: Python deps are missing (need yaml, pydantic, fastapi, uvicorn, httpx)." >&2
1077
+ ...
1078
+ exit 1
1079
+ fi
1080
+
1081
+ if ! python -c "import app.space_app" 2>/dev/null; then
1082
+ echo "ERROR: cannot import the 'app' package even with PYTHONPATH=$REPO_ROOT" >&2
1083
+ ...
1084
+ exit 1
1085
+ fi
1086
+ ```
1087
+
1088
+ **Stage 1 — the inference server, with a staleness guard.** The script records a **stamp** of the git
1089
+ revision and the asset-upload environment, because `_port_open` alone cannot tell you what the running
1090
+ process was started from:
1091
+
1092
+ ```bash
1093
+ _current_stamp() {
1094
+ printf 'rev=%s asset_enabled=%s asset_dir=%s\n' \
1095
+ "$(git rev-parse HEAD 2>/dev/null || echo nogit)" \
1096
+ "${SATQUERY_ASSET_ENABLED:-}" \
1097
+ "${SATQUERY_ASSET_DIR:-}"
1098
+ }
1099
+ ```
1100
+
1101
+ and the comment explains the failure a stale process causes:
1102
+
1103
+ > *"A stale serve process is worse than no process: it answers `/v1/health` and `/v1/capabilities` from OLD
1104
+ > code, so the deployment looks alive while reporting the previous revision's capabilities."*
1105
+
1106
+ The restart uses the real invocation, not the file path — `pkill -f "python deploy/codespace/serve.py"`,
1107
+ because *"`pgrep -f serve.py` would also match an editor or this script's own argv."* If SIGTERM is not
1108
+ enough it escalates to `pkill -9`, and the process is launched detached:
1109
+
1110
+ ```bash
1111
+ setsid nohup python deploy/codespace/serve.py > "$SERVE_LOG" 2>&1 < /dev/null &
1112
+ ```
1113
+
1114
+ **Stage 2 — the outbound tunnel agent, supervised.** The header comment records the production incident
1115
+ that shaped the launch:
1116
+
1117
+ > *"`setsid` alone is NOT enough in Codespaces. The lifecycle shell that runs `postStartCommand` can still
1118
+ > reap the process group, which showed up in production as 'the agent announced once, then vanished' — the
1119
+ > hub then reported `agent_connected=false` and `/api/infer` fell back to the dead forwarded-port path
1120
+ > (401 -> wake_timeout)."*
1121
+
1122
+ The remedy is `setsid + nohup + </dev/null` **plus a supervising wrapper** that relaunches the agent if it
1123
+ ever exits:
1124
+
1125
+ ```bash
1126
+ setsid nohup bash -c '
1127
+ while true; do
1128
+ echo "[supervisor $(date +%H:%M:%S)] starting tunnel agent" >> "'"$TUNNEL_LOG"'"
1129
+ python deploy/codespace/tunnel_agent.py >> "'"$TUNNEL_LOG"'" 2>&1
1130
+ rc=$?
1131
+ echo "[supervisor $(date +%H:%M:%S)] tunnel agent exited rc=$rc �� restarting in 5s" >> "'"$TUNNEL_LOG"'"
1132
+ sleep 5
1133
+ done
1134
+ ' > /dev/null 2>&1 < /dev/null &
1135
+ ```
1136
+
1137
+ The guard is on the **process, not a port** — *"the agent listens on nothing"* — and the supervisor itself
1138
+ is what gets detached, so *"the agent is effectively immortal for the life of the Codespace."*
1139
+
1140
+ **Stage 3 — verify the agent actually connected.** This stage exists because backgrounding with all output
1141
+ discarded makes a crashing agent invisible:
1142
+
1143
+ > *"Backgrounding with all output discarded means a crashing agent is completely invisible — that is
1144
+ > exactly how a missing `httpx` hid itself. So we wait, then check: the process is alive, and the log shows
1145
+ > a successful announce."*
1146
+
1147
+ ```bash
1148
+ sleep 4
1149
+
1150
+ if ! pgrep -f "deploy/codespace/tunnel_agent.py" > /dev/null 2>&1; then
1151
+ echo "WARNING: the tunnel agent is not running. Last log lines:" >&2
1152
+ tail -n 20 "$TUNNEL_LOG" 2>/dev/null >&2 || echo " (no log at $TUNNEL_LOG)" >&2
1153
+ ...
1154
+ else
1155
+ echo "tunnel agent process is up (pid $(pgrep -f 'deploy/codespace/tunnel_agent.py' | head -1))"
1156
+ if grep -q "announced to hub" "$TUNNEL_LOG" 2>/dev/null; then
1157
+ echo "tunnel agent announced to the hub successfully"
1158
+ ...
1159
+ ```
1160
+
1161
+ The hub URL is a defaulted variable, so a renamed Render service can be overridden in the Codespace:
1162
+
1163
+ ```bash
1164
+ export SATQUERY_HUB_URL="${SATQUERY_HUB_URL:-https://satquery-backend-m4yv.onrender.com}"
1165
+ ```
1166
+
1167
+ ### 9.3 The asset-upload environment, and why `/tmp` is correct
1168
+
1169
+ `launch.sh` sets the asset variables on every start, and its comment argues the choice rather than
1170
+ asserting it:
1171
+
1172
+ > *"`/tmp` is correct here and not a compromise: the Codespace filesystem is ephemeral, handles are TTL'd
1173
+ > (900s), and `cache_max_models: 1` means an uploaded asset is consumed within one analysis, so nothing
1174
+ > needs to outlive the process. The store creates the directory if absent."*
1175
+
1176
+ ```bash
1177
+ export SATQUERY_ASSET_ENABLED="${SATQUERY_ASSET_ENABLED:-1}"
1178
+ export SATQUERY_ASSET_DIR="${SATQUERY_ASSET_DIR:-/tmp/satquery-assets}"
1179
+ ```
1180
+
1181
+ The comment also cross-references the exact fallback path in code — *"the directory is intentionally the
1182
+ same path the code falls back to (`space_app.py:310`)"* — which is
1183
+ `Path(tempfile.gettempdir()) / "satquery-assets"` in `app/space_app.py::_asset_root()`. That is a
1184
+ deliberate alignment: *"a deployment that set only the flag — or neither — cannot silently start writing
1185
+ to a barely-chosen location."*
1186
+
1187
+ ### 9.4 The tunnel agent file itself
1188
+
1189
+ | Question | Answer | Status |
1190
+ |---|---|---|
1191
+ | Is `deploy/codespace/tunnel_agent.py` in the monorepo working tree? | **No** — `Glob **/tunnel_agent*` finds nothing | `MEASURED` |
1192
+ | Is it tracked by git? | **No** — `git ls-files deploy/` is empty; the whole `deploy/` tree is untracked | `MEASURED` |
1193
+ | Where does it exist? | `Anish-lab-blip/SatQuery-Inference` (private) — *"Codespace FastAPI + `deploy/codespace/tunnel_agent.py`"* (`docs/FINAL_DELIVERY_TODO.md` §1.3) | `VERIFIED` |
1194
+ | What are its function names, arguments, payload shapes? | `UNKNOWN — not established from the available evidence` | `OPEN` |
1195
+ | What *is* established about it? | it imports `httpx`; it dials `SATQUERY_HUB_URL`; it executes against `http://127.0.0.1:8000`; it logs `announced to hub`; it is supervised by `launch.sh` | `VERIFIED` (from `launch.sh` comments and greps) |
1196
+
1197
+ > **This is the single largest evidence gap in this chapter**, and it is recorded rather than filled in
1198
+ > with a plausible guess. A reader who needs the agent's protocol should read
1199
+ > `SatQuery-Inference/deploy/codespace/tunnel_agent.py`.
1200
+
1201
+ ### 9.5 The "stale working copy" trap — B-03
1202
+
1203
+ **`B-03` is `KNOWN`**, and it is the reason §9.4 has a gap at all:
1204
+
1205
+ > *"`B-03` | Local `deploy/` stale + untracked | Edits there do not deploy | all deploy tasks | edit the 3
1206
+ > real repos instead | KNOWN"* (`docs/FINAL_DELIVERY_TODO.md` §5)
1207
+
1208
+ > *"Local `deploy/render/main.py` (532 lines, no tunnel) is superseded by `SatQuery-Backend/main.py` (768
1209
+ > lines, tunnel)."* (`docs/FINAL_DELIVERY_TODO.md` §1.1)
1210
+
1211
+ > *"**Critical:** the deployed backend is **not** this working copy."* (`docs/FINAL_DELIVERY_TODO.md` §1.1)
1212
+
1213
+ > *"Local `deploy/` | stale/untracked | Edit the 3 real repos, not this copy."*
1214
+ > (`docs/FINAL_DELIVERY_REPORT.md` §6)
1215
+
1216
+ ### 9.6 Repositories of record
1217
+
1218
+ `docs/FINAL_DELIVERY_TODO.md` §1.3:
1219
+
1220
+ | Repo | Role | Deployed from |
1221
+ |---|---|---|
1222
+ | `Anish-lab-blip/SatQuery-Frontend` (private) | Cloudflare Pages (static) | root = local `frontend/` contents |
1223
+ | `Anish-lab-blip/SatQuery-Backend` (private) | Render hub + `tunnel.py` + `codespaces.py` | Render `satquery-orchestrator` |
1224
+ | `Anish-lab-blip/SatQuery-Inference` (private) | Codespace FastAPI + `deploy/codespace/tunnel_agent.py` | Codespace |
1225
+ | `Anish-lab-blip/SatQuery-AI` (**public**) | umbrella / monorepo mirror | — |
1226
+
1227
+ ### 9.7 Deployed revisions
1228
+
1229
+ | Component | Repository | Branch | Revision | Host |
1230
+ |---|---|---|---|---|
1231
+ | Frontend | `SatQuery-Frontend` | `main` | **`2d7ae53b482d`** | Cloudflare Pages → `satquery.pages.dev` |
1232
+ | Backend / orchestrator | `SatQuery-Backend` | `main` | **`89d80eaddec5`** | Render → `satquery-backend-m4yv.onrender.com` |
1233
+ | Inference | `SatQuery-Inference` | `main` | **`5a0936ace491`** | Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port 8000 |
1234
+ | Public umbrella | `SatQuery-AI` | `main` | `3dcabd32da41` | the release home |
1235
+ | Monorepo working copy | `C:/Users/anish/satquery-ai` | `master` | `9d57aed` | local only, **no remote**, 334 dirty entries |
1236
+
1237
+ Source: `release/repo/docs/DEPLOYMENT.md` §1. This is the correct place to look up a deployed revision;
1238
+ **the monorepo HEAD is not the deployed revision.**
1239
+
1240
+ ---
1241
+
1242
+ ## 10. Deployment mechanics
1243
+
1244
+ ### 10.1 Frontend → Cloudflare Pages
1245
+
1246
+ Staged by `scripts/stage_pages.mjs`, deployed with `npx wrangler pages deploy`. The staging run measured on
1247
+ 2026-09-25 (`docs/DEPLOYMENT_DECISION.md` §7):
1248
+
1249
+ ```
1250
+ files staged : 60
1251
+ total bytes : 39,173,936 (37.36 MiB)
1252
+ largest file : assets/video/satquery-launch-50s.mp4 22,710,313 B (21.66 MiB)
1253
+ 25 MiB headroom left : 3,504,087 B on the largest file
1254
+ missing refs in staged : 0
1255
+ external network deps : 0 (HERMETIC)
1256
+ exit : 0
1257
+ ```
1258
+
1259
+ `_headers` and `robots.txt` must be **force-included** because no page references them; `provenance.json`
1260
+ and `CREDITS.md` likewise, because they are provenance records rather than assets
1261
+ (`docs/DEPLOYMENT_DECISION.md` §7).
1262
+
1263
+ ### 10.2 Backend → Render
1264
+
1265
+ `render.yaml` is the blueprint (§8.2); `main.py` exposes `app`
1266
+ (`uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT`).
1267
+
1268
+ ### 10.3 Inference → Codespace
1269
+
1270
+ `deploy/codespace/serve.py` serves `build_space_app()` on `$PORT`; `.devcontainer/` forwards port `8000`
1271
+ and runs the tunnel agent on start via `postStartCommand` (`release/repo/docs/DEPLOYMENT.md` §4).
1272
+
1273
+ ### 10.4 Repository writes use the GitHub Git Data API, not `git push`
1274
+
1275
+ > *"Repository writes are performed through the **GitHub Git Data API** (blob → tree → commit → `PATCH`
1276
+ > ref) with **sha256 byte-verification** of every uploaded blob. Deletions are expressed as `sha: null`
1277
+ > tree entries. This is used instead of `git push` so each deployed file is verified by content hash."*
1278
+ > (`release/repo/docs/DEPLOYMENT.md` §4)
1279
+
1280
+ The integrity check is recorded: *"Deployed files were re-read from the GitHub API and compared
1281
+ byte-for-byte against the local copies: **9 files sha256 byte-identical**, and the deployed HEAD re-read
1282
+ from the API."* (`release/repo/docs/DEPLOYMENT.md` §4.1; `docs/FINAL_DELIVERY_TODO.md` §6 E-10.)
1283
+
1284
+ ---
1285
+
1286
+ ## 11. The five historical backend blockers, and how the design closes them
1287
+
1288
+ `docs/DEPLOYMENT_DECISION.md` §8 enumerated **five** verified backend blockers that had to be closed
1289
+ before any backend could boot. `docs/DEPLOYMENT_TOPOLOGY.md` §4 carries them forward with the active
1290
+ design's response.
1291
+
1292
+ | # | Blocker (verified, old doc) | How the new topology addresses it |
1293
+ |---|---|---|
1294
+ | 1 | `requirements.txt` declared no `fastapi` / `uvicorn` / `httpx` / `starlette` | the Codespace/Render runtime installs the ASGI stack so `build_space_app()` and the gateway `app` can import |
1295
+ | 2 | No code read `$PORT` — a platform port would be ignored | `deploy/codespace/serve.py` binds `build_space_app()` to `$PORT`; Render reads its own `$PORT` |
1296
+ | 3 | Hand-rolled CORS; `OPTIONS` raised `405`, so browser preflight failed | the gateway registers `OPTIONS` explicitly (or relies on Starlette's CORS middleware) so preflight succeeds |
1297
+ | 4 | Module-level `app = create_app()` swallowed config errors into `app = None` | construction errors propagate (fail-fast) instead of silently leaving a dead `app` |
1298
+ | 5 | Adapter integrity unverified on load (`_adapter_sha256` computed but never compared) | the load path compares the computed digest against an expected value, or fails startup |
1299
+
1300
+ ### 11.1 The blockers, with their original verification
1301
+
1302
+ `docs/DEPLOYMENT_DECISION.md` §8 is the primary record, and it is more specific than the summary table:
1303
+
1304
+ | # | Blocker | State (as recorded) |
1305
+ |---|---|---|
1306
+ | 1 | `requirements.txt` declares no fastapi / uvicorn / httpx / starlette | **VERIFIED** |
1307
+ | 2 | No code reads `$PORT` — a platform-assigned port would be ignored | **VERIFIED** |
1308
+ | 3 | CORS is hand-rolled (`gateway/policy.py:429-451`); `policy.py:591` admits OPTIONS but routes register only GET/HEAD/POST (`gateway/app.py:334-346`), so Starlette raises **405** and browser preflight fails | **VERIFIED** |
1309
+ | 4 | Module-level `app = create_app()` swallows config errors into `app = None` (`gateway/app.py:683-690`) | **VERIFIED** |
1310
+ | 5 | Adapter integrity unverified on load — `_adapter_sha256` is computed and stored (`:258`, `:273`) but never compared against an expected digest | **VERIFIED** |
1311
+
1312
+ Blockers 3 and 4 are visible in the code this document cites. `gateway/app.py`'s own tail is blocker 4
1313
+ exactly:
1314
+
1315
+ ```python
1316
+ try: # pragma: no cover - depends on FastAPI being importable
1317
+ app = create_app()
1318
+ except Exception: # pragma: no cover - the sandbox path
1319
+ app = None # type: ignore[assignment]
1320
+ ```
1321
+
1322
+ and `deploy/render/main.py` is the fail-fast counterpart — its `create_app()` is called at module scope
1323
+ with no `try`, so a misconfiguration raises at import:
1324
+
1325
+ ```python
1326
+ # The ASGI object uvicorn imports: `uvicorn deploy.render.main:app`.
1327
+ app = create_app()
1328
+ ```
1329
+
1330
+ The CORS half of blocker 3 is closed in `deploy/render/main.py` by registering
1331
+ `CORSMiddleware`, whose comment names the defect it fixes:
1332
+
1333
+ > *"CORS fix: `CORSMiddleware` answers OPTIONS preflight itself, which resolves the earlier 405 on
1334
+ > preflight."*
1335
+
1336
+ ### 11.2 Status: closed by construction, not proven in production
1337
+
1338
+ `docs/DEPLOYMENT_TOPOLOGY.md` §4 is careful about the claim, and this document keeps that caution:
1339
+
1340
+ > *"They are recorded honestly here — the new infra (`deploy/render/`, `deploy/codespace/`) is **in
1341
+ > progress**, so treat these as *closed by construction / to be verified on first live run*, not as
1342
+ > already proven in production."*
1343
+
1344
+ **However**, the live deployment has since been exercised end-to-end: `docs/FINAL_DELIVERY_REPORT.md` §3
1345
+ records `/api/health` 200, `/api/capabilities` 200 with 6× `available:true`, `/api/infer {}` → 422
1346
+ `invalid_request` with `x-satquery-transport: tunnel`, and real inference for all six tasks. So the honest
1347
+ composite statement is: **the five blockers are closed in the deployed system as evidenced by the live
1348
+ behaviour recorded in the delivery documents, while `docs/DEPLOYMENT_TOPOLOGY.md` §4's own text still
1349
+ carries the earlier "in progress" framing.** Where the two disagree, the dated measurement is the stronger
1350
+ evidence, and it is cited here rather than substituted for the source's own words.
1351
+
1352
+ ---
1353
+
1354
+ ## 12. The superseded design, and exactly what did NOT change
1355
+
1356
+ ### 12.1 The historical topology
1357
+
1358
+ The superseded design ran inference on an **HF Space with ZeroGPU** (5 GPU-min/day,
1359
+ `@spaces.GPU(duration=…)` decoration) behind a **Railway** gateway
1360
+ (`docs/DEPLOYMENT_TOPOLOGY.md` §5; `docs/DEPLOYMENT_ARCHITECTURE.md` §1, §3).
1361
+
1362
+ | Old (superseded) | New (active) |
1363
+ |---|---|
1364
+ | Railway (gateway/API) | **Render** (orchestrator / API gateway) |
1365
+ | Hugging Face Space (inference) | **GitHub Codespace** (FastAPI inference) |
1366
+ | Cloudflare Pages | Cloudflare Pages (**unchanged**) |
1367
+ | Hugging Face (project/models) | Hugging Face (project card + pinned model references) |
1368
+
1369
+ Source: `docs/DEPLOYMENT_TOPOLOGY.md` §1.
1370
+
1371
+ ### 12.2 The three things that changed
1372
+
1373
+ `docs/DEPLOYMENT_TOPOLOGY.md` §5 enumerates them:
1374
+
1375
+ 1. **CPU-first instead of ZeroGPU.** No code change was required — `device_preference` honours
1376
+ `SATQUERY_DEVICE` and defaults to CPU, every specialist defaults to `device="cpu"`, and all placement
1377
+ is `.to(device)` (never `.cuda()`). ZeroGPU's GPU-minute quota and `@spaces.GPU` decoration are no
1378
+ longer on the critical path.
1379
+ 2. **A real, always-buildable inference environment.** A GitHub Codespace gives a reproducible container
1380
+ that builds and runs `build_space_app()` without a GPU quota or a Space's ephemeral-cold-start
1381
+ constraint. The wake flow (§4) replaces ZeroGPU lazy-loading as the cold-start story.
1382
+ 3. **No GPU quota to protect at the gateway.** Because there is no ZeroGPU budget, the gateway's
1383
+ rate/size limits remain as *fairness* controls, but the "never spend GPU quota on a shape-rejected
1384
+ request" rationale no longer dominates the design.
1385
+
1386
+ The CPU adaptation is independently verified in `docs/DEPLOYMENT_DECISION.md` §5, which lists the
1387
+ specific sites: `core/config.py:87-91` (`device_preference`), `specialists/vqa/model.py:234`
1388
+ (`float16` on cuda, **`float32` on cpu**), the per-specialist `device: str = "cpu"` defaults
1389
+ (`change/specialist.py:153`, `change/stanet.py:641`, `change/vqa_specialist.py:116`,
1390
+ `grounding/remoteclip.py:111`, `grounding/specialist.py:682`), `configs/base.yaml:293`
1391
+ (`cpu_mode_required: true`), and the fact that **no `.cuda()` call exists anywhere** — all placement is
1392
+ `.to(device)`.
1393
+
1394
+ ### 12.3 What did NOT change
1395
+
1396
+ `docs/DEPLOYMENT_TOPOLOGY.md` §5 closes with the list, and it is the most important part of this section:
1397
+
1398
+ > *"**What did NOT change:** the 4-endpoint contract, the gateway responsibility table, the env-var
1399
+ > vocabulary (only host names moved: `SATQUERY_SPACE_URL` → `SATQUERY_UPSTREAM_URL`), and the
1400
+ > `Config.hash == 78f1e3700da15aa1` freeze. The backend contract in `DEPLOYMENT_ARCHITECTURE.md` §1.1,
1401
+ > §2, §3.3, §4, §5 remains authoritative."*
1402
+
1403
+ Expanded:
1404
+
1405
+ | Unchanged artefact | Where it lives | Why it survived the host change |
1406
+ |---|---|---|
1407
+ | **The 4-endpoint contract** | `docs/API_CONTRACT.md`; `app/space_app.py`; `gateway/app.py::PROXIED_ROUTES` | it is a *client-facing* contract; hosts are an implementation detail |
1408
+ | **The gateway responsibility table** | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 | the responsibilities are the same regardless of who hosts the upstream |
1409
+ | **The env-var vocabulary** | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 | only `SATQUERY_SPACE_URL` → `SATQUERY_UPSTREAM_URL` moved |
1410
+ | **The config freeze `78f1e3700da15aa1`** | `core/config.py::Config.hash`; `configs/base.yaml` | the deployment was changed *around* the config, never inside it |
1411
+ | **The gateway failure-mode table** | `docs/DEPLOYMENT_ARCHITECTURE.md` §5 | still governs, host names aside |
1412
+ | **The entrypoint requirements** | `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3 | import cheaply without torch; reuse `app.serving`; degrade don't crash; never load a model for a metadata request; honour the config hash |
1413
+
1414
+ ### 12.4 The frozen paperwork
1415
+
1416
+ `configs/deploy.yaml` still describes an **HF Space + Gradio + ZeroGPU** target, and it is left
1417
+ **undisturbed** (`docs/DEPLOYMENT_TOPOLOGY.md` §3.4; `docs/DEPLOYMENT_DECISION.md` §4). The reasoning is
1418
+ structural, not sentimental, and it is a good example of why the config freeze matters:
1419
+
1420
+ 1. **`Config.hash` cannot move.** `core/config.py` reads only `configs/base.yaml`. `configs/deploy.yaml`
1421
+ carries `registry: false` and is never loaded — **but** `scripts/validate_deploy_config.py` hard-fails
1422
+ if the `deployment:` block in `deploy.yaml` differs key-for-key from `base.yaml`'s (assertions at
1423
+ `:114-131`). So changing `zerogpu: true` → `false` in `deploy.yaml` alone fails the validator, and
1424
+ moving `base.yaml` to match moves the frozen hash. **Both paths are closed.**
1425
+ 2. **There is no Gradio runtime to conflict with.** No `import gradio`, no `gr.Blocks`, no `gr.Interface`
1426
+ and no Gradio entrypoint exists anywhere. Gradio appears only as `requirements.txt:36` and the manifest
1427
+ value `sdk: gradio` (`configs/base.yaml:284`). The one ZeroGPU code path —
1428
+ `spaces.GPU(duration=duration)` at `app/space_app.py:165` — sits inside `decorate_gpu()`, which **is
1429
+ never applied to any route**; routes use plain `@api.get`/`@api.post` at `:521/:549/:555/:661`. The real
1430
+ entrypoint is FastAPI: `build_space_app()` at `app/space_app.py:409`.
1431
+
1432
+ > *"Conclusion: the frozen contract describes a Gradio Space that does not exist in code. It is frozen
1433
+ > paperwork, not a competing deployment."* (`docs/DEPLOYMENT_DECISION.md` §4)
1434
+
1435
+ `configs/base.yaml` still carries the frozen ZeroGPU declarations, and `app/space_app.py` transcribes the
1436
+ durations into `GPU_DURATIONS`:
1437
+
1438
+ ```python
1439
+ GPU_DURATIONS: dict[str, int] = {
1440
+ "vqa": 20,
1441
+ "caption": 20,
1442
+ "grounding": 45,
1443
+ "change": 30,
1444
+ "optical_sar": 45,
1445
+ "change_vqa": 30,
1446
+ }
1447
+ ```
1448
+
1449
+ with `change_vqa` reusing the `change` budget **because adding a key of its own would move `Config.hash`**
1450
+ (`docs/DEPLOYMENT_ARCHITECTURE.md` §3.4; `app/space_app.py`). And the decoration is applied conditionally,
1451
+ because `spaces` is not installed on a CPU host:
1452
+
1453
+ ```python
1454
+ def decorate_gpu(task: str) -> Callable[[Callable[..., Any]], Callable[..., Any]]:
1455
+ ...
1456
+ spaces = _spaces_module()
1457
+ if spaces is None or not hasattr(spaces, "GPU"):
1458
+ def _identity(fn): return fn
1459
+ return _identity
1460
+ return spaces.GPU(duration=duration)
1461
+ ```
1462
+
1463
+ > **The honest status of the ZeroGPU path:** *"the ZeroGPU decoration has **never executed** here. It is
1464
+ > specified from finding C-8 and the frozen `gpu_duration_*` values, and that is all it is."*
1465
+ > (`app/space_app.py` docstring; `docs/PHASE19_FINAL_HARDENING.md`.)
1466
+
1467
+ ---
1468
+
1469
+ ## 13. Failure modes and their handling
1470
+
1471
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §5 is the authoritative table. Reproduced, with the active host names:
1472
+
1473
+ | Failure | Detected by | Surface | Recovery |
1474
+ |---|---|---|---|
1475
+ | Upstream cold start | gateway upstream timeout | `504` with `recoverable: true` | client retries once, manually |
1476
+ | Model absent | `capabilities[].available: false` | `503 model_unavailable` | capability disabled in the UI |
1477
+ | Model corrupt | `ModelLoadError` | `503 model_load_error` | **defect** — report it |
1478
+ | GPU quota exhausted | allocation error | `503` | wait for the daily reset |
1479
+ | Request too large | gateway size check | `413` | client re-encodes |
1480
+ | Upload content type absent or refused | content-type allowlist | `415` | client sends a supported type; **the server does not guess** |
1481
+ | Uploaded handle expired or unknown | store lookup on read | `400 input_error` | re-upload; handles are ephemeral by design |
1482
+ | Asset store not configured or full | store construction / capacity check | `503` | ⚠️ **distinguishing these two needs an instrument the deployment does not expose** |
1483
+ | Asset root configured but unusable | **nothing** — `get_asset_store()` raises outside the route's `try` | **`500 text/plain`** on the upstream directly; the gateway masks it as `502` | ⚠️ bounded defect (F-12) |
1484
+ | Framework error upstream (`404`/`405`) | **nothing** on the upstream | `{"detail": …}` upstream; the gateway masks it as an envelope | ⚠️ bounded defect (F-12b) — **now fixed upstream too** (`app/space_app.py` registers the handler) |
1485
+ | Malformed body | gateway schema validation | `422` | client bug |
1486
+ | `trace.inputs` echoing a path | **nothing** | `200` with a server-side path | ✅ fixed (F-13) — `core/controller.py::_asset_label` |
1487
+ | `trace.steps[PARSE].detail["inputs"]` echoing the same path | **nothing** | `200` with a path in the `PARSE` step record | ✅ fixed (F-14) |
1488
+ | A construction failure's exception string reaching the client | **nothing** | `200` with a path in `result.warnings[]`, `evidence[].payload["message"]`, and the registry block **twice** | ✅ fixed (F-15) — path-scrubbed to a **basename**, raw detail logged server-side; **four** live carriers, not three |
1489
+ | `artifact_ref` / `result.change_map` carrying a path | **nothing**, and only when `artifact_dir` is configured | `200` with a path where the contract documents an `artifact://` URI | ✅ fixed (F-16) — refs are `null`, no `artifact://` fabricated, explicit non-retrievable warning |
1490
+ | `change_vqa.artifact_dir` configured but never read | **nothing** — the key is accepted and silently ignored | **no surface at all** | ⚠️ documented, not patched (F-17) |
1491
+ | Upstream unreachable | gateway connection error | `502` | report; do not silently retry analyze |
1492
+ | Analysis exceeds budget | `SpecialistTimeoutError` | `504`, `recoverable: true` | offer a retry |
1493
+ | Non-JSON response upstream | gateway parse check | `502` with the upstream body logged | **defect** |
1494
+ | Per-IP rate limit bypassed | **not detected** | no `429` is produced | fairness only; **not** a protection control (§2.4) |
1495
+
1496
+ ### 13.1 Two failure modes that the gateway and the upstream now agree on
1497
+
1498
+ The `404`/`405` and unhandled-exception rows were originally *upstream* holes that the gateway masked.
1499
+ Both are now closed **on the upstream as well**, so a client following the runbook to the upstream's own
1500
+ URL gets the same envelope as a client going through the gateway. `app/space_app.py` registers both
1501
+ handlers, and its comment records the measurement that forced it:
1502
+
1503
+ > *"Measured, direct to the Space, before this fix: `GET /v1/whocares -> 404 {"detail":"Not Found"}`,
1504
+ > `GET /v1/assets -> 405 {"detail":"Method Not Allowed"}`, an unwrapped failure -> `500 text/plain`, no
1505
+ > envelope at all."*
1506
+
1507
+ ### 13.2 A saturated asset store is indistinguishable from a misconfigured one
1508
+
1509
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 records finding F-11 and its resolution. `POST /v1/assets` answers
1510
+ `503` in two unrelated situations — the store is **not configured**, or the store is **full** — with the
1511
+ same status and the same envelope shape, so *"a client and an operator cannot tell them apart from a
1512
+ response."*
1513
+
1514
+ The one value that would have separated them (`capacity_refusals` from `AssetStore.stats()`) was computed
1515
+ on every request and read by nothing. **RESOLVED 2026-09-23 by owner ruling — the unused computation was
1516
+ REMOVED, not given a consumer.** The owner's reasoning: *"a metrics surface with no reader is a cost paid
1517
+ on every request for an instrument nobody holds."* The ambiguity itself **remains**, and the document says
1518
+ so:
1519
+
1520
+ > *"Removing the counter did NOT remove the ambiguity. The two `503` causes remain indistinguishable from a
1521
+ > response, and the deployment still **does not expose** an instrument that tells them apart."*
1522
+
1523
+ The remedy is unchanged: `SATQUERY_ASSET_MAX_FILES` / `SATQUERY_ASSET_TTL_S` if the store is saturating,
1524
+ and those two variables if it is unconfigured — but **confirming which requires inspecting the
1525
+ deployment**, because the response will not say.
1526
+
1527
+ ### 13.3 A deployment precondition list, carried forward
1528
+
1529
+ `docs/DEPLOYMENT_TOPOLOGY.md` §6, with host names updated:
1530
+
1531
+ 1. Cloudflare Pages project name / domain (needed for the deploy command and `robots.txt` sitemap).
1532
+ 2. Artifacts present, or capabilities shipped `available: false` (change head, change_vqa head,
1533
+ calibration JSON) — degrades honestly, not broken.
1534
+ 3. `HF_TOKEN` set on Render **if** the HF proxy path is used (not used in the live config).
1535
+ 4. Codespace `.devcontainer/` forwarding `:8000` **and** starting the tunnel agent.
1536
+ 5. The five blockers in §11 closed and verified on the first live run.
1537
+
1538
+ ---
1539
+
1540
+ ## 14. What is deliberately absent from the deployment
1541
+
1542
+ `docs/DEPLOYMENT_ARCHITECTURE.md` §6 records the exclusions so that omission is not mistaken for
1543
+ oversight. From plan §73/§74:
1544
+
1545
+ * **No Kubernetes, no Docker swarm.** Render plus one Codespace is the whole fleet.
1546
+ * **No Kafka, no Redis cluster, no queue.** Requests are synchronous.
1547
+ * **No autoscaling.** The free tier has a fixed quota; autoscaling cannot raise it.
1548
+ * **No multi-tenant isolation, no auth, no user accounts.**
1549
+ * **No second VLM and no foundation-model retraining.**
1550
+ * **No vector database.** The retriever-free RAG decision is separate and upstream.
1551
+ * **No database, no session store** at the gateway (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.2).
1552
+
1553
+ `docs/FRONTEND_INTEGRATION.md` §7 adds the client-side counterpart: **no login screen**, because *"There
1554
+ is none to build (plan §74)."*
1555
+
1556
+ ---
1557
+
1558
+ ## 15. What is `NOT RUN` / `OPEN` / `BLOCKED` for this topic
1559
+
1560
+ | Item | Status | Note |
1561
+ |---|---|---|
1562
+ | **B-07** transient tunnel-agent gaps | **OPEN** | patch prepared, **not deployed**; worst case ≈ 249 s (§6) |
1563
+ | **B-02** `codespace_name` trailing `\n` | **OPEN (cosmetic)** | reporting only; the wake path strips (§7.4) |
1564
+ | **B-03** local `deploy/` stale + untracked | **KNOWN** | the working copy is not the deployed source (§9.5) |
1565
+ | **B-06** Render free-tier sleep / Codespace idle 30 min | **KNOWN** | cold start delay; documented, not hidden (§5) |
1566
+ | Tunnel agent source (`tunnel_agent.py`) | **UNKNOWN** | not in the monorepo; `UNKNOWN — not established from the available evidence` (§9.4) |
1567
+ | Cold-start latency distribution | **NOT MEASURED** | "tens of seconds" is a documented expectation; no distribution exists (§5) |
1568
+ | Throughput / concurrency characterisation | **NOT RUN** | `docs/FRONTEND_INTEGRATION.md` §9: *"Latency is not characterized."* |
1569
+ | ZeroGPU decoration execution | **NOT RUN** | never executed anywhere; CPU path only (`app/space_app.py`; `docs/PHASE19_FINAL_HARDENING.md`) |
1570
+ | Sequential-request test under `cache_max_models=1` | **NOT DONE — environment-blocked** | requires a reachable upstream (`docs/DEPLOYMENT_ARCHITECTURE.md` §7) |
1571
+ | Gateway's rate limiter as an abuse control | **REJECTED** | ruled fairness-only, 2026-09-23 (§2.4) |
1572
+ | `HF_TOKEN` proxy path | **not used live** | absent from the live Render config (§8.1) |
1573
+ | `SATQUERY_UPSTREAM_URL` | **not used live** | absent from the live Render config (§8.1) |
1574
+ | A second copy of the capability table at the gateway | **REJECTED** | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 |
1575
+ | End-to-end benchmark of the deployed stack | **does not exist** | no system-level accuracy is claimed anywhere |
1576
+ | Asset-store 503 disambiguation instrument | **absent** | removed by ruling; ambiguity remains (§13.2) |
1577
+
1578
+ ---
1579
+
1580
+ ## 16. Where the evidence lives
1581
+
1582
+ | Claim | Source |
1583
+ |---|---|
1584
+ | four tiers, host names, tunnel direction | `docs/DEPLOYMENT_TOPOLOGY.md` §1, §2; `docs/FINAL_DELIVERY_TODO.md` §1.2 |
1585
+ | gateway rationale (three reasons) | `docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 |
1586
+ | gateway responsibility table | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 |
1587
+ | what the gateway must NOT do | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 |
1588
+ | CORS assembly + wildcard refusal | `deploy/render/main.py::_allowed_origins`, `_DEV_ORIGINS`, `_PRODUCTION_ORIGINS` |
1589
+ | CORS header filtering assertion | `gateway/app.py::_proxy` (F-2) |
1590
+ | per-file cap 4,194,304 B | `gateway/policy.py:221`; `app/space_app.py::_asset_max_file_bytes` |
1591
+ | body cap 8 MiB + F-6 measurement | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 |
1592
+ | rate-limiter ruling + measurement | `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2 |
1593
+ | no-retry rule | `gateway/app.py::_proxy`; `docs/DEPLOYMENT_TOPOLOGY.md` §2 |
1594
+ | request-id injection | `gateway/app.py::_proxy` |
1595
+ | 404/405 envelope handler | `gateway/app.py`; `app/space_app.py` |
1596
+ | error-envelope shape | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.3 |
1597
+ | orchestrator error classes + statuses | `deploy/render/main.py` |
1598
+ | transport-failure classification | `gateway/app.py::_transport_failure_detail` (F-15c) |
1599
+ | proxied / costly / blocked route tuples | `gateway/app.py::PROXIED_ROUTES`, `COSTLY_ROUTES`, `BLOCKED_ROUTES` |
1600
+ | forwarded port returns 302 | `docs/DEPLOYMENT_TOPOLOGY.md` measured note; `release/repo/docs/DEPLOYMENT.md` §7 |
1601
+ | tunnel rationale for private repos | `deploy/codespace/launch.sh` header comment |
1602
+ | `x-satquery-transport` as proof | `tests/unit/test_frontend_live_wiring.py`; `docs/FINAL_DELIVERY_TODO.md` §6 E-03 |
1603
+ | wake flow | `docs/DEPLOYMENT_TOPOLOGY.md` §2; `deploy/render/main.py::ensure_codespace_up` |
1604
+ | `X-SatQuery-State` header | `deploy/render/main.py::infer` |
1605
+ | cold start, documented not hidden | `docs/DEPLOYMENT_TOPOLOGY.md` §2; `docs/FRONTEND_INTEGRATION.md` §6 |
1606
+ | Codespace idle 30 min | `docs/FINAL_DELIVERY_TODO.md` §6 E-04 |
1607
+ | B-07 root shape + ≈249 s | `docs/FINAL_DELIVERY_TODO.md` §5; `DELIVERY_REPORT_2026-09-25.md` §4 |
1608
+ | B-07 patch contents + verification | `DELIVERY_REPORT_2026-09-25.md` §4; `docs/FINAL_DELIVERY_TODO.md` §6 E-12 |
1609
+ | live health payload | `release/repo/docs/DEPLOYMENT.md` §2 |
1610
+ | `completed` readings 97 / 314 / 338 | `release/repo/docs/DEPLOYMENT.md` §2; `docs/FINAL_DELIVERY_TODO.md` §1.4; `docs/FINAL_DELIVERY_REPORT.md` §3 |
1611
+ | B-02 cosmetic | `docs/FINAL_DELIVERY_REPORT.md` §6; `docs/FINAL_DELIVERY_TODO.md` §5 |
1612
+ | live Render env vars | `docs/DEPLOYMENT_TOPOLOGY.md` header note; `release/repo/docs/DEPLOYMENT.md` §3.1 |
1613
+ | blueprint env vars | `render.yaml` |
1614
+ | Codespace env vars | `.devcontainer/devcontainer.json`; `deploy/codespace/launch.sh` |
1615
+ | env-var vocabulary + F-7/F-8/F-9 | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 |
1616
+ | `serve.py` entrypoint | `deploy/codespace/serve.py` |
1617
+ | launcher stages 0–3 | `deploy/codespace/launch.sh` |
1618
+ | tunnel agent existence + gap | `docs/FINAL_DELIVERY_TODO.md` §1.3; `Glob`/`git ls-files` on the monorepo |
1619
+ | repos of record | `docs/FINAL_DELIVERY_TODO.md` §1.3 |
1620
+ | deployed revisions | `release/repo/docs/DEPLOYMENT.md` §1 |
1621
+ | staging measurement | `docs/DEPLOYMENT_DECISION.md` §7 |
1622
+ | Git Data API + sha256 verification | `release/repo/docs/DEPLOYMENT.md` §4, §4.1 |
1623
+ | five blockers + original verification | `docs/DEPLOYMENT_DECISION.md` §8; `docs/DEPLOYMENT_TOPOLOGY.md` §4 |
1624
+ | superseded design + what did not change | `docs/DEPLOYMENT_TOPOLOGY.md` §5 |
1625
+ | frozen paperwork | `docs/DEPLOYMENT_DECISION.md` §4; `docs/DEPLOYMENT_TOPOLOGY.md` §3.4 |
1626
+ | `GPU_DURATIONS` | `app/space_app.py` |
1627
+ | failure modes | `docs/DEPLOYMENT_ARCHITECTURE.md` §5 |
1628
+ | asset-store ambiguity (F-11) | `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 |
1629
+ | deliberate exclusions | `docs/DEPLOYMENT_ARCHITECTURE.md` §6; `docs/FRONTEND_INTEGRATION.md` §7 |
1630
+
1631
+ ---
1632
+
1633
+ *Continue to [03 — Request lifecycle](03-request-lifecycle.md).*