thundercode commited on
Commit
d8660af
·
verified ·
1 Parent(s): 6df4c6e

release: add docs/OPERATIONS.md

Browse files
Files changed (1) hide show
  1. docs/OPERATIONS.md +1185 -0
docs/OPERATIONS.md ADDED
@@ -0,0 +1,1185 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Operations Manual
2
+
3
+ **Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` ·
4
+ `DEFERRED` · `OPEN` · `RESOLVED` · `BY DESIGN`.
5
+
6
+ This is the operator-facing manual for the **live SatQuery AI stack**. It answers four questions that
7
+ the architecture chapters deliberately do not: *what is running, right now, and who owns it*; *how do I
8
+ bring it up and keep it up*; *how do I tell a transient transport gap from a real failure*; and *what
9
+ monitoring, capacity and cost machinery is **absent** so I do not assume it exists*.
10
+
11
+ It is written for the person who has to make the system answer a question in front of an audience, and
12
+ for the person who has to diagnose it at 23:00 when it does not.
13
+
14
+ > **Read this first.** The live system runs across **three private repositories** plus the public
15
+ > umbrella. The monorepo working copy — including the `deploy/` directory inside it — is **not** the
16
+ > deployed source. `deploy/` in the monorepo is **stale and untracked** (`git status` reports
17
+ > `?? deploy/`; verified in the working copy). Any operational fix must be applied to the real
18
+ > repositories, never to the monorepo copy (`docs/DEPLOYMENT.md` §1, §10).
19
+
20
+ > **Second rule.** Two known defects are `OPEN` and are **not** fixed in production: **B-07** (transient
21
+ > tunnel gaps) and **B-02** (a cosmetic trailing newline in one health field). Neither may be described
22
+ > as resolved, and the B-07 patch is **prepared but NOT deployed**. This document never upgrades them.
23
+
24
+ ---
25
+
26
+ ## Table of contents
27
+
28
+ **Part I — The operational model**
29
+ 1. What runs where
30
+ 2. The tier inventory, with the deployed revision of each
31
+ 3. Who owns what
32
+ 4. The operational invariants (four rules that must never be broken)
33
+ 5. The stale-copy problem, in operational terms
34
+
35
+ **Part II — The runbook**
36
+ 6. Warm the stack before a demo
37
+ 7. Restart after an idle-stop
38
+ 8. Tell a tunnel gap from a real failure
39
+ 9. What the client shows while waking
40
+
41
+ **Part III — Cold start and the timing budget**
42
+ 10. The cold-start shape
43
+ 11. The four timeouts, and why the relationship matters
44
+ 12. The worst case: ≈249 s under B-07
45
+
46
+ **Part IV — The known defects, in operational terms**
47
+ 13. B-07 — transient tunnel gaps (`OPEN`)
48
+ 14. B-02 — the trailing newline (`OPEN`, cosmetic)
49
+
50
+ **Part V — Monitoring and alerting**
51
+ 15. What exists
52
+ 16. What does **not** exist
53
+ 17. Why "no alerting" is a design fact, not an oversight
54
+
55
+ **Part VI — Capacity and cost**
56
+ 18. The capacity shape
57
+ 19. The cost shape, and the one absence that has a code artifact
58
+
59
+ **Part VII — Incident triage**
60
+ 20. The symptom → cause → check → action table
61
+ 21. The consolidated decision tree
62
+ 22. The machine-code reference
63
+
64
+ **Part VIII — Routine maintenance**
65
+ 23. Rotating credentials (procedure only)
66
+ 24. Restarting the tunnel agent
67
+ 25. Re-warming the model cache
68
+ 26. Deploying a change (the Git Data API path)
69
+
70
+ **Part IX — Known operational gaps**
71
+ 27. The explicit gaps list
72
+ 28. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for operations
73
+
74
+ **Part X — Evidence**
75
+ 29. Where the evidence lives
76
+
77
+ ---
78
+
79
+ # Part I — The operational model
80
+
81
+ ## 1. What runs where
82
+
83
+ The live stack is four tiers in a straight line, plus a model tier that is reached *through* the
84
+ inference tier rather than by the user (`release/repo/docs/DEPLOYMENT.md` §2):
85
+
86
+ ```
87
+ Browser
88
+ │ HTTPS
89
+ ▼
90
+ Cloudflare Pages — satquery.pages.dev (static frontend, 11 pages)
91
+ │ HTTPS / JSON → /api/*
92
+ ▼
93
+ Render — satquery-backend-m4yv.onrender.com (orchestrator / API gateway)
94
+ │ outbound long-poll POST /tunnel/agent
95
+ ▼
96
+ GitHub Codespace — FastAPI inference, CPU, port 8000
97
+ │ build_space_app()
98
+ ▼
99
+ specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
100
+ │
101
+ ▼
102
+ ResultEnvelope → tunnel → Render → browser
103
+ ```
104
+
105
+ ```mermaid
106
+ flowchart LR
107
+ U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"]
108
+ CF -->|"HTTPS JSON<br/>/api/health · /api/capabilities · /api/infer · /api/assets"| R["Render<br/>orchestrator / gateway"]
109
+ R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"]
110
+ C --> S[(SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet)]
111
+ C -->|ResultEnvelope| R
112
+ R -->|"envelope + error translation"| CF
113
+ ```
114
+
115
+ Three properties of this diagram matter operationally, and each is the subject of a section below:
116
+
117
+ 1. **The transport is an outbound tunnel, not an inbound port.** The Codespace dials *out* to Render.
118
+ Render never dials into the Codespace. The transport is therefore alive only while an agent process
119
+ is polling — which is why "is the agent connected?" is the single most important operational
120
+ question (§6, §8).
121
+ 2. **There is exactly one inference host.** One Codespace, one Render service, no replicas, no
122
+ autoscaling (`render.yaml` declares a single web service with `plan: free`; plan §74 lists
123
+ `autoscaling` under **Not included**). Capacity is therefore bounded by that one host (§18).
124
+ 3. **Inference is CPU-only.** `SATQUERY_DEVICE=cpu` is set on both the orchestrator and the Codespace
125
+ (`render.yaml`, `.devcontainer/devcontainer.json`), and every specialist defaults to `device="cpu"`
126
+ (`docs/DEPLOYMENT_DECISION.md` §5). No GPU path is on the live critical path.
127
+
128
+ ## 2. The tier inventory, with the deployed revision of each
129
+
130
+ Read from the GitHub API during the release reconnaissance (`release/CURRENT_RELEASE_STATE.md` §1;
131
+ `release/repo/docs/DEPLOYMENT.md` §1):
132
+
133
+ | Component | Repository | Visibility | Branch | Revision | Host |
134
+ |---|---|---|---|---|---|
135
+ | Frontend | `Anish-lab-blip/SatQuery-Frontend` | **private** | `main` | **`2d7ae53b482d`** | Cloudflare Pages → `satquery.pages.dev` |
136
+ | Backend / orchestrator | `Anish-lab-blip/SatQuery-Backend` | **private** | `main` | **`89d80eaddec5`** | Render → `satquery-backend-m4yv.onrender.com` |
137
+ | Inference | `Anish-lab-blip/SatQuery-Inference` | **private** | `main` | **`5a0936ace491`** | Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port 8000, via outbound tunnel |
138
+ | Public umbrella | `Anish-lab-blip/SatQuery-AI` | **public** | `main` | `3dcabd32da41` ("Initial commit") | this release home |
139
+ | Monorepo (working copy) | `C:/Users/anish/satquery-ai` | local only | `master` | `9d57aed` | **no git remote**; 334 dirty entries |
140
+ | Hugging Face | `thundercode/SatQuery` | **public** | `main` | lastModified `2026-09-25T16:26:53Z` | model tier |
141
+
142
+ The three private repositories are private **by design**; their links return 404 for an outside
143
+ audience (`release/CURRENT_RELEASE_STATE.md` §6). An operator therefore cannot browse the deployed
144
+ source from a public URL — the deployed files must be fetched with an authenticated API call
145
+ (`release/repo/docs/DEPLOYMENT.md` §7.1, §10).
146
+
147
+ > **The monorepo's `deploy/` is not the deployed source.** This is the single most important trap in
148
+ > the whole system (§5).
149
+
150
+ ## 3. Who owns what
151
+
152
+ The ownership table below is derived from the code and the deployment records. "Owner" means *the
153
+ person or role that must act when this tier misbehaves*.
154
+
155
+ | Tier | Owner | What they own | What they must never do |
156
+ |---|---|---|---|
157
+ | Cloudflare Pages (frontend) | Frontend maintainer | the static bundle, `_headers`, `robots.txt`, the Analyze console | add a server-side secret — the tier holds none |
158
+ | Render (gateway) | Backend maintainer | the orchestrator revision, the env-var set, the CORS allowlist, the tunnel hub state | retry `POST /api/infer` (§4) |
159
+ | Codespace (inference) | Inference maintainer | the Codespace, the tunnel agent, the asset directory, the HF cache | let a stale serve process keep answering (§5) |
160
+ | Hugging Face (model tier) | Release owner | model cards, the pinned model references, the released checksums | treat the Hub as the runtime inference host — it is not |
161
+ | Credentials | Owner (human) | the GitHub PAT and the account tokens | record any credential value in a public document (§23) |
162
+
163
+ Two decisions are explicitly **not** an agent's to make, and both gate operational change
164
+ (`docs/PHASE19_FINAL_HARDENING.md` §7): the **SDK choice** (irrelevant on the live path, but still
165
+ unmade for the frozen manifest) and the **rate-limit / size-limit values** (which bound one client's
166
+ share of capacity). Neither is needed to operate the system as deployed.
167
+
168
+ ## 4. The operational invariants
169
+
170
+ Four rules are load-bearing. Each is enforced somewhere in code or config, and each has a documented
171
+ failure if broken.
172
+
173
+ ### 4.1 The config hash is frozen
174
+
175
+ `Config.hash == 78f1e3700da15aa1`. The loader computes a sha256 over the whole registry
176
+ (`core/config.py:76-80`) and every evaluation run records it. **Editing `configs/base.yaml` moves the
177
+ hash and invalidates every artifact keyed to it** (`release/repo/docs/DEPLOYMENT.md` §11;
178
+ `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §6.1). Deployment state that must *not* move the hash — asset-store
179
+ capacity, TTL, the per-file cap — is read from the **environment**, not from the YAML
180
+ (`release/repo/README.md` §Installation).
181
+
182
+ Operational consequence: **never edit `configs/base.yaml` to point at a deployment artifact.** The
183
+ serving path wires checkpoints through the registry's `builders=` override precisely so it does not have
184
+ to (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §6.1).
185
+
186
+ ### 4.2 The gateway never retries `POST /api/infer`
187
+
188
+ > *"Render must not retry `POST /api/infer` on its own — a retry would consume inference a second
189
+ > time. The client decides on retry."* (`docs/DEPLOYMENT_TOPOLOGY.md` §2)
190
+
191
+ This is stated in three places (`docs/DEPLOYMENT_TOPOLOGY.md` §2, `docs/DEPLOYMENT.md` §7,
192
+ `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.3) because a retry is the natural thing to add and the wrong
193
+ thing to add. On the live CPU deployment it wastes compute; on the historical ZeroGPU target it spent a
194
+ metered GPU-minute twice.
195
+
196
+ ### 4.3 The CORS allowlist is explicit and never a wildcard
197
+
198
+ The gateway assembles its allowlist from `SATQUERY_ALLOWED_ORIGINS` plus a hard-coded production origin
199
+ plus a fixed list of development origins (`deploy/render/main.py:139-216`). A `*` raises
200
+ (`deploy/render/main.py:204-208`). The live value is `https://satquery.pages.dev`
201
+ (`release/repo/docs/DEPLOYMENT.md` §6.1).
202
+
203
+ Operational consequence: a new frontend origin must be **added** to the env var; it will not work by
204
+ accident.
205
+
206
+ ### 4.4 One inference host, one asset store
207
+
208
+ There is one Codespace, and its filesystem is **ephemeral** (`deploy/codespace/launch.sh:44-47`). Uploaded
209
+ assets are written under `SATQUERY_ASSET_DIR` (default `/tmp/satquery-assets`) and are TTL'd (900 s
210
+ default). A Codespace restart empties the store and makes every previously issued handle unresolvable
211
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §3.1.1).
212
+
213
+ Operational consequence: a handle that worked seconds ago may return `400 input_error` after a restart.
214
+ That is documented behaviour, not a bug (§20).
215
+
216
+ ## 5. The stale-copy problem, in operational terms
217
+
218
+ Three copies of the deployment code exist, and confusing them is the most expensive operational
219
+ mistake in the system.
220
+
221
+ | Copy | What it is | Trustworthy? |
222
+ |---|---|---|
223
+ | the monorepo `deploy/` | local, **untracked** (`?? deploy/`), stale | **no** — it is not the deployed source |
224
+ | the session scratch copy | a local copy used to author and verify the B-07 patch | **no** — it is "deployed + patch", not deployed |
225
+ | the private repositories | the real deployed source | **yes** — fetch it before editing |
226
+
227
+ Evidence for the divergence is direct. The monorepo's `deploy/render/main.py` (532 lines) exposes
228
+ `/api/health` with a `config` block that has **no** `tunnel` field and **no** `transport_mode`,
229
+ `tunnel_timeout_s` or `wake_timeout_s` keys (`deploy/render/main.py:444-466`), whereas the **live**
230
+ payload carries all of them (`release/repo/docs/DEPLOYMENT.md` §5). The monorepo copy also contains no
231
+ `tunnel_agent.py` and no `doctor.sh`, even though `deploy/codespace/launch.sh` invokes both
232
+ (`deploy/codespace/launch.sh:83,89,153,160-168,179`). The two are different programs.
233
+
234
+ > **Operational rule.** Before changing anything, fetch the deployed `main.py` from the private
235
+ > repository and diff it against what you are about to edit. The monorepo copy will silently disagree.
236
+
237
+ The `deploy/codespace/launch.sh` file *is* useful as documentation of intent — its header explains why
238
+ it is defensive (`deploy/codespace/launch.sh:14-24`) — but it is a copy, and its references to
239
+ `tunnel_agent.py` resolve only in the deployed repository.
240
+
241
+ ---
242
+
243
+ # Part II — The runbook
244
+
245
+ Every step in this part is grounded in a file. Commands are quoted as they appear in the sources.
246
+
247
+ ## 6. Warm the stack before a demo
248
+
249
+ Three things can be cold, and all three are warmed differently (`docs/architecture/10-observability-and-ops.md`
250
+ §7.1):
251
+
252
+ | What is cold | How it warms | Bound |
253
+ |---|---|---|
254
+ | the Render orchestrator | the first request to any `/api/*` route | Render free tier sleep/wake cycle |
255
+ | the Codespace | `ensure_codespace_up()` starts it and polls | `wake_timeout_s = 120` |
256
+ | the HF model cache | `warm_cache.py`, or the first model-touching request | one download per model |
257
+
258
+ ### Step 1 — confirm the tunnel agent is connected
259
+
260
+ This is the check that matters most, because the tunnel is the live transport and the forwarded-port
261
+ path is dead for a private repository (`release/CURRENT_RELEASE_STATE.md` §2: *"The forwarded-port path
262
+ is **dead** (302 for a private repo); the tunnel is the live transport."*).
263
+
264
+ ```bash
265
+ curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health
266
+ ```
267
+
268
+ Read `tunnel.agent_connected`. `true` means an agent has polled recently; `false` means **no agent has
269
+ polled recently**, and under `transport_mode: auto` a request will then take the forward path, which for
270
+ a private repository fails after burning the wake timeout (`docs/architecture/10-observability-and-ops.md`
271
+ §7.1).
272
+
273
+ > **`--noproxy '*'` is not optional in the authoring sandbox.** The sandbox proxy is dead; without the
274
+ > flag the request fails before reaching Render. In a normal environment the flag is harmless
275
+ > (`release/repo/docs/REPRODUCIBILITY.md` §10.1).
276
+
277
+ ### Step 2 — if `agent_connected` is false, start the Codespace
278
+
279
+ The agent is launched by the devcontainer's `postStartCommand`, which runs `launch.sh`:
280
+
281
+ ```json
282
+ "postStartCommand": "bash deploy/codespace/launch.sh"
283
+ ```
284
+
285
+ (`.devcontainer/devcontainer.json:18`)
286
+
287
+ Starting the Codespace is what re-runs `postStartCommand` and therefore reconnects the agent (§7).
288
+ `launch.sh` is defensive by design, and its own header explains why:
289
+
290
+ > *"`setsid` alone is NOT enough in Codespaces. The lifecycle shell that runs postStartCommand can still
291
+ > reap the process group, which showed up in production as 'the agent announced once, then vanished' —
292
+ > the hub then reported agent_connected=false and /api/infer fell back to the dead forwarded-port path
293
+ > (401 -> wake_timeout)."* (`deploy/codespace/launch.sh:14-19`)
294
+
295
+ The script therefore uses `setsid + nohup + </dev/null` plus a **supervising wrapper** that restarts the
296
+ agent if it exits (`deploy/codespace/launch.sh:20-23,160-168`), and then **verifies** the agent came up:
297
+
298
+ ```bash
299
+ sleep 4
300
+
301
+ if ! pgrep -f "deploy/codespace/tunnel_agent.py" > /dev/null 2>&1; then
302
+ echo "WARNING: the tunnel agent is not running. Last log lines:" >&2
303
+ ...
304
+ else
305
+ echo "tunnel agent process is up (pid $(pgrep -f 'deploy/codespace/tunnel_agent.py' | head -1))"
306
+ if grep -q "announced to hub" "$TUNNEL_LOG" 2>/dev/null; then
307
+ echo "tunnel agent announced to the hub successfully"
308
+ ...
309
+ ```
310
+
311
+ (`deploy/codespace/launch.sh:177-191`)
312
+
313
+ The verification exists because of a specific failure:
314
+
315
+ > *"Backgrounding with all output discarded means a crashing agent is completely invisible — that is
316
+ > exactly how a missing `httpx` hid itself."* (`deploy/codespace/launch.sh:174-176`)
317
+
318
+ ### Step 3 — warm the HF cache, if the Codespace was rebuilt
319
+
320
+ `warm_cache.py` pre-downloads the four pinned models:
321
+
322
+ ```bash
323
+ python deploy/codespace/warm_cache.py
324
+ ```
325
+
326
+ It reports per-model `OK` / `SKIPPED` / `FAILED`, never aborts on a single miss, and **always exits 0** so
327
+ a cache miss cannot fail a build (`deploy/codespace/warm_cache.py:8-10,110-112`). The four models and
328
+ their pinned revisions are transcribed verbatim from `configs/base.yaml`
329
+ (`deploy/codespace/warm_cache.py:12-16,35-60`):
330
+
331
+ | key | repo | revision | file |
332
+ |---|---|---|---|
333
+ | `vlm` | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | (snapshot) |
334
+ | `grounding` | `chendelong/RemoteCLIP` | `bf1d8a3ccf2d` | `RemoteCLIP-ViT-B-32.pt` |
335
+ | `router` | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | (snapshot) |
336
+ | `croma` | `antofuller/CROMA` | `0dd28e3d633b` | `CROMA_base.pt` |
337
+
338
+ It is idempotent and safe to re-run; a second run is a no-op because the blob is already on disk
339
+ (`deploy/codespace/warm_cache.py:3-6`). Two environment switches skip work:
340
+ `SATQUERY_WARM_OFFLINE` / `HF_HUB_OFFLINE` (skip everything) and `SATQUERY_WARM_SKIP="croma,grounding"`
341
+ (skip named models) (`deploy/codespace/warm_cache.py:23-25`).
342
+
343
+ > **Note.** `warm_cache.py` warms four models. The **fifth** pinned backbone in the specialist stack —
344
+ `STANet`/change — is a local trained head, not a Hub backbone, and is not in the warm list. If the change
345
+ head is absent the capability degrades honestly rather than failing the warm step.
346
+
347
+ ### Step 4 — run one throwaway analysis
348
+
349
+ A single cheap query confirms the whole chain end to end. Do this **before** the demo, not during it.
350
+ Watch for: a `run_id`, a `mock_nodes` count of 0, and an answer carrying a `[task]` tag
351
+ (`release/repo/docs/REPRODUCIBILITY.md` §6.3).
352
+
353
+ ### The warm-state checklist
354
+
355
+ | Check | Command | Expected |
356
+ |---|---|---|
357
+ | orchestrator up | `curl --noproxy '*' .../api/health` | `status: ok`, `service: satquery-orchestrator` |
358
+ | tunnel connected | same payload | `tunnel.agent_connected: true` |
359
+ | capabilities | `curl --noproxy '*' .../api/capabilities` | six tasks, all `available: true` |
360
+ | one live run | drive the Analyze console | a `run_id`, `mock_nodes: 0` |
361
+
362
+ ## 7. Restart after an idle-stop
363
+
364
+ A Codespace stops after an idle period. When it stops, the tunnel agent stops polling, and
365
+ `GET /api/health` reports `tunnel.agent_connected: false`
366
+ (`docs/DEPLOYMENT_TOPOLOGY.md` §2).
367
+
368
+ **The reconnect is automatic on start, because `postStartCommand` runs `launch.sh`.** The chain is:
369
+
370
+ ```
371
+ Codespace start
372
+ → devcontainer postStartCommand: bash deploy/codespace/launch.sh (.devcontainer/devcontainer.json:18)
373
+ → launch.sh: start serve.py on $PORT (if not already current) (launch.sh:110-147)
374
+ → launch.sh: start the supervised tunnel agent -> $SATQUERY_HUB_URL (launch.sh:149-169)
375
+ → launch.sh: sleep 4, verify the agent process and the announce line (launch.sh:171-191)
376
+ → the agent dials POST /tunnel/agent and long-polls (launch.sh:150-156)
377
+ → GET /api/health: tunnel.agent_connected becomes true
378
+ ```
379
+
380
+ The hub URL the agent dials is `SATQUERY_HUB_URL`, defaulting to
381
+ `https://satquery-backend-m4yv.onrender.com` (`deploy/codespace/launch.sh:59-61`).
382
+
383
+ **What the operator does:**
384
+
385
+ 1. Start the Codespace (or let the wake path start it — `ensure_codespace_up()` calls the GitHub
386
+ Codespaces `POST .../start` API when `state != "available"`, `deploy/render/main.py:299-356`).
387
+ 2. Wait for `postStartCommand` to run.
388
+ 3. Re-read `GET /api/health` and confirm `tunnel.agent_connected: true` (§6 step 1).
389
+
390
+ **Two things that make a restart go wrong, and their mitigation:**
391
+
392
+ | Failure | Symptom | Mitigation in `launch.sh` |
393
+ |---|---|---|
394
+ | The agent is launched but immediately reaped by the lifecycle shell | agent announces once, then vanishes; hub reports `agent_connected: false`; `/api/infer` falls back to the dead forward path (401 → wake_timeout) | `setsid + nohup + </dev/null` plus a supervising restart loop (`launch.sh:14-24,160-168`) |
395
+ | A **stale** serve process keeps answering from OLD code | `/v1/health` and `/v1/capabilities` answer, but from the previous revision's capabilities | a **stamp** recording the revision + asset config; a mismatch restarts the server (`launch.sh:100-147`) |
396
+
397
+ > *"A stale serve process is worse than no process: it answers /v1/health and /v1/capabilities from OLD
398
+ > code, so the deployment looks alive while reporting the previous revision's capabilities."*
399
+ > (`deploy/codespace/launch.sh:111-113`)
400
+
401
+ The stamp is the closest thing in the system to a deployment-identity check:
402
+
403
+ ```bash
404
+ _current_stamp() {
405
+ printf 'rev=%s asset_enabled=%s asset_dir=%s\n' \
406
+ "$(git rev-parse HEAD 2>/dev/null || echo nogit)" \
407
+ "${SATQUERY_ASSET_ENABLED:-}" \
408
+ "${SATQUERY_ASSET_DIR:-}"
409
+ }
410
+ ```
411
+
412
+ (`deploy/codespace/launch.sh:103-108`)
413
+
414
+ It is **local to the Codespace** and is not exposed on any HTTP route — an operator on the orchestrator
415
+ side cannot see it (`docs/architecture/10-observability-and-ops.md` §7.5).
416
+
417
+ ### The preflight that refuses a half-configured start
418
+
419
+ `launch.sh` refuses to start if the Python dependencies or the `app` package cannot be imported
420
+ (`deploy/codespace/launch.sh:70-91`). The dependency check is explicit about `httpx`, because a missing
421
+ `httpx` once made the agent die instantly and the supervised loop hid the error in a log file
422
+ (`deploy/codespace/launch.sh:76-79`):
423
+
424
+ ```bash
425
+ if ! python -c "import yaml, pydantic, fastapi, uvicorn, httpx" 2>/dev/null; then
426
+ echo "ERROR: Python deps are missing (need yaml, pydantic, fastapi, uvicorn, httpx)." >&2
427
+ ...
428
+ exit 1
429
+ fi
430
+ ```
431
+
432
+ > **Caveat — the referenced repair tool is not in this tree.** `launch.sh` points the operator at
433
+ > `bash deploy/codespace/doctor.sh --install` (`launch.sh:83,89,182`), but **`doctor.sh` does not exist in
434
+ > the monorepo working copy**, and neither does `tunnel_agent.py` (§5). Whether `doctor.sh` exists in the
435
+ > deployed `SatQuery-Inference` repository is `UNKNOWN — not established from the available evidence`.
436
+
437
+ ## 8. Tell a tunnel gap from a real failure
438
+
439
+ This is the runbook's most useful procedure, and it is grounded in a measured timing.
440
+
441
+ ### The signature
442
+
443
+ A request that hangs for **≈249 seconds** and then returns **`504`** is the tunnel-gap signature, not a
444
+ broken model. The arithmetic is exact:
445
+
446
+ ```
447
+ tunnel_timeout_s (150) + wake_timeout_s (120) = 270 s (nominal)
448
+ measured ≈ 249 s
449
+ ```
450
+
451
+ > *"B-07 root shape: in `auto` mode a tunnel timeout **falls through** to the forward path
452
+ > (`main.py:546`), burning `wake_timeout_s = 120` on a `302` (~249 s ≈ 150 + 120)."*
453
+ > (`release/CURRENT_RELEASE_STATE.md` §6; `release/repo/docs/DEPLOYMENT.md` §8.1)
454
+
455
+ ### The three signals to read, in order
456
+
457
+ 1. **`tunnel.agent_connected` on a fresh `/api/health`.** If `false`, it is a tunnel gap and the remedy is
458
+ §6 step 2. Do not trust a single reading — the flag is a freshness window, so re-read.
459
+ 2. **The elapsed time.** ≈249 s is the B-07 signature. A fast failure is something else.
460
+ 3. **The machine code.** The codes are disjoint and each implies a different action (§22).
461
+
462
+ ### The decision tree
463
+
464
+ ```mermaid
465
+ flowchart TD
466
+ A["a request hung, or returned 5xx"] --> B{"re-read /api/health<br/>tunnel.agent_connected?"}
467
+ B -->|"true"| C{"was the wait ≈249 s?"}
468
+ B -->|"false"| D["TUNNEL GAP<br/>the agent is not polling.<br/>Start the Codespace (§6)."]
469
+ C -->|"yes, 504"| E["TUNNEL GAP<br/>agent went stale mid-request,<br/>or the Codespace stopped.<br/>B-07. Mitigate operationally."]
470
+ C -->|"no"| F{"what was the code?"}
471
+ F -->|"invalid_request 422"| G["a CLIENT bug — the body<br/>did not match AnalysisRequest"]
472
+ F -->|"model_load_error / model_unavailable"| H["an ARTIFACT defect —<br/>surface it, do not retry blindly"]
473
+ F -->|"upstream_timeout 504"| I["tunnel healthy but slow —<br/>the agent did not complete in 150 s"]
474
+ F -->|"other"| J["read the code and<br/>trace.errors[]"]
475
+ ```
476
+
477
+ > **Note which branches exist only in the patch.** `forward_unavailable` and `upstream_timeout` do **not**
478
+ > exist in the deployed revision (§13). An operator on the deployed system will not see them; they will
479
+ > see `wake_timeout` after ≈249 s instead.
480
+
481
+ ## 9. What the client shows while waking
482
+
483
+ The frontend is not silent during a cold start. It shows *"Waking inference engine…"* while Render starts
484
+ the Codespace (`docs/DEPLOYMENT_TOPOLOGY.md` §2, §2 mermaid; `release/repo/docs/DEPLOYMENT.md` §8).
485
+
486
+ The orchestrator's side of this is a response header. The wake path is **blocking wake-then-proxy**: the
487
+ client waits and receives the result, and the response is tagged:
488
+
489
+ ```python
490
+ out.headers["X-SatQuery-State"] = "waking" if woke else "ready"
491
+ ```
492
+
493
+ (`deploy/render/main.py:486-489`)
494
+
495
+ The module docstring states the intent plainly:
496
+
497
+ > *"The client simply **waits** (blocking, wake-then-proxy) and receives the result; the frontend
498
+ > independently shows 'Waking inference engine...' on slow responses. We optionally tag the response with
499
+ > `X-SatQuery-State: waking` so the frontend can confirm the delay was a cold start."*
500
+ > (`deploy/render/main.py:13-20`)
501
+
502
+ So an operator watching a demo sees: the client's "Waking inference engine…" message, a wait, and then
503
+ either a result or a `504`. **A `504` after ≈249 s is the B-07 shape, not a broken model** (§8).
504
+
505
+ ---
506
+
507
+ # Part III — Cold start and the timing budget
508
+
509
+ ## 10. The cold-start shape
510
+
511
+ The cold start is **blocking and documented**, not hidden. The shape an operator should expect
512
+ (`docs/architecture/10-observability-and-ops.md` §7.2):
513
+
514
+ | Phase | What happens | Bound |
515
+ |---|---|---|
516
+ | 0 | the client sends `POST /api/infer` and **waits** | — |
517
+ | 1 | `GET` the Codespace via the GitHub API; `POST .../start` if not `available` | one API round trip |
518
+ | 2 | poll `GET {base}/v1/health` every 2 s, each with a 10 s timeout | ≤ `wake_timeout_s = 120` |
519
+ | 3 | on the first `200`, proxy the original request | ≤ `upstream_timeout_s = 90` |
520
+ | 4 | the response carries `X-SatQuery-State: waking` | — |
521
+
522
+ The poll interval and per-poll timeout are module constants:
523
+
524
+ ```python
525
+ # Polling knobs for the wake loop.
526
+ _WAKE_POLL_INTERVAL_S = 2.0
527
+ _WAKE_HEALTH_TIMEOUT_S = 10.0
528
+ ```
529
+
530
+ (`deploy/render/main.py:76-78`)
531
+
532
+ > **Do not quote a single cold-start number.** The documented statement is *"tens of seconds"*
533
+ > (`release/repo/docs/DEPLOYMENT.md` §8), and the bounded worst case is `wake_timeout_s = 120` before a
534
+ > `504`. **No measured cold-start distribution exists:**
535
+ > `UNKNOWN — not established from the available evidence`
536
+ > (`docs/architecture/10-observability-and-ops.md` §7.2).
537
+
538
+ ## 11. The four timeouts, and why the relationship matters
539
+
540
+ The live payload reports three of them; the fourth comes from the frozen config.
541
+
542
+ | Timeout | Value | Where it lives | What it bounds |
543
+ |---|---|---|---|
544
+ | `tunnel_timeout_s` | **150.0 s** | live config (`release/repo/docs/DEPLOYMENT.md` §5) | how long a request waits on the tunnel before falling through |
545
+ | `wake_timeout_s` | **120.0 s** | live config | how long the wake loop waits for `/v1/health` |
546
+ | `upstream_timeout_s` | **90.0 s** | live config | the gateway → upstream proxy budget |
547
+ | `agent.timeout_seconds` | **120 s** | `configs/base.yaml` (`agent.timeout_seconds: 120`) | the Space's own total request budget |
548
+
549
+ The **relationship** between the last two is the one that must not be inverted
550
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.2):
551
+
552
+ ```
553
+ gateway upstream timeout < agent.timeout_seconds ≤ the Space's own request budget
554
+ 90 s < 120 s
555
+ ```
556
+
557
+ Both failure directions are documented:
558
+
559
+ - **Gateway timeout too short** — it kills a legitimately running `grounding` or `optical_sar` call and
560
+ reports it as an upstream failure. The Space's error, not the client's.
561
+ - **Gateway timeout too long** — it holds a connection past the point the Space itself has given up,
562
+ converting a clean upstream timeout into a client-side hang.
563
+
564
+ > The lower bound of the *historical* window was 45 s (the longest single `gpu_duration_*`). On the live
565
+ > **CPU** deployment the ZeroGPU durations are frozen paperwork (§19), so the binding upper constraint is
566
+ > the 120 s agent timeout and the live gateway value is 90 s.
567
+
568
+ ## 12. The worst case: ≈249 s under B-07
569
+
570
+ `SATQUERY_TRANSPORT=auto` means **try the tunnel; on timeout, fall through to the forward path**
571
+ (`release/repo/docs/DEPLOYMENT.md` §8.1). The forward path to a private repository returns `302` quickly,
572
+ but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S` (120 s) first. So a worst-case failed request
573
+ takes roughly:
574
+
575
+ ```
576
+ 150 s (tunnel timeout) + 120 s (wake timeout on a 302) ≈ 249 s
577
+ ```
578
+
579
+ This is the **root shape** of the observed transient tunnel gap, and it is why a request can appear to
580
+ hang and then fail (`release/repo/docs/DEPLOYMENT.md` §8.1). It is `OPEN` (§13).
581
+
582
+ ---
583
+
584
+ # Part IV — The known defects, in operational terms
585
+
586
+ ## 13. B-07 — transient tunnel gaps (`OPEN`)
587
+
588
+ **B-07 is `OPEN`.** The patch is prepared and **not deployed**. This is the single most important
589
+ operational fact in this manual.
590
+
591
+ > *"B-07 | Transient tunnel-agent gaps → a request can hang or return 504. Patch prepared, **NOT
592
+ > deployed**. | **OPEN**"* (`release/CURRENT_RELEASE_STATE.md` §6)
593
+
594
+ ### What an operator experiences
595
+
596
+ - The tunnel agent is briefly absent (a restart, a reap, a gap).
597
+ - A request issued during the gap either hangs or returns `504`.
598
+ - In `auto` mode the hang lasts up to ≈249 s before the `504` (§12).
599
+ - Once the agent reconnects, the next request succeeds.
600
+
601
+ ### The root cause, exactly
602
+
603
+ In `auto` mode a tunnel timeout **falls through to the forward path** (`main.py:546`), and the forward
604
+ path to a private repository returns `302`. The wake step burns `wake_timeout_s = 120` on that `302`
605
+ before the request fails (`release/CURRENT_RELEASE_STATE.md` §6).
606
+
607
+ ### What the patch does
608
+
609
+ The patch was authored and verified (`py_compile` clean, applies cleanly to the deployed `main.py`)
610
+ (`release/repo/docs/DEPLOYMENT.md` §8.1). It makes three changes:
611
+
612
+ | Change | Code | Effect |
613
+ |---|---|---|
614
+ | A | `forward_unavailable` (`503`, `recoverable: true`) on a **terminal** `302`/`401`/`403` | converts a 504-after-249 s into a 503-early with an actionable code |
615
+ | B | `upstream_timeout` (`504`) for "tunnel healthy but slow" | distinguishes a slow agent from a dead forward path |
616
+ | C | `/api/health` `codespace_name` `.strip()` | fixes B-02 |
617
+
618
+ ### The honesty note attached to the patch
619
+
620
+ > *"The report records that an earlier claim that the patch 'would not have prevented' the observed 504
621
+ > 'was wrong and was retracted'. The corrected position: 'Change A is genuinely **on the failing path** —
622
+ > it converts a 504-after-249 s into a 503-early with an actionable code.'"*
623
+ > (`release/CURRENT_RELEASE_STATE.md` §8 evidence list; `docs/architecture/02-deployment-topology.md` §6.4)
624
+
625
+ Do not repeat the retracted version.
626
+
627
+ ### Why it is not deployed
628
+
629
+ > *"the patch is not needed for the demo and touches the live backend. The residual is better mitigated
630
+ > operationally (keep the Codespace warm, raise the idle timeout)."*
631
+ > (`DELIVERY_REPORT_2026-09-25.md` §4, quoted in `docs/architecture/10-observability-and-ops.md` §7.4)
632
+
633
+ ### The operational mitigation (this is what an operator actually does)
634
+
635
+ 1. Keep the Codespace **warm** before and during any demo (§6).
636
+ 2. **Raise the Codespace idle timeout** so it does not stop mid-session.
637
+ 3. Re-read `/api/health` before a run, and treat `agent_connected: false` as "warm the stack now".
638
+ 4. Expect a ≈249 s hang followed by a `504` if the agent goes stale mid-request — and **retry**, because
639
+ the residual is transient.
640
+
641
+ > **Do not upgrade B-07.** It is `OPEN`. It is not `RESOLVED`, and the patch is not deployed — the live
642
+ > payload's `codespace_name` trailing `\n` is the witness (§14).
643
+
644
+ ## 14. B-02 — the trailing newline (`OPEN`, cosmetic)
645
+
646
+ `GET /api/health` reports the Codespace name with a trailing newline:
647
+
648
+ ```json
649
+ "codespace_name": "potential-space-trout-r4ppw969w45j2pvvw\n"
650
+ ```
651
+
652
+ (`release/repo/docs/DEPLOYMENT.md` §5; `release/CURRENT_RELEASE_STATE.md` §1)
653
+
654
+ **It is cosmetic.** The wake path strips it — `_codespace_name()` calls `.strip()` before using the value
655
+ (`deploy/render/main.py:99-106`) — so only the health payload reports the raw value
656
+ (`release/repo/docs/DEPLOYMENT.md` §5).
657
+
658
+ **It is `OPEN`.** Its presence is also the operational **witness** that the B-07 patch is not deployed:
659
+ change C of that patch is the `.strip()` fix, so a live payload still showing the trailing `\n` proves the
660
+ patch is absent (`docs/architecture/10-observability-and-ops.md` §7.4).
661
+
662
+ > **Do not "fix" B-02 by editing the health payload on the live service.** The fix ships with the B-07
663
+ > patch, which is deliberately not deployed. A cosmetic newline is not worth a live-backend change on
664
+ > its own.
665
+
666
+ ---
667
+
668
+ # Part V — Monitoring and alerting
669
+
670
+ ## 15. What exists
671
+
672
+ SatQuery AI observes **exactly three things** (`docs/architecture/10-observability-and-ops.md` §6): whether
673
+ the transport is up, what one run did, and what failed inside the server.
674
+
675
+ ### 15.1 The health payload
676
+
677
+ The measured live payload (`release/repo/docs/DEPLOYMENT.md` §5):
678
+
679
+ ```json
680
+ {
681
+ "status": "ok",
682
+ "service": "satquery-orchestrator",
683
+ "tunnel": {
684
+ "agent_connected": true,
685
+ "agent_id": "codespaces-fd1038",
686
+ "pending": 0,
687
+ "completed": 97
688
+ },
689
+ "config": {
690
+ "codespace_name": "potential-space-trout-r4ppw969w45j2pvvw\n",
691
+ "codespace_port": 8000,
692
+ "transport_mode": "auto",
693
+ "tunnel_timeout_s": 150.0,
694
+ "wake_timeout_s": 120.0,
695
+ "upstream_timeout_s": 90.0,
696
+ "device": "cpu",
697
+ "has_github_token": true
698
+ }
699
+ }
700
+ ```
701
+
702
+ Field by field, for the operator:
703
+
704
+ | Field | Meaning | Operational use |
705
+ |---|---|---|
706
+ | `status` | orchestrator liveness | a single up/down bit for the gateway tier |
707
+ | `tunnel.agent_connected` | an agent polled within the freshness window | **the** check before a demo (§6) |
708
+ | `tunnel.agent_id` | which agent | identifies the Codespace agent that is connected |
709
+ | `tunnel.pending` | in-flight tunnel requests | a rising value means work is queueing |
710
+ | `tunnel.completed` | a monotonic delivery counter, per Render process | trend only — it **resets on restart** |
711
+ | `config.codespace_name` | the target Codespace | carries B-02's trailing `\n` (§14) |
712
+ | `config.codespace_port` | the inference port | `8000` |
713
+ | `config.transport_mode` | `auto` / `tunnel` / `forward` | `auto` is what makes B-07 reachable (§13) |
714
+ | `config.tunnel_timeout_s` | the tunnel wait | `150.0` (§11) |
715
+ | `config.wake_timeout_s` | the cold-start wait | `120.0` (§11) |
716
+ | `config.upstream_timeout_s` | the proxy budget | `90.0` (§11) |
717
+ | `config.device` | device preference | `cpu` |
718
+ | `config.has_github_token` | whether a token is present (boolean only) | a `false` here means the wake path cannot start the Codespace |
719
+
720
+ > **`completed` was measured at three different values** across probes (`97`, `314`, and others). It is a
721
+ > counter that resets when the Render process restarts, not a constant. **Do not treat any single reading
722
+ > as the value** (`docs/architecture/10-observability-and-ops.md` §2.3;
723
+ > `docs/architecture/02-deployment-topology.md` §7.3).
724
+
725
+ ### 15.2 The per-run `ExecutionTrace`
726
+
727
+ Every run carries an `ExecutionTrace` (`core/schemas.py:296-319`) with: `run_id`, `task`, `query`,
728
+ `inputs`, `modalities`, `intent`, `validation`, `workflow`, `steps`, `selected_models`, `parameters`,
729
+ `outputs`, `confidence`, `timings`, `fallbacks`, `errors`, `contradiction`, `config_hash`, `started_at`,
730
+ `finished_at`.
731
+
732
+ This is the **only** diagnostic surface for a *wrong but successful* answer — no log line records a
733
+ successful run (`docs/architecture/10-observability-and-ops.md` §7.6). The operator reads:
734
+
735
+ | Trace field | What it answers |
736
+ |---|---|
737
+ | `intent` | what the router thought the query meant |
738
+ | `task` | what was dispatched |
739
+ | `selected_models` | what actually ran |
740
+ | `errors` | what failed inside the run |
741
+ | `fallbacks` | what degraded |
742
+ | `config_hash` | which config produced the result |
743
+
744
+ ### 15.3 The Codespace's own `/v1/health`
745
+
746
+ The inference tier answers its own health route, derived rather than asserted, with a torch-free device
747
+ probe (`app/space_app.py:521-547`; `core/schemas.py:430-437`; `docs/architecture/10-observability-and-ops.md`
748
+ §2.10). `gpu_available: false` is **expected** on a CPU host and must never be surfaced as a fault
749
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2).
750
+
751
+ ## 16. What does **not** exist
752
+
753
+ This is the honest inventory. None of it is aspirational — the list exists so that a reader does not
754
+ assume a monitoring facility that was never built (`docs/architecture/10-observability-and-ops.md` §6):
755
+
756
+ | Capability | Present? | Evidence |
757
+ |---|---|---|
758
+ | **APM** (application performance monitoring) | **no** | no APM client, SDK or agent in any source read |
759
+ | **Distributed tracing** | **no** | no trace-context propagation; the three id namespaces do not join |
760
+ | **Cost accounting** | **no** | `GPU_DURATIONS` declares durations but nothing meters or reports consumption |
761
+ | **Metrics endpoint** (Prometheus / OpenMetrics) | **no** | no `/metrics` route in any app factory |
762
+ | **Per-model latency histogram** | **no** | `trace.timings` is per-run, per-step, never aggregated |
763
+ | **Error-rate counter** | **no** | no counter exists; `trace.errors` is per-run only |
764
+ | **Request counter** | **no** | the tunnel's `completed` counts only tunnel deliveries, per Render process |
765
+ | **Uptime / restart tracking** | **no** | no uptime field; the hub's counters reset on restart |
766
+ | **Alerting** | **no** | no alerting rule, webhook or threshold anywhere |
767
+ | **Structured / JSON logs** | **no** | all log calls use `%s`-style free text |
768
+ | **Log shipping / aggregation** | **no** | logs are per-host; the Codespace's are on an ephemeral filesystem |
769
+ | **Dashboards** | **no** | none exists |
770
+ | **SLO / SLA definition** | **no** | none exists |
771
+ | **A system-level end-to-end benchmark** | **no** | `DOCS_STYLE_GUIDE.md` §3: *"does not exist; no system-level accuracy is claimed"* |
772
+
773
+ **There is no pager, no alert, and no dashboard.** An operator learns the system is down by trying to use
774
+ it. This is stated as a fact, not a complaint.
775
+
776
+ ## 17. Why "no alerting" is a design fact, not an oversight
777
+
778
+ The plan's §74 lists what is **deliberately not included**, and the monitoring gaps are downstream of
779
+ that list:
780
+
781
+ ```
782
+ authentication
783
+ multi-tenant isolation
784
+ distributed queues
785
+ autoscaling
786
+ observability platform
787
+ Kubernetes
788
+ service mesh
789
+ distributed storage
790
+ horizontal worker orchestration
791
+ enterprise security
792
+ billing
793
+ SLA infrastructure
794
+ ```
795
+
796
+ (`Implementation and Architecture plan.md` §74)
797
+
798
+ The architecture is described there as **scale-compatible, but not a production implementation**. The
799
+ absence of an observability platform, billing and SLA infrastructure is therefore intentional at this
800
+ stage. An operator should not expect — and must not claim — production-grade monitoring.
801
+
802
+ ---
803
+
804
+ # Part VI — Capacity and cost
805
+
806
+ ## 18. The capacity shape
807
+
808
+ Three facts bound capacity, and none of them is elastic:
809
+
810
+ | Property | Value | Evidence |
811
+ |---|---|---|
812
+ | Render plan | **free tier** — sleeps when idle | `render.yaml` (`plan: free`); `release/repo/docs/DEPLOYMENT.md` §8 |
813
+ | Inference hosts | **one** Codespace | `release/CURRENT_RELEASE_STATE.md` §1 |
814
+ | Device | **CPU-only** | `render.yaml`; `.devcontainer/devcontainer.json`; `docs/DEPLOYMENT_DECISION.md` §5 |
815
+ | Autoscaling | **absent** | plan §74 lists `autoscaling` under **Not included** |
816
+ | Horizontal workers | **absent** | plan §74 lists `horizontal worker orchestration` under **Not included** |
817
+ | Database / queue | **absent** | the gateway has *"no database, no auth, no queue"* (`deploy/render/main.py:4-6`) |
818
+
819
+ The consequence for an operator:
820
+
821
+ - **A single client can occupy the system.** The per-IP rate limit is a fairness control, **not** a
822
+ security control (§19.2). There is no queue to absorb a burst.
823
+ - **A restart is a full outage.** There is no replica to fail over to. The Codespace filesystem is
824
+ ephemeral (`deploy/codespace/launch.sh:44-47`), so a restart also empties the asset store.
825
+ - **Cold starts are unavoidable.** Render's free tier sleeps, so the first request after idle pays the
826
+ cold-start cost (§10).
827
+
828
+ ## 19. The cost shape, and the one absence that has a code artifact
829
+
830
+ ### 19.1 What is declared
831
+
832
+ `app/space_app.py` declares a per-task ZeroGPU duration, transcribed from the frozen config
833
+ (`app/space_app.py:105-116`, quoted in `docs/architecture/10-observability-and-ops.md` §6.1):
834
+
835
+ ```python
836
+ GPU_DURATIONS: dict[str, int] = {
837
+ "vqa": 20,
838
+ "caption": 20,
839
+ "grounding": 45,
840
+ "change": 30,
841
+ "optical_sar": 45,
842
+ "change_vqa": 30,
843
+ }
844
+ ```
845
+
846
+ These values are used **only** to decorate a handler with a ZeroGPU reservation. On the live **CPU**
847
+ deployment the decoration is a **no-op**: `_spaces_module()` returns `None` when the `spaces` package is
848
+ absent, so `decorate_gpu` returns the identity decorator (`app/space_app.py:143-166`).
849
+
850
+ > **Do not present `GPU_DURATIONS` as a cost model.** It is a declaration of intended reservation, and on
851
+ > the CPU deployment it reserves nothing (`docs/architecture/10-observability-and-ops.md` §6.1). The
852
+ > ZeroGPU 5-GPU-minute/day quota and the `@spaces.GPU` decoration are **frozen paperwork** — no Gradio
853
+ > runtime exists in code, and the manifest is left undisturbed because editing it would move the config
854
+ > hash (`release/repo/docs/DEPLOYMENT.md` §11).
855
+
856
+ ### 19.2 What is **not** metered
857
+
858
+ Nothing reads `GPU_DURATIONS` back out to compute, record or report consumption. There is no per-run
859
+ GPU-second field, no cumulative counter, and no budget-exhaustion signal
860
+ (`docs/architecture/10-observability-and-ops.md` §6.1).
861
+
862
+ **No cost is observed.** There is no cost accounting for:
863
+
864
+ | Cost | Metered? | Note |
865
+ |---|---|---|
866
+ | Render compute | **no** | free tier; no usage field is read |
867
+ | Codespace compute | **no** | no usage field is read; the quota is a platform property |
868
+ | Hugging Face model hosting | **no** | no usage field is read |
869
+ | Model download volume | **no** | `warm_cache.py` reports per-model status but not bytes or cost |
870
+ | Per-request inference cost | **no** | no field exists |
871
+
872
+ The rate-limit and size-limit values are the implementation's choices, not the plan's, and the maintainer
873
+ should confirm them because they bound one client's share of capacity
874
+ (`docs/PHASE19_FINAL_HARDENING.md` §4.4).
875
+
876
+ ### 19.3 The rate limit is fairness, not protection
877
+
878
+ The limiter keys on the first hop of `X-Forwarded-For`, which is **client-supplied**. A caller that varies
879
+ the header is never throttled. Measured in-process at 3 requests / 60 s, 8 requests sent: `5/8` throttled
880
+ without the header, **`0/8` with a fresh value per request** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2).
881
+
882
+ > **Do not treat `SATQUERY_RATE_LIMIT_PER_IP` as protecting capacity.** It bounds accidental loops and
883
+ > honest clients. Fixing it correctly depends on how many proxy hops Render inserts, which must be
884
+ > measured on a deployed gateway — hard-coding a guess would replace a documented weakness with an
885
+ > undocumented one (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2).
886
+
887
+ ---
888
+
889
+ # Part VII — Incident triage
890
+
891
+ ## 20. The symptom → cause → check → action table
892
+
893
+ One table, ordered by how often each symptom is seen. "First check" is the single cheapest command that
894
+ distinguishes the cases.
895
+
896
+ | Symptom | Likely cause | First check | Action |
897
+ |---|---|---|---|
898
+ | Request hangs ~249 s, then `504` | **B-07 tunnel gap** | re-read `tunnel.agent_connected` | retry; keep the Codespace warm (§13) |
899
+ | Request returns `503` quickly, `recoverable: true` | no agent connected and `mode == tunnel` | `GET /api/health` | start the Codespace (§7) |
900
+ | `GET /api/health` unreachable | Render sleeping or down | re-issue the request | the first request wakes Render; retry |
901
+ | `GET /api/health` OK but `agent_connected: false` | Codespace stopped, or the agent was reaped | — | start the Codespace; confirm `postStartCommand` ran (§7) |
902
+ | `504 wake_timeout` | the Codespace did not come up within 120 s | `tunnel.agent_connected` | the Codespace is not coming up; check it directly |
903
+ | `502 upstream_unreachable` | a connection error to the Codespace | `tunnel.agent_connected` | check the transport |
904
+ | `422 invalid_request` | the body did not match `AnalysisRequest` | the client's request body | fix the client; **not** a server fault |
905
+ | `503 model_unavailable` | an artifact is absent | `GET /api/capabilities` reasons | expected if not uploaded; ship degraded or upload it |
906
+ | `503 model_load_error` | an artifact is present but corrupt | the capability's `reason` | **a defect** — replace the artifact and report |
907
+ | `confidence.method: "uncalibrated"` | the calibration artifact was not uploaded | the run's trace | honest, not broken |
908
+ | `status: "degraded"` on `/v1/health` | at least one capability is not servable | `GET /api/capabilities` | not an error; the service is up |
909
+ | `gpu_available: false` | CPU host | — | expected; never surface as a fault |
910
+ | a handle that worked now returns `400 input_error` | the handle lapsed, or the Codespace restarted and its asset dir is ephemeral | re-upload | do not assume handles persist |
911
+ | a capability `available: true` **with** a non-null `reason` | hub-backed, no local checkpoint — the first call will be slow | the reason string | not a fault; do not surface as an error |
912
+ | a **wrong but successful** answer | a router/dispatch/artifact issue | the run's `ExecutionTrace` | read `intent`, `task`, `selected_models`, `errors`, `fallbacks` (§15.2) |
913
+
914
+ The first four rows are the ones an operator hits in practice. Rows 8–15 are from
915
+ `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2, which is the canonical degraded-state reference.
916
+
917
+ ## 21. The consolidated decision tree
918
+
919
+ For the case where something is wrong and the operator has only the client-side report
920
+ (`docs/architecture/10-observability-and-ops.md` §7.6):
921
+
922
+ ```mermaid
923
+ flowchart TD
924
+ A["something is wrong"] --> B["GET /api/health"]
925
+ B --> C{"tunnel.agent_connected?"}
926
+ C -->|"false"| D["warm the stack<br/>§6 steps 1-2"]
927
+ C -->|"true"| E{"what did the client see?"}
928
+
929
+ E -->|"nothing, request hung"| F{"waited ≈249 s?"}
930
+ F -->|"yes"| G["B-07 tunnel gap<br/>retry + warm<br/>§13"]
931
+ F -->|"no"| H["still waiting —<br/>within budget"]
932
+
933
+ E -->|"a 5xx"| I["read the error code<br/>§22"]
934
+ E -->|"a 4xx"| J["a client bug —<br/>the body did not match the contract"]
935
+
936
+ I --> K{"code present in the<br/>DEPLOYED revision?"}
937
+ K -->|"no"| L["the patch is not deployed<br/>§13"]
938
+ K -->|"yes"| M["act on the code"]
939
+
940
+ E -->|"a result, but wrong"| N["read trace:<br/>intent, task, selected_models,<br/>errors, fallbacks<br/>§15.2"]
941
+ ```
942
+
943
+ ## 22. The machine-code reference
944
+
945
+ | Code | Status | Meaning | Action |
946
+ |---|---|---|---|
947
+ | `tunnel_offline` | 503 | no agent is connected and `mode == "tunnel"` | start the Codespace |
948
+ | `forward_unavailable` | 503 | the forwarded port is not anonymously reachable | start the tunnel agent (**patch only**) |
949
+ | `wake_timeout` | 504 | the wake loop exhausted `wake_timeout_s` | the Codespace is not coming up |
950
+ | `upstream_timeout` | 504 | the tunnel was healthy but no agent completed in `tunnel_timeout_s` | retry; check the agent (**patch only**) |
951
+ | `upstream_unreachable` | 502 | a connection error to the Codespace | check the transport |
952
+ | `invalid_request` | 422 | the body was not a valid `AnalysisRequest` | fix the client |
953
+ | `model_unavailable` | 502 | the gateway could not reach the analysis service | retry |
954
+ | `orchestrator_config_error` | 500 | missing `GITHUB_TOKEN` or `CODESPACE_NAME` | fix the env vars (`deploy/render/main.py:256-262`) |
955
+ | `upstream_error` | 502 | a non-connection httpx error | check the upstream (`deploy/render/main.py:392-400`) |
956
+ | `schema_validation_error` | 502 | a non-JSON upstream body | check the upstream (`deploy/render/main.py:404-414`) |
957
+
958
+ (`deploy/render/main.py:247-291,383-414`; the full taxonomy is in
959
+ [08 — The API Contract](architecture/08-api-contract.md) §12.)
960
+
961
+ > **The two codes marked "patch only" do not exist in the deployed revision.** An operator on the
962
+ > deployed system will not see `forward_unavailable` or `upstream_timeout`; they will see `wake_timeout`
963
+ > after ≈249 s instead (§13).
964
+
965
+ ---
966
+
967
+ # Part VIII — Routine maintenance
968
+
969
+ ## 23. Rotating credentials (procedure only)
970
+
971
+ > **This document records no credential value, and no path to a credential file.** The repository's own
972
+ > evidence index records where credentials are held (`release/CURRENT_RELEASE_STATE.md` §7) by
973
+ > **location and kind only**; that index is not reproduced here. The procedure below is deliberately
974
+ > value-free.
975
+
976
+ Four credential purposes exist, and each is rotated by changing a value the operator holds — never a
977
+ value in this document:
978
+
979
+ | Purpose | Where it is consumed | What rotation changes |
980
+ |---|---|---|
981
+ | Codespace control (wake) | the Render env var `GITHUB_TOKEN` | the token the orchestrator uses to `GET`/`POST .../start` a Codespace (`deploy/render/codespaces.py:65-70`) |
982
+ | Repository writes | the operator's local tooling | the token used for the Git Data API deploy path (`release/repo/docs/DEPLOYMENT.md` §7.1) |
983
+ | Codespace account access | the GitHub account session | the account credential used to open the Codespace |
984
+ | Hugging Face model access | the Hub download path | the token used to resolve pinned model revisions |
985
+
986
+ **The procedure, in order:**
987
+
988
+ 1. **Create the replacement credential** in the provider's UI, with the minimum scope the purpose needs.
989
+ For the wake path the scope is `codespace` (`deploy/render/codespaces.py:50`, `:90-92`).
990
+ 2. **Update the consumer's configuration.** For the wake path this is the Render environment variable
991
+ `GITHUB_TOKEN` (`render.yaml:14-15`; `release/repo/docs/DEPLOYMENT.md` §6.1).
992
+ 3. **Redeploy / restart the consumer** so it reads the new value. The orchestrator reads its environment
993
+ once at process start, so a value change without a restart is invisible.
994
+ 4. **Verify.** Call `GET /api/health` and confirm `config.has_github_token: true`
995
+ (`deploy/render/main.py:454`). Then trigger one wake to confirm the token is accepted end to end.
996
+ 5. **Revoke the old credential** at the provider, once step 4 passes.
997
+
998
+ **Two rules that must not be broken:**
999
+
1000
+ - **Never set a token variable to an empty string.** An empty `HF_TOKEN` produced
1001
+ `Authorization: Bearer `, which httpx rejects with a `LocalProtocolError` that is misreported as an
1002
+ upstream failure. **Omit the variable entirely rather than setting it empty**
1003
+ (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1). The same reasoning applies to `GITHUB_TOKEN`: the
1004
+ orchestrator raises a clear `orchestrator_config_error` when it is absent
1005
+ (`deploy/render/main.py:90-96`), which is better than a confusing transport failure.
1006
+ - **Never commit a credential.** The Render variables are `sync: false` in `render.yaml` and must be set
1007
+ in the dashboard (`deploy/render/README.md` §Environment variables).
1008
+
1009
+ > **`has_github_token` is a boolean, never the value.** The health payload reports only whether a token is
1010
+ > present (`deploy/render/main.py:454`), so a health probe is a safe way to confirm rotation without
1011
+ > exposing the credential.
1012
+
1013
+ ## 24. Restarting the tunnel agent
1014
+
1015
+ The agent is supervised and should not need a manual restart, but the procedure is:
1016
+
1017
+ 1. **Check whether it is running** — the Codespace-side check is
1018
+ `pgrep -f "deploy/codespace/tunnel_agent.py"` (`deploy/codespace/launch.sh:153,179`).
1019
+ 2. **Re-run the launcher** — `bash deploy/codespace/launch.sh`. It is **guarded** on the process, so
1020
+ re-running is a no-op if the agent is already up (`deploy/codespace/launch.sh:24,153-155`).
1021
+ 3. **Confirm the reconnect** — read the agent log for the announce line
1022
+ (`grep -q "announced to hub" "$TUNNEL_LOG"`, `deploy/codespace/launch.sh:185`) and then confirm
1023
+ `tunnel.agent_connected: true` from outside (§6 step 1).
1024
+
1025
+ > **The launcher restarts the serve process too, if it is stale** (`deploy/codespace/launch.sh:100-147`).
1026
+ > That is intentional: a stale server answering from old code is worse than a restart (§7).
1027
+
1028
+ > **Caveat.** The agent log path is `/tmp/satquery-tunnel.log` and the serve log is
1029
+ > `/tmp/satquery-serve.log` (`deploy/codespace/launch.sh:63-64`). Both are on the **ephemeral** Codespace
1030
+ > filesystem, so they vanish with the Codespace (`deploy/codespace/launch.sh:44-47`).
1031
+
1032
+ ## 25. Re-warming the model cache
1033
+
1034
+ Re-warm after a Codespace rebuild, or whenever the first analysis is slower than expected:
1035
+
1036
+ ```bash
1037
+ python deploy/codespace/warm_cache.py
1038
+ ```
1039
+
1040
+ `post_create.sh` runs this automatically on container **creation** (`deploy/codespace/post_create.sh:5-6`),
1041
+ and it is wired as the devcontainer `postCreateCommand` (`.devcontainer/devcontainer.json:17`).
1042
+
1043
+ > **Creation vs. start.** `postCreateCommand` runs only when the container is **created**, whereas
1044
+ > `postStartCommand` runs on every start (`.devcontainer/devcontainer.json:17-18`). The `containerEnv`
1045
+ > block is likewise applied only at creation, which is why `launch.sh` re-exports the asset-upload
1046
+ > variables on every start — *"`containerEnv` is only applied when the container is CREATED"*
1047
+ > (`deploy/codespace/launch.sh:49-52`). An operator who changes an env var must restart, not just reload.
1048
+
1049
+ ## 26. Deploying a change (the Git Data API path)
1050
+
1051
+ Deployment does **not** use `git push`. Every deployed file is uploaded as a **blob** whose sha256 is
1052
+ computed locally and verified against the uploaded blob, then assembled into a tree, committed, and the
1053
+ branch ref patched (`release/repo/docs/DEPLOYMENT.md` §7.1).
1054
+
1055
+ Why this matters operationally:
1056
+
1057
+ - each file is **content-verified** rather than trusted;
1058
+ - deletions are expressed explicitly as `sha: null` tree entries;
1059
+ - the deploy is **idempotent** — re-running it with identical content produces no change.
1060
+
1061
+ **Measured:** 9 deployed files were re-read from the API and found **sha256 byte-identical** to the local
1062
+ copies, with the deployed HEAD re-read independently (`verify_deployed_head.py`,
1063
+ `release/CURRENT_RELEASE_STATE.md` §5).
1064
+
1065
+ **Frontend deploy** (Cloudflare Pages) uses a staging step, not the Git Data API:
1066
+
1067
+ ```bash
1068
+ node scripts/stage_pages.mjs \
1069
+ --out=.deploy/dist-final \
1070
+ --include=_headers \
1071
+ --include=robots.txt \
1072
+ --include=assets/img/eo/provenance.json \
1073
+ --include=assets/img/eo/CREDITS.md
1074
+
1075
+ npx wrangler pages deploy "C:/Users/anish/satquery-ai/.deploy/dist-final" --project-name <name>
1076
+ ```
1077
+
1078
+ (`docs/DEPLOYMENT_DECISION.md` §7)
1079
+
1080
+ `_headers` and `robots.txt` must be **force-included** because no page references them; `provenance.json`
1081
+ and `CREDITS.md` likewise (`docs/DEPLOYMENT_DECISION.md` §7).
1082
+
1083
+ ---
1084
+
1085
+ # Part IX — Known operational gaps
1086
+
1087
+ ## 27. The explicit gaps list
1088
+
1089
+ Each row is something an operator might reasonably expect and that does **not** exist. None is a
1090
+ regression; each is a boundary of the current release.
1091
+
1092
+ | # | Gap | Consequence | Status |
1093
+ |---|---|---|---|
1094
+ | 1 | **No alerting** of any kind | an outage is discovered by trying to use the system | **not implemented** (§16) |
1095
+ | 2 | **No APM / metrics / distributed tracing** | no latency, error-rate or throughput trend exists | **not implemented** |
1096
+ | 3 | **No cost accounting** | consumption is unmeasured | **not implemented** (§19.2) |
1097
+ | 4 | **No dashboard / SLO / SLA** | no shared view of health; no target defined | **not implemented** |
1098
+ | 5 | **No structured logs / log shipping** | logs are per-host free text; the Codespace's are ephemeral | **not implemented** |
1099
+ | 6 | **B-07 is unfixed in production** | a request can hang ≈249 s then `504` | **`OPEN`** (§13) |
1100
+ | 7 | **B-02 trailing `\n`** | cosmetic; a wrong-looking field in health | **`OPEN` (cosmetic)** (§14) |
1101
+ | 8 | **No autoscaling, no replicas** | a restart is a full outage | **BY DESIGN** (plan §74) |
1102
+ | 9 | **No database, queue or persistence** | no run history survives a restart | **BY DESIGN** (`deploy/render/main.py:4-6`) |
1103
+ | 10 | **One Codespace** | capacity is bounded by one CPU host | **BY DESIGN** (§18) |
1104
+ | 11 | **No auth** | the contract documents *"no auth in v1"*; paths are scrubbed from client-visible fields as a partial mitigation | **BY DESIGN** |
1105
+ | 12 | **A system-level E2E benchmark does not exist** | no single system accuracy number can be quoted | **NOT RUN** |
1106
+ | 13 | **No measured cold-start distribution** | only *"tens of seconds"* is documented | **`UNKNOWN`** |
1107
+ | 14 | **The effective log level / destination per host** | not established | **`UNKNOWN`** |
1108
+ | 15 | **Whether `doctor.sh` / `tunnel_agent.py` exist in the deployed repo** | the monorepo copy is stale and lacks them | **`UNKNOWN`** (§7) |
1109
+ | 16 | **The B-07 patch is not deployed** | the fast-fail codes are absent in production | **`OPEN`** (§13) |
1110
+
1111
+ ## 28. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for operations
1112
+
1113
+ | # | Item | Status |
1114
+ |---|---|---|
1115
+ | 1 | B-07 — tunnel gaps; patch prepared, **not deployed** | **`OPEN`** |
1116
+ | 2 | B-02 — `/api/health` `codespace_name` trailing `\n` | **`OPEN` (cosmetic)** |
1117
+ | 3 | A deployed system-level load test | **`NOT RUN`** |
1118
+ | 4 | A measured cold-start distribution | **`NOT RUN`** — only "tens of seconds" is documented |
1119
+ | 5 | Multi-region / HA deployment | **`NOT RUN`** |
1120
+ | 6 | A system-level end-to-end benchmark | **`NOT RUN`** — none exists |
1121
+ | 7 | Any APM / metrics / distributed tracing / alerting | **not implemented** |
1122
+ | 8 | Cost accounting | **not implemented** — `GPU_DURATIONS` is declared, not metered |
1123
+ | 9 | The effective log level and destination on each host | **`UNKNOWN`** |
1124
+ | 10 | Whether the deployed `SatQuery-Inference` repo carries `doctor.sh` / `tunnel_agent.py` | **`UNKNOWN`** |
1125
+ | 11 | Log retention on Render | **`UNKNOWN`** — a platform property, not observable from the code |
1126
+ | 12 | The ZeroGPU/Gradio deployment target | **`REJECTED`** (superseded; frozen paperwork only) |
1127
+ | 13 | The five historical backend blockers | **closed by construction, not proven in production** |
1128
+
1129
+ > Rows 9 and 11 are copied from `docs/architecture/10-observability-and-ops.md` §8.1 so that the two
1130
+ > documents cannot drift. Row 13 is the honest framing from `release/repo/docs/DEPLOYMENT.md` §9: the
1131
+ > design closes the blockers, and the first live run is what would *verify* them.
1132
+
1133
+ ---
1134
+
1135
+ # Part X — Evidence
1136
+
1137
+ ## 29. Where the evidence lives
1138
+
1139
+ | What | Where |
1140
+ |---|---|
1141
+ | the live health payload | `release/CURRENT_RELEASE_STATE.md` §1; `release/repo/docs/DEPLOYMENT.md` §5 |
1142
+ | the live capability contract | `release/CURRENT_RELEASE_STATE.md` §1 |
1143
+ | the deployed revisions | `release/CURRENT_RELEASE_STATE.md` §1; `release/repo/docs/DEPLOYMENT.md` §1 |
1144
+ | the B-07 root shape and ≈249 s | `release/CURRENT_RELEASE_STATE.md` §6; `release/repo/docs/DEPLOYMENT.md` §8.1 |
1145
+ | the undeployed B-07 patch | session scratch: `fix-b07-forward-unavailable.patch` |
1146
+ | B-02's status and witness role | `release/repo/docs/DEPLOYMENT.md` §5; `docs/architecture/10-observability-and-ops.md` §7.4 |
1147
+ | the cold-start shape and the wake loop | `deploy/render/main.py:76-78,299-356,486-489` |
1148
+ | the Codespace launcher | `deploy/codespace/launch.sh`; `.devcontainer/devcontainer.json:18` |
1149
+ | the warm-up contract | `deploy/codespace/warm_cache.py` |
1150
+ | the asset-upload environment | `deploy/codespace/launch.sh:36-54` |
1151
+ | the deploy mechanics (Git Data API) | `release/repo/docs/DEPLOYMENT.md` §7 |
1152
+ | the platform traps | `release/repo/docs/DEPLOYMENT.md` §10; `release/repo/docs/REPRODUCIBILITY.md` §10 |
1153
+ | the historical backend blockers | `docs/DEPLOYMENT_DECISION.md` §8; `release/repo/docs/DEPLOYMENT.md` §9 |
1154
+ | the frozen config and hash | `configs/base.yaml`; `core/config.py:76-80` |
1155
+ | the observability inventory | `docs/architecture/10-observability-and-ops.md` §6 |
1156
+ | the triage tables | `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2, §7 |
1157
+ | the live validation (3 passes, 24 runs) | `.workbuddy-ai/scratch/live_validation/` |
1158
+ | the rate-limit finding (F-5) | `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2 |
1159
+
1160
+ ### Cross-references
1161
+
1162
+ | For | See |
1163
+ |---|---|
1164
+ | the four tiers, the tunnel, the wake flow, `transport_mode` | [02 — Deployment Topology](architecture/02-deployment-topology.md) |
1165
+ | health, counters, traces, what is and is not observed | [10 — Observability and Operations](architecture/10-observability-and-ops.md) |
1166
+ | the four endpoints, the envelopes, the error taxonomy | [08 — The API Contract](architecture/08-api-contract.md) |
1167
+ | the request lifecycle and the nine-state spine in motion | [03 — Request Lifecycle](architecture/03-request-lifecycle.md) |
1168
+ | the frozen config and `Config.hash == 78f1e3700da15aa1` | [07 — Configuration and Freeze](architecture/07-configuration-freeze.md) |
1169
+ | live revisions, env vars, deploy mechanics, platform traps | [../DEPLOYMENT.md](DEPLOYMENT.md) |
1170
+ | what a third party can and cannot reproduce | [../REPRODUCIBILITY.md](REPRODUCIBILITY.md) |
1171
+ | how to build, test and extend the codebase | [../DEVELOPMENT.md](DEVELOPMENT.md) |
1172
+
1173
+ ---
1174
+
1175
+ > **Chapter summary.** SatQuery AI runs four tiers — a static frontend, a thin Render orchestrator, one
1176
+ > CPU Codespace reached over an outbound tunnel, and a Hugging Face model tier — with exactly one
1177
+ > inference host and no replicas. The operator's single most important check is
1178
+ > `GET /api/health` → `tunnel.agent_connected: true`; the single most important diagnostic is the
1179
+ > ≈249 s-then-`504` signature of a B-07 tunnel gap. **B-07 is `OPEN` and the patch is not deployed; B-02
1180
+ > is `OPEN` and cosmetic.** The system observes three things — the transport, one run's trace, and the
1181
+ > server logs — and has **no** alerting, APM, distributed tracing or cost accounting. Capacity is one
1182
+ > free-tier Render service and one CPU Codespace; nothing is metered. Four things are genuinely
1183
+ > `UNKNOWN — not established from the available evidence`: the effective log level and destination per
1184
+ > host, log retention on Render, whether the deployed inference repository carries the tools the monorepo
1185
+ > launcher references, and any measured cold-start distribution.