Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
release: add docs/OPERATIONS.md
Browse files- docs/OPERATIONS.md +1185 -0
docs/OPERATIONS.md
ADDED
|
@@ -0,0 +1,1185 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Operations Manual
|
| 2 |
+
|
| 3 |
+
**Status tags:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · `NOT RUN` · `BLOCKED` ·
|
| 4 |
+
`DEFERRED` · `OPEN` · `RESOLVED` · `BY DESIGN`.
|
| 5 |
+
|
| 6 |
+
This is the operator-facing manual for the **live SatQuery AI stack**. It answers four questions that
|
| 7 |
+
the architecture chapters deliberately do not: *what is running, right now, and who owns it*; *how do I
|
| 8 |
+
bring it up and keep it up*; *how do I tell a transient transport gap from a real failure*; and *what
|
| 9 |
+
monitoring, capacity and cost machinery is **absent** so I do not assume it exists*.
|
| 10 |
+
|
| 11 |
+
It is written for the person who has to make the system answer a question in front of an audience, and
|
| 12 |
+
for the person who has to diagnose it at 23:00 when it does not.
|
| 13 |
+
|
| 14 |
+
> **Read this first.** The live system runs across **three private repositories** plus the public
|
| 15 |
+
> umbrella. The monorepo working copy — including the `deploy/` directory inside it — is **not** the
|
| 16 |
+
> deployed source. `deploy/` in the monorepo is **stale and untracked** (`git status` reports
|
| 17 |
+
> `?? deploy/`; verified in the working copy). Any operational fix must be applied to the real
|
| 18 |
+
> repositories, never to the monorepo copy (`docs/DEPLOYMENT.md` §1, §10).
|
| 19 |
+
|
| 20 |
+
> **Second rule.** Two known defects are `OPEN` and are **not** fixed in production: **B-07** (transient
|
| 21 |
+
> tunnel gaps) and **B-02** (a cosmetic trailing newline in one health field). Neither may be described
|
| 22 |
+
> as resolved, and the B-07 patch is **prepared but NOT deployed**. This document never upgrades them.
|
| 23 |
+
|
| 24 |
+
---
|
| 25 |
+
|
| 26 |
+
## Table of contents
|
| 27 |
+
|
| 28 |
+
**Part I — The operational model**
|
| 29 |
+
1. What runs where
|
| 30 |
+
2. The tier inventory, with the deployed revision of each
|
| 31 |
+
3. Who owns what
|
| 32 |
+
4. The operational invariants (four rules that must never be broken)
|
| 33 |
+
5. The stale-copy problem, in operational terms
|
| 34 |
+
|
| 35 |
+
**Part II — The runbook**
|
| 36 |
+
6. Warm the stack before a demo
|
| 37 |
+
7. Restart after an idle-stop
|
| 38 |
+
8. Tell a tunnel gap from a real failure
|
| 39 |
+
9. What the client shows while waking
|
| 40 |
+
|
| 41 |
+
**Part III — Cold start and the timing budget**
|
| 42 |
+
10. The cold-start shape
|
| 43 |
+
11. The four timeouts, and why the relationship matters
|
| 44 |
+
12. The worst case: ≈249 s under B-07
|
| 45 |
+
|
| 46 |
+
**Part IV — The known defects, in operational terms**
|
| 47 |
+
13. B-07 — transient tunnel gaps (`OPEN`)
|
| 48 |
+
14. B-02 — the trailing newline (`OPEN`, cosmetic)
|
| 49 |
+
|
| 50 |
+
**Part V — Monitoring and alerting**
|
| 51 |
+
15. What exists
|
| 52 |
+
16. What does **not** exist
|
| 53 |
+
17. Why "no alerting" is a design fact, not an oversight
|
| 54 |
+
|
| 55 |
+
**Part VI — Capacity and cost**
|
| 56 |
+
18. The capacity shape
|
| 57 |
+
19. The cost shape, and the one absence that has a code artifact
|
| 58 |
+
|
| 59 |
+
**Part VII — Incident triage**
|
| 60 |
+
20. The symptom → cause → check → action table
|
| 61 |
+
21. The consolidated decision tree
|
| 62 |
+
22. The machine-code reference
|
| 63 |
+
|
| 64 |
+
**Part VIII — Routine maintenance**
|
| 65 |
+
23. Rotating credentials (procedure only)
|
| 66 |
+
24. Restarting the tunnel agent
|
| 67 |
+
25. Re-warming the model cache
|
| 68 |
+
26. Deploying a change (the Git Data API path)
|
| 69 |
+
|
| 70 |
+
**Part IX — Known operational gaps**
|
| 71 |
+
27. The explicit gaps list
|
| 72 |
+
28. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for operations
|
| 73 |
+
|
| 74 |
+
**Part X — Evidence**
|
| 75 |
+
29. Where the evidence lives
|
| 76 |
+
|
| 77 |
+
---
|
| 78 |
+
|
| 79 |
+
# Part I — The operational model
|
| 80 |
+
|
| 81 |
+
## 1. What runs where
|
| 82 |
+
|
| 83 |
+
The live stack is four tiers in a straight line, plus a model tier that is reached *through* the
|
| 84 |
+
inference tier rather than by the user (`release/repo/docs/DEPLOYMENT.md` §2):
|
| 85 |
+
|
| 86 |
+
```
|
| 87 |
+
Browser
|
| 88 |
+
│ HTTPS
|
| 89 |
+
▼
|
| 90 |
+
Cloudflare Pages — satquery.pages.dev (static frontend, 11 pages)
|
| 91 |
+
│ HTTPS / JSON → /api/*
|
| 92 |
+
▼
|
| 93 |
+
Render — satquery-backend-m4yv.onrender.com (orchestrator / API gateway)
|
| 94 |
+
│ outbound long-poll POST /tunnel/agent
|
| 95 |
+
▼
|
| 96 |
+
GitHub Codespace — FastAPI inference, CPU, port 8000
|
| 97 |
+
│ build_space_app()
|
| 98 |
+
▼
|
| 99 |
+
specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
|
| 100 |
+
│
|
| 101 |
+
▼
|
| 102 |
+
ResultEnvelope → tunnel → Render → browser
|
| 103 |
+
```
|
| 104 |
+
|
| 105 |
+
```mermaid
|
| 106 |
+
flowchart LR
|
| 107 |
+
U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"]
|
| 108 |
+
CF -->|"HTTPS JSON<br/>/api/health · /api/capabilities · /api/infer · /api/assets"| R["Render<br/>orchestrator / gateway"]
|
| 109 |
+
R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"]
|
| 110 |
+
C --> S[(SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet)]
|
| 111 |
+
C -->|ResultEnvelope| R
|
| 112 |
+
R -->|"envelope + error translation"| CF
|
| 113 |
+
```
|
| 114 |
+
|
| 115 |
+
Three properties of this diagram matter operationally, and each is the subject of a section below:
|
| 116 |
+
|
| 117 |
+
1. **The transport is an outbound tunnel, not an inbound port.** The Codespace dials *out* to Render.
|
| 118 |
+
Render never dials into the Codespace. The transport is therefore alive only while an agent process
|
| 119 |
+
is polling — which is why "is the agent connected?" is the single most important operational
|
| 120 |
+
question (§6, §8).
|
| 121 |
+
2. **There is exactly one inference host.** One Codespace, one Render service, no replicas, no
|
| 122 |
+
autoscaling (`render.yaml` declares a single web service with `plan: free`; plan §74 lists
|
| 123 |
+
`autoscaling` under **Not included**). Capacity is therefore bounded by that one host (§18).
|
| 124 |
+
3. **Inference is CPU-only.** `SATQUERY_DEVICE=cpu` is set on both the orchestrator and the Codespace
|
| 125 |
+
(`render.yaml`, `.devcontainer/devcontainer.json`), and every specialist defaults to `device="cpu"`
|
| 126 |
+
(`docs/DEPLOYMENT_DECISION.md` §5). No GPU path is on the live critical path.
|
| 127 |
+
|
| 128 |
+
## 2. The tier inventory, with the deployed revision of each
|
| 129 |
+
|
| 130 |
+
Read from the GitHub API during the release reconnaissance (`release/CURRENT_RELEASE_STATE.md` §1;
|
| 131 |
+
`release/repo/docs/DEPLOYMENT.md` §1):
|
| 132 |
+
|
| 133 |
+
| Component | Repository | Visibility | Branch | Revision | Host |
|
| 134 |
+
|---|---|---|---|---|---|
|
| 135 |
+
| Frontend | `Anish-lab-blip/SatQuery-Frontend` | **private** | `main` | **`2d7ae53b482d`** | Cloudflare Pages → `satquery.pages.dev` |
|
| 136 |
+
| Backend / orchestrator | `Anish-lab-blip/SatQuery-Backend` | **private** | `main` | **`89d80eaddec5`** | Render → `satquery-backend-m4yv.onrender.com` |
|
| 137 |
+
| Inference | `Anish-lab-blip/SatQuery-Inference` | **private** | `main` | **`5a0936ace491`** | Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port 8000, via outbound tunnel |
|
| 138 |
+
| Public umbrella | `Anish-lab-blip/SatQuery-AI` | **public** | `main` | `3dcabd32da41` ("Initial commit") | this release home |
|
| 139 |
+
| Monorepo (working copy) | `C:/Users/anish/satquery-ai` | local only | `master` | `9d57aed` | **no git remote**; 334 dirty entries |
|
| 140 |
+
| Hugging Face | `thundercode/SatQuery` | **public** | `main` | lastModified `2026-09-25T16:26:53Z` | model tier |
|
| 141 |
+
|
| 142 |
+
The three private repositories are private **by design**; their links return 404 for an outside
|
| 143 |
+
audience (`release/CURRENT_RELEASE_STATE.md` §6). An operator therefore cannot browse the deployed
|
| 144 |
+
source from a public URL — the deployed files must be fetched with an authenticated API call
|
| 145 |
+
(`release/repo/docs/DEPLOYMENT.md` §7.1, §10).
|
| 146 |
+
|
| 147 |
+
> **The monorepo's `deploy/` is not the deployed source.** This is the single most important trap in
|
| 148 |
+
> the whole system (§5).
|
| 149 |
+
|
| 150 |
+
## 3. Who owns what
|
| 151 |
+
|
| 152 |
+
The ownership table below is derived from the code and the deployment records. "Owner" means *the
|
| 153 |
+
person or role that must act when this tier misbehaves*.
|
| 154 |
+
|
| 155 |
+
| Tier | Owner | What they own | What they must never do |
|
| 156 |
+
|---|---|---|---|
|
| 157 |
+
| Cloudflare Pages (frontend) | Frontend maintainer | the static bundle, `_headers`, `robots.txt`, the Analyze console | add a server-side secret — the tier holds none |
|
| 158 |
+
| Render (gateway) | Backend maintainer | the orchestrator revision, the env-var set, the CORS allowlist, the tunnel hub state | retry `POST /api/infer` (§4) |
|
| 159 |
+
| Codespace (inference) | Inference maintainer | the Codespace, the tunnel agent, the asset directory, the HF cache | let a stale serve process keep answering (§5) |
|
| 160 |
+
| Hugging Face (model tier) | Release owner | model cards, the pinned model references, the released checksums | treat the Hub as the runtime inference host — it is not |
|
| 161 |
+
| Credentials | Owner (human) | the GitHub PAT and the account tokens | record any credential value in a public document (§23) |
|
| 162 |
+
|
| 163 |
+
Two decisions are explicitly **not** an agent's to make, and both gate operational change
|
| 164 |
+
(`docs/PHASE19_FINAL_HARDENING.md` §7): the **SDK choice** (irrelevant on the live path, but still
|
| 165 |
+
unmade for the frozen manifest) and the **rate-limit / size-limit values** (which bound one client's
|
| 166 |
+
share of capacity). Neither is needed to operate the system as deployed.
|
| 167 |
+
|
| 168 |
+
## 4. The operational invariants
|
| 169 |
+
|
| 170 |
+
Four rules are load-bearing. Each is enforced somewhere in code or config, and each has a documented
|
| 171 |
+
failure if broken.
|
| 172 |
+
|
| 173 |
+
### 4.1 The config hash is frozen
|
| 174 |
+
|
| 175 |
+
`Config.hash == 78f1e3700da15aa1`. The loader computes a sha256 over the whole registry
|
| 176 |
+
(`core/config.py:76-80`) and every evaluation run records it. **Editing `configs/base.yaml` moves the
|
| 177 |
+
hash and invalidates every artifact keyed to it** (`release/repo/docs/DEPLOYMENT.md` §11;
|
| 178 |
+
`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §6.1). Deployment state that must *not* move the hash — asset-store
|
| 179 |
+
capacity, TTL, the per-file cap — is read from the **environment**, not from the YAML
|
| 180 |
+
(`release/repo/README.md` §Installation).
|
| 181 |
+
|
| 182 |
+
Operational consequence: **never edit `configs/base.yaml` to point at a deployment artifact.** The
|
| 183 |
+
serving path wires checkpoints through the registry's `builders=` override precisely so it does not have
|
| 184 |
+
to (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §6.1).
|
| 185 |
+
|
| 186 |
+
### 4.2 The gateway never retries `POST /api/infer`
|
| 187 |
+
|
| 188 |
+
> *"Render must not retry `POST /api/infer` on its own — a retry would consume inference a second
|
| 189 |
+
> time. The client decides on retry."* (`docs/DEPLOYMENT_TOPOLOGY.md` §2)
|
| 190 |
+
|
| 191 |
+
This is stated in three places (`docs/DEPLOYMENT_TOPOLOGY.md` §2, `docs/DEPLOYMENT.md` §7,
|
| 192 |
+
`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.3) because a retry is the natural thing to add and the wrong
|
| 193 |
+
thing to add. On the live CPU deployment it wastes compute; on the historical ZeroGPU target it spent a
|
| 194 |
+
metered GPU-minute twice.
|
| 195 |
+
|
| 196 |
+
### 4.3 The CORS allowlist is explicit and never a wildcard
|
| 197 |
+
|
| 198 |
+
The gateway assembles its allowlist from `SATQUERY_ALLOWED_ORIGINS` plus a hard-coded production origin
|
| 199 |
+
plus a fixed list of development origins (`deploy/render/main.py:139-216`). A `*` raises
|
| 200 |
+
(`deploy/render/main.py:204-208`). The live value is `https://satquery.pages.dev`
|
| 201 |
+
(`release/repo/docs/DEPLOYMENT.md` §6.1).
|
| 202 |
+
|
| 203 |
+
Operational consequence: a new frontend origin must be **added** to the env var; it will not work by
|
| 204 |
+
accident.
|
| 205 |
+
|
| 206 |
+
### 4.4 One inference host, one asset store
|
| 207 |
+
|
| 208 |
+
There is one Codespace, and its filesystem is **ephemeral** (`deploy/codespace/launch.sh:44-47`). Uploaded
|
| 209 |
+
assets are written under `SATQUERY_ASSET_DIR` (default `/tmp/satquery-assets`) and are TTL'd (900 s
|
| 210 |
+
default). A Codespace restart empties the store and makes every previously issued handle unresolvable
|
| 211 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §3.1.1).
|
| 212 |
+
|
| 213 |
+
Operational consequence: a handle that worked seconds ago may return `400 input_error` after a restart.
|
| 214 |
+
That is documented behaviour, not a bug (§20).
|
| 215 |
+
|
| 216 |
+
## 5. The stale-copy problem, in operational terms
|
| 217 |
+
|
| 218 |
+
Three copies of the deployment code exist, and confusing them is the most expensive operational
|
| 219 |
+
mistake in the system.
|
| 220 |
+
|
| 221 |
+
| Copy | What it is | Trustworthy? |
|
| 222 |
+
|---|---|---|
|
| 223 |
+
| the monorepo `deploy/` | local, **untracked** (`?? deploy/`), stale | **no** — it is not the deployed source |
|
| 224 |
+
| the session scratch copy | a local copy used to author and verify the B-07 patch | **no** — it is "deployed + patch", not deployed |
|
| 225 |
+
| the private repositories | the real deployed source | **yes** — fetch it before editing |
|
| 226 |
+
|
| 227 |
+
Evidence for the divergence is direct. The monorepo's `deploy/render/main.py` (532 lines) exposes
|
| 228 |
+
`/api/health` with a `config` block that has **no** `tunnel` field and **no** `transport_mode`,
|
| 229 |
+
`tunnel_timeout_s` or `wake_timeout_s` keys (`deploy/render/main.py:444-466`), whereas the **live**
|
| 230 |
+
payload carries all of them (`release/repo/docs/DEPLOYMENT.md` §5). The monorepo copy also contains no
|
| 231 |
+
`tunnel_agent.py` and no `doctor.sh`, even though `deploy/codespace/launch.sh` invokes both
|
| 232 |
+
(`deploy/codespace/launch.sh:83,89,153,160-168,179`). The two are different programs.
|
| 233 |
+
|
| 234 |
+
> **Operational rule.** Before changing anything, fetch the deployed `main.py` from the private
|
| 235 |
+
> repository and diff it against what you are about to edit. The monorepo copy will silently disagree.
|
| 236 |
+
|
| 237 |
+
The `deploy/codespace/launch.sh` file *is* useful as documentation of intent — its header explains why
|
| 238 |
+
it is defensive (`deploy/codespace/launch.sh:14-24`) — but it is a copy, and its references to
|
| 239 |
+
`tunnel_agent.py` resolve only in the deployed repository.
|
| 240 |
+
|
| 241 |
+
---
|
| 242 |
+
|
| 243 |
+
# Part II — The runbook
|
| 244 |
+
|
| 245 |
+
Every step in this part is grounded in a file. Commands are quoted as they appear in the sources.
|
| 246 |
+
|
| 247 |
+
## 6. Warm the stack before a demo
|
| 248 |
+
|
| 249 |
+
Three things can be cold, and all three are warmed differently (`docs/architecture/10-observability-and-ops.md`
|
| 250 |
+
§7.1):
|
| 251 |
+
|
| 252 |
+
| What is cold | How it warms | Bound |
|
| 253 |
+
|---|---|---|
|
| 254 |
+
| the Render orchestrator | the first request to any `/api/*` route | Render free tier sleep/wake cycle |
|
| 255 |
+
| the Codespace | `ensure_codespace_up()` starts it and polls | `wake_timeout_s = 120` |
|
| 256 |
+
| the HF model cache | `warm_cache.py`, or the first model-touching request | one download per model |
|
| 257 |
+
|
| 258 |
+
### Step 1 — confirm the tunnel agent is connected
|
| 259 |
+
|
| 260 |
+
This is the check that matters most, because the tunnel is the live transport and the forwarded-port
|
| 261 |
+
path is dead for a private repository (`release/CURRENT_RELEASE_STATE.md` §2: *"The forwarded-port path
|
| 262 |
+
is **dead** (302 for a private repo); the tunnel is the live transport."*).
|
| 263 |
+
|
| 264 |
+
```bash
|
| 265 |
+
curl --noproxy '*' https://satquery-backend-m4yv.onrender.com/api/health
|
| 266 |
+
```
|
| 267 |
+
|
| 268 |
+
Read `tunnel.agent_connected`. `true` means an agent has polled recently; `false` means **no agent has
|
| 269 |
+
polled recently**, and under `transport_mode: auto` a request will then take the forward path, which for
|
| 270 |
+
a private repository fails after burning the wake timeout (`docs/architecture/10-observability-and-ops.md`
|
| 271 |
+
§7.1).
|
| 272 |
+
|
| 273 |
+
> **`--noproxy '*'` is not optional in the authoring sandbox.** The sandbox proxy is dead; without the
|
| 274 |
+
> flag the request fails before reaching Render. In a normal environment the flag is harmless
|
| 275 |
+
> (`release/repo/docs/REPRODUCIBILITY.md` §10.1).
|
| 276 |
+
|
| 277 |
+
### Step 2 — if `agent_connected` is false, start the Codespace
|
| 278 |
+
|
| 279 |
+
The agent is launched by the devcontainer's `postStartCommand`, which runs `launch.sh`:
|
| 280 |
+
|
| 281 |
+
```json
|
| 282 |
+
"postStartCommand": "bash deploy/codespace/launch.sh"
|
| 283 |
+
```
|
| 284 |
+
|
| 285 |
+
(`.devcontainer/devcontainer.json:18`)
|
| 286 |
+
|
| 287 |
+
Starting the Codespace is what re-runs `postStartCommand` and therefore reconnects the agent (§7).
|
| 288 |
+
`launch.sh` is defensive by design, and its own header explains why:
|
| 289 |
+
|
| 290 |
+
> *"`setsid` alone is NOT enough in Codespaces. The lifecycle shell that runs postStartCommand can still
|
| 291 |
+
> reap the process group, which showed up in production as 'the agent announced once, then vanished' —
|
| 292 |
+
> the hub then reported agent_connected=false and /api/infer fell back to the dead forwarded-port path
|
| 293 |
+
> (401 -> wake_timeout)."* (`deploy/codespace/launch.sh:14-19`)
|
| 294 |
+
|
| 295 |
+
The script therefore uses `setsid + nohup + </dev/null` plus a **supervising wrapper** that restarts the
|
| 296 |
+
agent if it exits (`deploy/codespace/launch.sh:20-23,160-168`), and then **verifies** the agent came up:
|
| 297 |
+
|
| 298 |
+
```bash
|
| 299 |
+
sleep 4
|
| 300 |
+
|
| 301 |
+
if ! pgrep -f "deploy/codespace/tunnel_agent.py" > /dev/null 2>&1; then
|
| 302 |
+
echo "WARNING: the tunnel agent is not running. Last log lines:" >&2
|
| 303 |
+
...
|
| 304 |
+
else
|
| 305 |
+
echo "tunnel agent process is up (pid $(pgrep -f 'deploy/codespace/tunnel_agent.py' | head -1))"
|
| 306 |
+
if grep -q "announced to hub" "$TUNNEL_LOG" 2>/dev/null; then
|
| 307 |
+
echo "tunnel agent announced to the hub successfully"
|
| 308 |
+
...
|
| 309 |
+
```
|
| 310 |
+
|
| 311 |
+
(`deploy/codespace/launch.sh:177-191`)
|
| 312 |
+
|
| 313 |
+
The verification exists because of a specific failure:
|
| 314 |
+
|
| 315 |
+
> *"Backgrounding with all output discarded means a crashing agent is completely invisible — that is
|
| 316 |
+
> exactly how a missing `httpx` hid itself."* (`deploy/codespace/launch.sh:174-176`)
|
| 317 |
+
|
| 318 |
+
### Step 3 — warm the HF cache, if the Codespace was rebuilt
|
| 319 |
+
|
| 320 |
+
`warm_cache.py` pre-downloads the four pinned models:
|
| 321 |
+
|
| 322 |
+
```bash
|
| 323 |
+
python deploy/codespace/warm_cache.py
|
| 324 |
+
```
|
| 325 |
+
|
| 326 |
+
It reports per-model `OK` / `SKIPPED` / `FAILED`, never aborts on a single miss, and **always exits 0** so
|
| 327 |
+
a cache miss cannot fail a build (`deploy/codespace/warm_cache.py:8-10,110-112`). The four models and
|
| 328 |
+
their pinned revisions are transcribed verbatim from `configs/base.yaml`
|
| 329 |
+
(`deploy/codespace/warm_cache.py:12-16,35-60`):
|
| 330 |
+
|
| 331 |
+
| key | repo | revision | file |
|
| 332 |
+
|---|---|---|---|
|
| 333 |
+
| `vlm` | `HuggingFaceTB/SmolVLM-500M-Instruct` | `a7da5b986cb5` | (snapshot) |
|
| 334 |
+
| `grounding` | `chendelong/RemoteCLIP` | `bf1d8a3ccf2d` | `RemoteCLIP-ViT-B-32.pt` |
|
| 335 |
+
| `router` | `sentence-transformers/all-MiniLM-L6-v2` | `1110a243fdf4` | (snapshot) |
|
| 336 |
+
| `croma` | `antofuller/CROMA` | `0dd28e3d633b` | `CROMA_base.pt` |
|
| 337 |
+
|
| 338 |
+
It is idempotent and safe to re-run; a second run is a no-op because the blob is already on disk
|
| 339 |
+
(`deploy/codespace/warm_cache.py:3-6`). Two environment switches skip work:
|
| 340 |
+
`SATQUERY_WARM_OFFLINE` / `HF_HUB_OFFLINE` (skip everything) and `SATQUERY_WARM_SKIP="croma,grounding"`
|
| 341 |
+
(skip named models) (`deploy/codespace/warm_cache.py:23-25`).
|
| 342 |
+
|
| 343 |
+
> **Note.** `warm_cache.py` warms four models. The **fifth** pinned backbone in the specialist stack —
|
| 344 |
+
`STANet`/change — is a local trained head, not a Hub backbone, and is not in the warm list. If the change
|
| 345 |
+
head is absent the capability degrades honestly rather than failing the warm step.
|
| 346 |
+
|
| 347 |
+
### Step 4 — run one throwaway analysis
|
| 348 |
+
|
| 349 |
+
A single cheap query confirms the whole chain end to end. Do this **before** the demo, not during it.
|
| 350 |
+
Watch for: a `run_id`, a `mock_nodes` count of 0, and an answer carrying a `[task]` tag
|
| 351 |
+
(`release/repo/docs/REPRODUCIBILITY.md` §6.3).
|
| 352 |
+
|
| 353 |
+
### The warm-state checklist
|
| 354 |
+
|
| 355 |
+
| Check | Command | Expected |
|
| 356 |
+
|---|---|---|
|
| 357 |
+
| orchestrator up | `curl --noproxy '*' .../api/health` | `status: ok`, `service: satquery-orchestrator` |
|
| 358 |
+
| tunnel connected | same payload | `tunnel.agent_connected: true` |
|
| 359 |
+
| capabilities | `curl --noproxy '*' .../api/capabilities` | six tasks, all `available: true` |
|
| 360 |
+
| one live run | drive the Analyze console | a `run_id`, `mock_nodes: 0` |
|
| 361 |
+
|
| 362 |
+
## 7. Restart after an idle-stop
|
| 363 |
+
|
| 364 |
+
A Codespace stops after an idle period. When it stops, the tunnel agent stops polling, and
|
| 365 |
+
`GET /api/health` reports `tunnel.agent_connected: false`
|
| 366 |
+
(`docs/DEPLOYMENT_TOPOLOGY.md` §2).
|
| 367 |
+
|
| 368 |
+
**The reconnect is automatic on start, because `postStartCommand` runs `launch.sh`.** The chain is:
|
| 369 |
+
|
| 370 |
+
```
|
| 371 |
+
Codespace start
|
| 372 |
+
→ devcontainer postStartCommand: bash deploy/codespace/launch.sh (.devcontainer/devcontainer.json:18)
|
| 373 |
+
→ launch.sh: start serve.py on $PORT (if not already current) (launch.sh:110-147)
|
| 374 |
+
→ launch.sh: start the supervised tunnel agent -> $SATQUERY_HUB_URL (launch.sh:149-169)
|
| 375 |
+
→ launch.sh: sleep 4, verify the agent process and the announce line (launch.sh:171-191)
|
| 376 |
+
→ the agent dials POST /tunnel/agent and long-polls (launch.sh:150-156)
|
| 377 |
+
→ GET /api/health: tunnel.agent_connected becomes true
|
| 378 |
+
```
|
| 379 |
+
|
| 380 |
+
The hub URL the agent dials is `SATQUERY_HUB_URL`, defaulting to
|
| 381 |
+
`https://satquery-backend-m4yv.onrender.com` (`deploy/codespace/launch.sh:59-61`).
|
| 382 |
+
|
| 383 |
+
**What the operator does:**
|
| 384 |
+
|
| 385 |
+
1. Start the Codespace (or let the wake path start it — `ensure_codespace_up()` calls the GitHub
|
| 386 |
+
Codespaces `POST .../start` API when `state != "available"`, `deploy/render/main.py:299-356`).
|
| 387 |
+
2. Wait for `postStartCommand` to run.
|
| 388 |
+
3. Re-read `GET /api/health` and confirm `tunnel.agent_connected: true` (§6 step 1).
|
| 389 |
+
|
| 390 |
+
**Two things that make a restart go wrong, and their mitigation:**
|
| 391 |
+
|
| 392 |
+
| Failure | Symptom | Mitigation in `launch.sh` |
|
| 393 |
+
|---|---|---|
|
| 394 |
+
| The agent is launched but immediately reaped by the lifecycle shell | agent announces once, then vanishes; hub reports `agent_connected: false`; `/api/infer` falls back to the dead forward path (401 → wake_timeout) | `setsid + nohup + </dev/null` plus a supervising restart loop (`launch.sh:14-24,160-168`) |
|
| 395 |
+
| A **stale** serve process keeps answering from OLD code | `/v1/health` and `/v1/capabilities` answer, but from the previous revision's capabilities | a **stamp** recording the revision + asset config; a mismatch restarts the server (`launch.sh:100-147`) |
|
| 396 |
+
|
| 397 |
+
> *"A stale serve process is worse than no process: it answers /v1/health and /v1/capabilities from OLD
|
| 398 |
+
> code, so the deployment looks alive while reporting the previous revision's capabilities."*
|
| 399 |
+
> (`deploy/codespace/launch.sh:111-113`)
|
| 400 |
+
|
| 401 |
+
The stamp is the closest thing in the system to a deployment-identity check:
|
| 402 |
+
|
| 403 |
+
```bash
|
| 404 |
+
_current_stamp() {
|
| 405 |
+
printf 'rev=%s asset_enabled=%s asset_dir=%s\n' \
|
| 406 |
+
"$(git rev-parse HEAD 2>/dev/null || echo nogit)" \
|
| 407 |
+
"${SATQUERY_ASSET_ENABLED:-}" \
|
| 408 |
+
"${SATQUERY_ASSET_DIR:-}"
|
| 409 |
+
}
|
| 410 |
+
```
|
| 411 |
+
|
| 412 |
+
(`deploy/codespace/launch.sh:103-108`)
|
| 413 |
+
|
| 414 |
+
It is **local to the Codespace** and is not exposed on any HTTP route — an operator on the orchestrator
|
| 415 |
+
side cannot see it (`docs/architecture/10-observability-and-ops.md` §7.5).
|
| 416 |
+
|
| 417 |
+
### The preflight that refuses a half-configured start
|
| 418 |
+
|
| 419 |
+
`launch.sh` refuses to start if the Python dependencies or the `app` package cannot be imported
|
| 420 |
+
(`deploy/codespace/launch.sh:70-91`). The dependency check is explicit about `httpx`, because a missing
|
| 421 |
+
`httpx` once made the agent die instantly and the supervised loop hid the error in a log file
|
| 422 |
+
(`deploy/codespace/launch.sh:76-79`):
|
| 423 |
+
|
| 424 |
+
```bash
|
| 425 |
+
if ! python -c "import yaml, pydantic, fastapi, uvicorn, httpx" 2>/dev/null; then
|
| 426 |
+
echo "ERROR: Python deps are missing (need yaml, pydantic, fastapi, uvicorn, httpx)." >&2
|
| 427 |
+
...
|
| 428 |
+
exit 1
|
| 429 |
+
fi
|
| 430 |
+
```
|
| 431 |
+
|
| 432 |
+
> **Caveat — the referenced repair tool is not in this tree.** `launch.sh` points the operator at
|
| 433 |
+
> `bash deploy/codespace/doctor.sh --install` (`launch.sh:83,89,182`), but **`doctor.sh` does not exist in
|
| 434 |
+
> the monorepo working copy**, and neither does `tunnel_agent.py` (§5). Whether `doctor.sh` exists in the
|
| 435 |
+
> deployed `SatQuery-Inference` repository is `UNKNOWN — not established from the available evidence`.
|
| 436 |
+
|
| 437 |
+
## 8. Tell a tunnel gap from a real failure
|
| 438 |
+
|
| 439 |
+
This is the runbook's most useful procedure, and it is grounded in a measured timing.
|
| 440 |
+
|
| 441 |
+
### The signature
|
| 442 |
+
|
| 443 |
+
A request that hangs for **≈249 seconds** and then returns **`504`** is the tunnel-gap signature, not a
|
| 444 |
+
broken model. The arithmetic is exact:
|
| 445 |
+
|
| 446 |
+
```
|
| 447 |
+
tunnel_timeout_s (150) + wake_timeout_s (120) = 270 s (nominal)
|
| 448 |
+
measured ≈ 249 s
|
| 449 |
+
```
|
| 450 |
+
|
| 451 |
+
> *"B-07 root shape: in `auto` mode a tunnel timeout **falls through** to the forward path
|
| 452 |
+
> (`main.py:546`), burning `wake_timeout_s = 120` on a `302` (~249 s ≈ 150 + 120)."*
|
| 453 |
+
> (`release/CURRENT_RELEASE_STATE.md` §6; `release/repo/docs/DEPLOYMENT.md` §8.1)
|
| 454 |
+
|
| 455 |
+
### The three signals to read, in order
|
| 456 |
+
|
| 457 |
+
1. **`tunnel.agent_connected` on a fresh `/api/health`.** If `false`, it is a tunnel gap and the remedy is
|
| 458 |
+
§6 step 2. Do not trust a single reading — the flag is a freshness window, so re-read.
|
| 459 |
+
2. **The elapsed time.** ≈249 s is the B-07 signature. A fast failure is something else.
|
| 460 |
+
3. **The machine code.** The codes are disjoint and each implies a different action (§22).
|
| 461 |
+
|
| 462 |
+
### The decision tree
|
| 463 |
+
|
| 464 |
+
```mermaid
|
| 465 |
+
flowchart TD
|
| 466 |
+
A["a request hung, or returned 5xx"] --> B{"re-read /api/health<br/>tunnel.agent_connected?"}
|
| 467 |
+
B -->|"true"| C{"was the wait ≈249 s?"}
|
| 468 |
+
B -->|"false"| D["TUNNEL GAP<br/>the agent is not polling.<br/>Start the Codespace (§6)."]
|
| 469 |
+
C -->|"yes, 504"| E["TUNNEL GAP<br/>agent went stale mid-request,<br/>or the Codespace stopped.<br/>B-07. Mitigate operationally."]
|
| 470 |
+
C -->|"no"| F{"what was the code?"}
|
| 471 |
+
F -->|"invalid_request 422"| G["a CLIENT bug — the body<br/>did not match AnalysisRequest"]
|
| 472 |
+
F -->|"model_load_error / model_unavailable"| H["an ARTIFACT defect —<br/>surface it, do not retry blindly"]
|
| 473 |
+
F -->|"upstream_timeout 504"| I["tunnel healthy but slow —<br/>the agent did not complete in 150 s"]
|
| 474 |
+
F -->|"other"| J["read the code and<br/>trace.errors[]"]
|
| 475 |
+
```
|
| 476 |
+
|
| 477 |
+
> **Note which branches exist only in the patch.** `forward_unavailable` and `upstream_timeout` do **not**
|
| 478 |
+
> exist in the deployed revision (§13). An operator on the deployed system will not see them; they will
|
| 479 |
+
> see `wake_timeout` after ≈249 s instead.
|
| 480 |
+
|
| 481 |
+
## 9. What the client shows while waking
|
| 482 |
+
|
| 483 |
+
The frontend is not silent during a cold start. It shows *"Waking inference engine…"* while Render starts
|
| 484 |
+
the Codespace (`docs/DEPLOYMENT_TOPOLOGY.md` §2, §2 mermaid; `release/repo/docs/DEPLOYMENT.md` §8).
|
| 485 |
+
|
| 486 |
+
The orchestrator's side of this is a response header. The wake path is **blocking wake-then-proxy**: the
|
| 487 |
+
client waits and receives the result, and the response is tagged:
|
| 488 |
+
|
| 489 |
+
```python
|
| 490 |
+
out.headers["X-SatQuery-State"] = "waking" if woke else "ready"
|
| 491 |
+
```
|
| 492 |
+
|
| 493 |
+
(`deploy/render/main.py:486-489`)
|
| 494 |
+
|
| 495 |
+
The module docstring states the intent plainly:
|
| 496 |
+
|
| 497 |
+
> *"The client simply **waits** (blocking, wake-then-proxy) and receives the result; the frontend
|
| 498 |
+
> independently shows 'Waking inference engine...' on slow responses. We optionally tag the response with
|
| 499 |
+
> `X-SatQuery-State: waking` so the frontend can confirm the delay was a cold start."*
|
| 500 |
+
> (`deploy/render/main.py:13-20`)
|
| 501 |
+
|
| 502 |
+
So an operator watching a demo sees: the client's "Waking inference engine…" message, a wait, and then
|
| 503 |
+
either a result or a `504`. **A `504` after ≈249 s is the B-07 shape, not a broken model** (§8).
|
| 504 |
+
|
| 505 |
+
---
|
| 506 |
+
|
| 507 |
+
# Part III — Cold start and the timing budget
|
| 508 |
+
|
| 509 |
+
## 10. The cold-start shape
|
| 510 |
+
|
| 511 |
+
The cold start is **blocking and documented**, not hidden. The shape an operator should expect
|
| 512 |
+
(`docs/architecture/10-observability-and-ops.md` §7.2):
|
| 513 |
+
|
| 514 |
+
| Phase | What happens | Bound |
|
| 515 |
+
|---|---|---|
|
| 516 |
+
| 0 | the client sends `POST /api/infer` and **waits** | — |
|
| 517 |
+
| 1 | `GET` the Codespace via the GitHub API; `POST .../start` if not `available` | one API round trip |
|
| 518 |
+
| 2 | poll `GET {base}/v1/health` every 2 s, each with a 10 s timeout | ≤ `wake_timeout_s = 120` |
|
| 519 |
+
| 3 | on the first `200`, proxy the original request | ≤ `upstream_timeout_s = 90` |
|
| 520 |
+
| 4 | the response carries `X-SatQuery-State: waking` | — |
|
| 521 |
+
|
| 522 |
+
The poll interval and per-poll timeout are module constants:
|
| 523 |
+
|
| 524 |
+
```python
|
| 525 |
+
# Polling knobs for the wake loop.
|
| 526 |
+
_WAKE_POLL_INTERVAL_S = 2.0
|
| 527 |
+
_WAKE_HEALTH_TIMEOUT_S = 10.0
|
| 528 |
+
```
|
| 529 |
+
|
| 530 |
+
(`deploy/render/main.py:76-78`)
|
| 531 |
+
|
| 532 |
+
> **Do not quote a single cold-start number.** The documented statement is *"tens of seconds"*
|
| 533 |
+
> (`release/repo/docs/DEPLOYMENT.md` §8), and the bounded worst case is `wake_timeout_s = 120` before a
|
| 534 |
+
> `504`. **No measured cold-start distribution exists:**
|
| 535 |
+
> `UNKNOWN — not established from the available evidence`
|
| 536 |
+
> (`docs/architecture/10-observability-and-ops.md` §7.2).
|
| 537 |
+
|
| 538 |
+
## 11. The four timeouts, and why the relationship matters
|
| 539 |
+
|
| 540 |
+
The live payload reports three of them; the fourth comes from the frozen config.
|
| 541 |
+
|
| 542 |
+
| Timeout | Value | Where it lives | What it bounds |
|
| 543 |
+
|---|---|---|---|
|
| 544 |
+
| `tunnel_timeout_s` | **150.0 s** | live config (`release/repo/docs/DEPLOYMENT.md` §5) | how long a request waits on the tunnel before falling through |
|
| 545 |
+
| `wake_timeout_s` | **120.0 s** | live config | how long the wake loop waits for `/v1/health` |
|
| 546 |
+
| `upstream_timeout_s` | **90.0 s** | live config | the gateway → upstream proxy budget |
|
| 547 |
+
| `agent.timeout_seconds` | **120 s** | `configs/base.yaml` (`agent.timeout_seconds: 120`) | the Space's own total request budget |
|
| 548 |
+
|
| 549 |
+
The **relationship** between the last two is the one that must not be inverted
|
| 550 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.2):
|
| 551 |
+
|
| 552 |
+
```
|
| 553 |
+
gateway upstream timeout < agent.timeout_seconds ≤ the Space's own request budget
|
| 554 |
+
90 s < 120 s
|
| 555 |
+
```
|
| 556 |
+
|
| 557 |
+
Both failure directions are documented:
|
| 558 |
+
|
| 559 |
+
- **Gateway timeout too short** — it kills a legitimately running `grounding` or `optical_sar` call and
|
| 560 |
+
reports it as an upstream failure. The Space's error, not the client's.
|
| 561 |
+
- **Gateway timeout too long** — it holds a connection past the point the Space itself has given up,
|
| 562 |
+
converting a clean upstream timeout into a client-side hang.
|
| 563 |
+
|
| 564 |
+
> The lower bound of the *historical* window was 45 s (the longest single `gpu_duration_*`). On the live
|
| 565 |
+
> **CPU** deployment the ZeroGPU durations are frozen paperwork (§19), so the binding upper constraint is
|
| 566 |
+
> the 120 s agent timeout and the live gateway value is 90 s.
|
| 567 |
+
|
| 568 |
+
## 12. The worst case: ≈249 s under B-07
|
| 569 |
+
|
| 570 |
+
`SATQUERY_TRANSPORT=auto` means **try the tunnel; on timeout, fall through to the forward path**
|
| 571 |
+
(`release/repo/docs/DEPLOYMENT.md` §8.1). The forward path to a private repository returns `302` quickly,
|
| 572 |
+
but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S` (120 s) first. So a worst-case failed request
|
| 573 |
+
takes roughly:
|
| 574 |
+
|
| 575 |
+
```
|
| 576 |
+
150 s (tunnel timeout) + 120 s (wake timeout on a 302) ≈ 249 s
|
| 577 |
+
```
|
| 578 |
+
|
| 579 |
+
This is the **root shape** of the observed transient tunnel gap, and it is why a request can appear to
|
| 580 |
+
hang and then fail (`release/repo/docs/DEPLOYMENT.md` §8.1). It is `OPEN` (§13).
|
| 581 |
+
|
| 582 |
+
---
|
| 583 |
+
|
| 584 |
+
# Part IV — The known defects, in operational terms
|
| 585 |
+
|
| 586 |
+
## 13. B-07 — transient tunnel gaps (`OPEN`)
|
| 587 |
+
|
| 588 |
+
**B-07 is `OPEN`.** The patch is prepared and **not deployed**. This is the single most important
|
| 589 |
+
operational fact in this manual.
|
| 590 |
+
|
| 591 |
+
> *"B-07 | Transient tunnel-agent gaps → a request can hang or return 504. Patch prepared, **NOT
|
| 592 |
+
> deployed**. | **OPEN**"* (`release/CURRENT_RELEASE_STATE.md` §6)
|
| 593 |
+
|
| 594 |
+
### What an operator experiences
|
| 595 |
+
|
| 596 |
+
- The tunnel agent is briefly absent (a restart, a reap, a gap).
|
| 597 |
+
- A request issued during the gap either hangs or returns `504`.
|
| 598 |
+
- In `auto` mode the hang lasts up to ≈249 s before the `504` (§12).
|
| 599 |
+
- Once the agent reconnects, the next request succeeds.
|
| 600 |
+
|
| 601 |
+
### The root cause, exactly
|
| 602 |
+
|
| 603 |
+
In `auto` mode a tunnel timeout **falls through to the forward path** (`main.py:546`), and the forward
|
| 604 |
+
path to a private repository returns `302`. The wake step burns `wake_timeout_s = 120` on that `302`
|
| 605 |
+
before the request fails (`release/CURRENT_RELEASE_STATE.md` §6).
|
| 606 |
+
|
| 607 |
+
### What the patch does
|
| 608 |
+
|
| 609 |
+
The patch was authored and verified (`py_compile` clean, applies cleanly to the deployed `main.py`)
|
| 610 |
+
(`release/repo/docs/DEPLOYMENT.md` §8.1). It makes three changes:
|
| 611 |
+
|
| 612 |
+
| Change | Code | Effect |
|
| 613 |
+
|---|---|---|
|
| 614 |
+
| A | `forward_unavailable` (`503`, `recoverable: true`) on a **terminal** `302`/`401`/`403` | converts a 504-after-249 s into a 503-early with an actionable code |
|
| 615 |
+
| B | `upstream_timeout` (`504`) for "tunnel healthy but slow" | distinguishes a slow agent from a dead forward path |
|
| 616 |
+
| C | `/api/health` `codespace_name` `.strip()` | fixes B-02 |
|
| 617 |
+
|
| 618 |
+
### The honesty note attached to the patch
|
| 619 |
+
|
| 620 |
+
> *"The report records that an earlier claim that the patch 'would not have prevented' the observed 504
|
| 621 |
+
> 'was wrong and was retracted'. The corrected position: 'Change A is genuinely **on the failing path** —
|
| 622 |
+
> it converts a 504-after-249 s into a 503-early with an actionable code.'"*
|
| 623 |
+
> (`release/CURRENT_RELEASE_STATE.md` §8 evidence list; `docs/architecture/02-deployment-topology.md` §6.4)
|
| 624 |
+
|
| 625 |
+
Do not repeat the retracted version.
|
| 626 |
+
|
| 627 |
+
### Why it is not deployed
|
| 628 |
+
|
| 629 |
+
> *"the patch is not needed for the demo and touches the live backend. The residual is better mitigated
|
| 630 |
+
> operationally (keep the Codespace warm, raise the idle timeout)."*
|
| 631 |
+
> (`DELIVERY_REPORT_2026-09-25.md` §4, quoted in `docs/architecture/10-observability-and-ops.md` §7.4)
|
| 632 |
+
|
| 633 |
+
### The operational mitigation (this is what an operator actually does)
|
| 634 |
+
|
| 635 |
+
1. Keep the Codespace **warm** before and during any demo (§6).
|
| 636 |
+
2. **Raise the Codespace idle timeout** so it does not stop mid-session.
|
| 637 |
+
3. Re-read `/api/health` before a run, and treat `agent_connected: false` as "warm the stack now".
|
| 638 |
+
4. Expect a ≈249 s hang followed by a `504` if the agent goes stale mid-request — and **retry**, because
|
| 639 |
+
the residual is transient.
|
| 640 |
+
|
| 641 |
+
> **Do not upgrade B-07.** It is `OPEN`. It is not `RESOLVED`, and the patch is not deployed — the live
|
| 642 |
+
> payload's `codespace_name` trailing `\n` is the witness (§14).
|
| 643 |
+
|
| 644 |
+
## 14. B-02 — the trailing newline (`OPEN`, cosmetic)
|
| 645 |
+
|
| 646 |
+
`GET /api/health` reports the Codespace name with a trailing newline:
|
| 647 |
+
|
| 648 |
+
```json
|
| 649 |
+
"codespace_name": "potential-space-trout-r4ppw969w45j2pvvw\n"
|
| 650 |
+
```
|
| 651 |
+
|
| 652 |
+
(`release/repo/docs/DEPLOYMENT.md` §5; `release/CURRENT_RELEASE_STATE.md` §1)
|
| 653 |
+
|
| 654 |
+
**It is cosmetic.** The wake path strips it — `_codespace_name()` calls `.strip()` before using the value
|
| 655 |
+
(`deploy/render/main.py:99-106`) — so only the health payload reports the raw value
|
| 656 |
+
(`release/repo/docs/DEPLOYMENT.md` §5).
|
| 657 |
+
|
| 658 |
+
**It is `OPEN`.** Its presence is also the operational **witness** that the B-07 patch is not deployed:
|
| 659 |
+
change C of that patch is the `.strip()` fix, so a live payload still showing the trailing `\n` proves the
|
| 660 |
+
patch is absent (`docs/architecture/10-observability-and-ops.md` §7.4).
|
| 661 |
+
|
| 662 |
+
> **Do not "fix" B-02 by editing the health payload on the live service.** The fix ships with the B-07
|
| 663 |
+
> patch, which is deliberately not deployed. A cosmetic newline is not worth a live-backend change on
|
| 664 |
+
> its own.
|
| 665 |
+
|
| 666 |
+
---
|
| 667 |
+
|
| 668 |
+
# Part V — Monitoring and alerting
|
| 669 |
+
|
| 670 |
+
## 15. What exists
|
| 671 |
+
|
| 672 |
+
SatQuery AI observes **exactly three things** (`docs/architecture/10-observability-and-ops.md` §6): whether
|
| 673 |
+
the transport is up, what one run did, and what failed inside the server.
|
| 674 |
+
|
| 675 |
+
### 15.1 The health payload
|
| 676 |
+
|
| 677 |
+
The measured live payload (`release/repo/docs/DEPLOYMENT.md` §5):
|
| 678 |
+
|
| 679 |
+
```json
|
| 680 |
+
{
|
| 681 |
+
"status": "ok",
|
| 682 |
+
"service": "satquery-orchestrator",
|
| 683 |
+
"tunnel": {
|
| 684 |
+
"agent_connected": true,
|
| 685 |
+
"agent_id": "codespaces-fd1038",
|
| 686 |
+
"pending": 0,
|
| 687 |
+
"completed": 97
|
| 688 |
+
},
|
| 689 |
+
"config": {
|
| 690 |
+
"codespace_name": "potential-space-trout-r4ppw969w45j2pvvw\n",
|
| 691 |
+
"codespace_port": 8000,
|
| 692 |
+
"transport_mode": "auto",
|
| 693 |
+
"tunnel_timeout_s": 150.0,
|
| 694 |
+
"wake_timeout_s": 120.0,
|
| 695 |
+
"upstream_timeout_s": 90.0,
|
| 696 |
+
"device": "cpu",
|
| 697 |
+
"has_github_token": true
|
| 698 |
+
}
|
| 699 |
+
}
|
| 700 |
+
```
|
| 701 |
+
|
| 702 |
+
Field by field, for the operator:
|
| 703 |
+
|
| 704 |
+
| Field | Meaning | Operational use |
|
| 705 |
+
|---|---|---|
|
| 706 |
+
| `status` | orchestrator liveness | a single up/down bit for the gateway tier |
|
| 707 |
+
| `tunnel.agent_connected` | an agent polled within the freshness window | **the** check before a demo (§6) |
|
| 708 |
+
| `tunnel.agent_id` | which agent | identifies the Codespace agent that is connected |
|
| 709 |
+
| `tunnel.pending` | in-flight tunnel requests | a rising value means work is queueing |
|
| 710 |
+
| `tunnel.completed` | a monotonic delivery counter, per Render process | trend only — it **resets on restart** |
|
| 711 |
+
| `config.codespace_name` | the target Codespace | carries B-02's trailing `\n` (§14) |
|
| 712 |
+
| `config.codespace_port` | the inference port | `8000` |
|
| 713 |
+
| `config.transport_mode` | `auto` / `tunnel` / `forward` | `auto` is what makes B-07 reachable (§13) |
|
| 714 |
+
| `config.tunnel_timeout_s` | the tunnel wait | `150.0` (§11) |
|
| 715 |
+
| `config.wake_timeout_s` | the cold-start wait | `120.0` (§11) |
|
| 716 |
+
| `config.upstream_timeout_s` | the proxy budget | `90.0` (§11) |
|
| 717 |
+
| `config.device` | device preference | `cpu` |
|
| 718 |
+
| `config.has_github_token` | whether a token is present (boolean only) | a `false` here means the wake path cannot start the Codespace |
|
| 719 |
+
|
| 720 |
+
> **`completed` was measured at three different values** across probes (`97`, `314`, and others). It is a
|
| 721 |
+
> counter that resets when the Render process restarts, not a constant. **Do not treat any single reading
|
| 722 |
+
> as the value** (`docs/architecture/10-observability-and-ops.md` §2.3;
|
| 723 |
+
> `docs/architecture/02-deployment-topology.md` §7.3).
|
| 724 |
+
|
| 725 |
+
### 15.2 The per-run `ExecutionTrace`
|
| 726 |
+
|
| 727 |
+
Every run carries an `ExecutionTrace` (`core/schemas.py:296-319`) with: `run_id`, `task`, `query`,
|
| 728 |
+
`inputs`, `modalities`, `intent`, `validation`, `workflow`, `steps`, `selected_models`, `parameters`,
|
| 729 |
+
`outputs`, `confidence`, `timings`, `fallbacks`, `errors`, `contradiction`, `config_hash`, `started_at`,
|
| 730 |
+
`finished_at`.
|
| 731 |
+
|
| 732 |
+
This is the **only** diagnostic surface for a *wrong but successful* answer — no log line records a
|
| 733 |
+
successful run (`docs/architecture/10-observability-and-ops.md` §7.6). The operator reads:
|
| 734 |
+
|
| 735 |
+
| Trace field | What it answers |
|
| 736 |
+
|---|---|
|
| 737 |
+
| `intent` | what the router thought the query meant |
|
| 738 |
+
| `task` | what was dispatched |
|
| 739 |
+
| `selected_models` | what actually ran |
|
| 740 |
+
| `errors` | what failed inside the run |
|
| 741 |
+
| `fallbacks` | what degraded |
|
| 742 |
+
| `config_hash` | which config produced the result |
|
| 743 |
+
|
| 744 |
+
### 15.3 The Codespace's own `/v1/health`
|
| 745 |
+
|
| 746 |
+
The inference tier answers its own health route, derived rather than asserted, with a torch-free device
|
| 747 |
+
probe (`app/space_app.py:521-547`; `core/schemas.py:430-437`; `docs/architecture/10-observability-and-ops.md`
|
| 748 |
+
§2.10). `gpu_available: false` is **expected** on a CPU host and must never be surfaced as a fault
|
| 749 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2).
|
| 750 |
+
|
| 751 |
+
## 16. What does **not** exist
|
| 752 |
+
|
| 753 |
+
This is the honest inventory. None of it is aspirational — the list exists so that a reader does not
|
| 754 |
+
assume a monitoring facility that was never built (`docs/architecture/10-observability-and-ops.md` §6):
|
| 755 |
+
|
| 756 |
+
| Capability | Present? | Evidence |
|
| 757 |
+
|---|---|---|
|
| 758 |
+
| **APM** (application performance monitoring) | **no** | no APM client, SDK or agent in any source read |
|
| 759 |
+
| **Distributed tracing** | **no** | no trace-context propagation; the three id namespaces do not join |
|
| 760 |
+
| **Cost accounting** | **no** | `GPU_DURATIONS` declares durations but nothing meters or reports consumption |
|
| 761 |
+
| **Metrics endpoint** (Prometheus / OpenMetrics) | **no** | no `/metrics` route in any app factory |
|
| 762 |
+
| **Per-model latency histogram** | **no** | `trace.timings` is per-run, per-step, never aggregated |
|
| 763 |
+
| **Error-rate counter** | **no** | no counter exists; `trace.errors` is per-run only |
|
| 764 |
+
| **Request counter** | **no** | the tunnel's `completed` counts only tunnel deliveries, per Render process |
|
| 765 |
+
| **Uptime / restart tracking** | **no** | no uptime field; the hub's counters reset on restart |
|
| 766 |
+
| **Alerting** | **no** | no alerting rule, webhook or threshold anywhere |
|
| 767 |
+
| **Structured / JSON logs** | **no** | all log calls use `%s`-style free text |
|
| 768 |
+
| **Log shipping / aggregation** | **no** | logs are per-host; the Codespace's are on an ephemeral filesystem |
|
| 769 |
+
| **Dashboards** | **no** | none exists |
|
| 770 |
+
| **SLO / SLA definition** | **no** | none exists |
|
| 771 |
+
| **A system-level end-to-end benchmark** | **no** | `DOCS_STYLE_GUIDE.md` §3: *"does not exist; no system-level accuracy is claimed"* |
|
| 772 |
+
|
| 773 |
+
**There is no pager, no alert, and no dashboard.** An operator learns the system is down by trying to use
|
| 774 |
+
it. This is stated as a fact, not a complaint.
|
| 775 |
+
|
| 776 |
+
## 17. Why "no alerting" is a design fact, not an oversight
|
| 777 |
+
|
| 778 |
+
The plan's §74 lists what is **deliberately not included**, and the monitoring gaps are downstream of
|
| 779 |
+
that list:
|
| 780 |
+
|
| 781 |
+
```
|
| 782 |
+
authentication
|
| 783 |
+
multi-tenant isolation
|
| 784 |
+
distributed queues
|
| 785 |
+
autoscaling
|
| 786 |
+
observability platform
|
| 787 |
+
Kubernetes
|
| 788 |
+
service mesh
|
| 789 |
+
distributed storage
|
| 790 |
+
horizontal worker orchestration
|
| 791 |
+
enterprise security
|
| 792 |
+
billing
|
| 793 |
+
SLA infrastructure
|
| 794 |
+
```
|
| 795 |
+
|
| 796 |
+
(`Implementation and Architecture plan.md` §74)
|
| 797 |
+
|
| 798 |
+
The architecture is described there as **scale-compatible, but not a production implementation**. The
|
| 799 |
+
absence of an observability platform, billing and SLA infrastructure is therefore intentional at this
|
| 800 |
+
stage. An operator should not expect — and must not claim — production-grade monitoring.
|
| 801 |
+
|
| 802 |
+
---
|
| 803 |
+
|
| 804 |
+
# Part VI — Capacity and cost
|
| 805 |
+
|
| 806 |
+
## 18. The capacity shape
|
| 807 |
+
|
| 808 |
+
Three facts bound capacity, and none of them is elastic:
|
| 809 |
+
|
| 810 |
+
| Property | Value | Evidence |
|
| 811 |
+
|---|---|---|
|
| 812 |
+
| Render plan | **free tier** — sleeps when idle | `render.yaml` (`plan: free`); `release/repo/docs/DEPLOYMENT.md` §8 |
|
| 813 |
+
| Inference hosts | **one** Codespace | `release/CURRENT_RELEASE_STATE.md` §1 |
|
| 814 |
+
| Device | **CPU-only** | `render.yaml`; `.devcontainer/devcontainer.json`; `docs/DEPLOYMENT_DECISION.md` §5 |
|
| 815 |
+
| Autoscaling | **absent** | plan §74 lists `autoscaling` under **Not included** |
|
| 816 |
+
| Horizontal workers | **absent** | plan §74 lists `horizontal worker orchestration` under **Not included** |
|
| 817 |
+
| Database / queue | **absent** | the gateway has *"no database, no auth, no queue"* (`deploy/render/main.py:4-6`) |
|
| 818 |
+
|
| 819 |
+
The consequence for an operator:
|
| 820 |
+
|
| 821 |
+
- **A single client can occupy the system.** The per-IP rate limit is a fairness control, **not** a
|
| 822 |
+
security control (§19.2). There is no queue to absorb a burst.
|
| 823 |
+
- **A restart is a full outage.** There is no replica to fail over to. The Codespace filesystem is
|
| 824 |
+
ephemeral (`deploy/codespace/launch.sh:44-47`), so a restart also empties the asset store.
|
| 825 |
+
- **Cold starts are unavoidable.** Render's free tier sleeps, so the first request after idle pays the
|
| 826 |
+
cold-start cost (§10).
|
| 827 |
+
|
| 828 |
+
## 19. The cost shape, and the one absence that has a code artifact
|
| 829 |
+
|
| 830 |
+
### 19.1 What is declared
|
| 831 |
+
|
| 832 |
+
`app/space_app.py` declares a per-task ZeroGPU duration, transcribed from the frozen config
|
| 833 |
+
(`app/space_app.py:105-116`, quoted in `docs/architecture/10-observability-and-ops.md` §6.1):
|
| 834 |
+
|
| 835 |
+
```python
|
| 836 |
+
GPU_DURATIONS: dict[str, int] = {
|
| 837 |
+
"vqa": 20,
|
| 838 |
+
"caption": 20,
|
| 839 |
+
"grounding": 45,
|
| 840 |
+
"change": 30,
|
| 841 |
+
"optical_sar": 45,
|
| 842 |
+
"change_vqa": 30,
|
| 843 |
+
}
|
| 844 |
+
```
|
| 845 |
+
|
| 846 |
+
These values are used **only** to decorate a handler with a ZeroGPU reservation. On the live **CPU**
|
| 847 |
+
deployment the decoration is a **no-op**: `_spaces_module()` returns `None` when the `spaces` package is
|
| 848 |
+
absent, so `decorate_gpu` returns the identity decorator (`app/space_app.py:143-166`).
|
| 849 |
+
|
| 850 |
+
> **Do not present `GPU_DURATIONS` as a cost model.** It is a declaration of intended reservation, and on
|
| 851 |
+
> the CPU deployment it reserves nothing (`docs/architecture/10-observability-and-ops.md` §6.1). The
|
| 852 |
+
> ZeroGPU 5-GPU-minute/day quota and the `@spaces.GPU` decoration are **frozen paperwork** — no Gradio
|
| 853 |
+
> runtime exists in code, and the manifest is left undisturbed because editing it would move the config
|
| 854 |
+
> hash (`release/repo/docs/DEPLOYMENT.md` §11).
|
| 855 |
+
|
| 856 |
+
### 19.2 What is **not** metered
|
| 857 |
+
|
| 858 |
+
Nothing reads `GPU_DURATIONS` back out to compute, record or report consumption. There is no per-run
|
| 859 |
+
GPU-second field, no cumulative counter, and no budget-exhaustion signal
|
| 860 |
+
(`docs/architecture/10-observability-and-ops.md` §6.1).
|
| 861 |
+
|
| 862 |
+
**No cost is observed.** There is no cost accounting for:
|
| 863 |
+
|
| 864 |
+
| Cost | Metered? | Note |
|
| 865 |
+
|---|---|---|
|
| 866 |
+
| Render compute | **no** | free tier; no usage field is read |
|
| 867 |
+
| Codespace compute | **no** | no usage field is read; the quota is a platform property |
|
| 868 |
+
| Hugging Face model hosting | **no** | no usage field is read |
|
| 869 |
+
| Model download volume | **no** | `warm_cache.py` reports per-model status but not bytes or cost |
|
| 870 |
+
| Per-request inference cost | **no** | no field exists |
|
| 871 |
+
|
| 872 |
+
The rate-limit and size-limit values are the implementation's choices, not the plan's, and the maintainer
|
| 873 |
+
should confirm them because they bound one client's share of capacity
|
| 874 |
+
(`docs/PHASE19_FINAL_HARDENING.md` §4.4).
|
| 875 |
+
|
| 876 |
+
### 19.3 The rate limit is fairness, not protection
|
| 877 |
+
|
| 878 |
+
The limiter keys on the first hop of `X-Forwarded-For`, which is **client-supplied**. A caller that varies
|
| 879 |
+
the header is never throttled. Measured in-process at 3 requests / 60 s, 8 requests sent: `5/8` throttled
|
| 880 |
+
without the header, **`0/8` with a fresh value per request** (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2).
|
| 881 |
+
|
| 882 |
+
> **Do not treat `SATQUERY_RATE_LIMIT_PER_IP` as protecting capacity.** It bounds accidental loops and
|
| 883 |
+
> honest clients. Fixing it correctly depends on how many proxy hops Render inserts, which must be
|
| 884 |
+
> measured on a deployed gateway — hard-coding a guess would replace a documented weakness with an
|
| 885 |
+
> undocumented one (`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2).
|
| 886 |
+
|
| 887 |
+
---
|
| 888 |
+
|
| 889 |
+
# Part VII — Incident triage
|
| 890 |
+
|
| 891 |
+
## 20. The symptom → cause → check → action table
|
| 892 |
+
|
| 893 |
+
One table, ordered by how often each symptom is seen. "First check" is the single cheapest command that
|
| 894 |
+
distinguishes the cases.
|
| 895 |
+
|
| 896 |
+
| Symptom | Likely cause | First check | Action |
|
| 897 |
+
|---|---|---|---|
|
| 898 |
+
| Request hangs ~249 s, then `504` | **B-07 tunnel gap** | re-read `tunnel.agent_connected` | retry; keep the Codespace warm (§13) |
|
| 899 |
+
| Request returns `503` quickly, `recoverable: true` | no agent connected and `mode == tunnel` | `GET /api/health` | start the Codespace (§7) |
|
| 900 |
+
| `GET /api/health` unreachable | Render sleeping or down | re-issue the request | the first request wakes Render; retry |
|
| 901 |
+
| `GET /api/health` OK but `agent_connected: false` | Codespace stopped, or the agent was reaped | — | start the Codespace; confirm `postStartCommand` ran (§7) |
|
| 902 |
+
| `504 wake_timeout` | the Codespace did not come up within 120 s | `tunnel.agent_connected` | the Codespace is not coming up; check it directly |
|
| 903 |
+
| `502 upstream_unreachable` | a connection error to the Codespace | `tunnel.agent_connected` | check the transport |
|
| 904 |
+
| `422 invalid_request` | the body did not match `AnalysisRequest` | the client's request body | fix the client; **not** a server fault |
|
| 905 |
+
| `503 model_unavailable` | an artifact is absent | `GET /api/capabilities` reasons | expected if not uploaded; ship degraded or upload it |
|
| 906 |
+
| `503 model_load_error` | an artifact is present but corrupt | the capability's `reason` | **a defect** — replace the artifact and report |
|
| 907 |
+
| `confidence.method: "uncalibrated"` | the calibration artifact was not uploaded | the run's trace | honest, not broken |
|
| 908 |
+
| `status: "degraded"` on `/v1/health` | at least one capability is not servable | `GET /api/capabilities` | not an error; the service is up |
|
| 909 |
+
| `gpu_available: false` | CPU host | — | expected; never surface as a fault |
|
| 910 |
+
| a handle that worked now returns `400 input_error` | the handle lapsed, or the Codespace restarted and its asset dir is ephemeral | re-upload | do not assume handles persist |
|
| 911 |
+
| a capability `available: true` **with** a non-null `reason` | hub-backed, no local checkpoint — the first call will be slow | the reason string | not a fault; do not surface as an error |
|
| 912 |
+
| a **wrong but successful** answer | a router/dispatch/artifact issue | the run's `ExecutionTrace` | read `intent`, `task`, `selected_models`, `errors`, `fallbacks` (§15.2) |
|
| 913 |
+
|
| 914 |
+
The first four rows are the ones an operator hits in practice. Rows 8–15 are from
|
| 915 |
+
`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2, which is the canonical degraded-state reference.
|
| 916 |
+
|
| 917 |
+
## 21. The consolidated decision tree
|
| 918 |
+
|
| 919 |
+
For the case where something is wrong and the operator has only the client-side report
|
| 920 |
+
(`docs/architecture/10-observability-and-ops.md` §7.6):
|
| 921 |
+
|
| 922 |
+
```mermaid
|
| 923 |
+
flowchart TD
|
| 924 |
+
A["something is wrong"] --> B["GET /api/health"]
|
| 925 |
+
B --> C{"tunnel.agent_connected?"}
|
| 926 |
+
C -->|"false"| D["warm the stack<br/>§6 steps 1-2"]
|
| 927 |
+
C -->|"true"| E{"what did the client see?"}
|
| 928 |
+
|
| 929 |
+
E -->|"nothing, request hung"| F{"waited ≈249 s?"}
|
| 930 |
+
F -->|"yes"| G["B-07 tunnel gap<br/>retry + warm<br/>§13"]
|
| 931 |
+
F -->|"no"| H["still waiting —<br/>within budget"]
|
| 932 |
+
|
| 933 |
+
E -->|"a 5xx"| I["read the error code<br/>§22"]
|
| 934 |
+
E -->|"a 4xx"| J["a client bug —<br/>the body did not match the contract"]
|
| 935 |
+
|
| 936 |
+
I --> K{"code present in the<br/>DEPLOYED revision?"}
|
| 937 |
+
K -->|"no"| L["the patch is not deployed<br/>§13"]
|
| 938 |
+
K -->|"yes"| M["act on the code"]
|
| 939 |
+
|
| 940 |
+
E -->|"a result, but wrong"| N["read trace:<br/>intent, task, selected_models,<br/>errors, fallbacks<br/>§15.2"]
|
| 941 |
+
```
|
| 942 |
+
|
| 943 |
+
## 22. The machine-code reference
|
| 944 |
+
|
| 945 |
+
| Code | Status | Meaning | Action |
|
| 946 |
+
|---|---|---|---|
|
| 947 |
+
| `tunnel_offline` | 503 | no agent is connected and `mode == "tunnel"` | start the Codespace |
|
| 948 |
+
| `forward_unavailable` | 503 | the forwarded port is not anonymously reachable | start the tunnel agent (**patch only**) |
|
| 949 |
+
| `wake_timeout` | 504 | the wake loop exhausted `wake_timeout_s` | the Codespace is not coming up |
|
| 950 |
+
| `upstream_timeout` | 504 | the tunnel was healthy but no agent completed in `tunnel_timeout_s` | retry; check the agent (**patch only**) |
|
| 951 |
+
| `upstream_unreachable` | 502 | a connection error to the Codespace | check the transport |
|
| 952 |
+
| `invalid_request` | 422 | the body was not a valid `AnalysisRequest` | fix the client |
|
| 953 |
+
| `model_unavailable` | 502 | the gateway could not reach the analysis service | retry |
|
| 954 |
+
| `orchestrator_config_error` | 500 | missing `GITHUB_TOKEN` or `CODESPACE_NAME` | fix the env vars (`deploy/render/main.py:256-262`) |
|
| 955 |
+
| `upstream_error` | 502 | a non-connection httpx error | check the upstream (`deploy/render/main.py:392-400`) |
|
| 956 |
+
| `schema_validation_error` | 502 | a non-JSON upstream body | check the upstream (`deploy/render/main.py:404-414`) |
|
| 957 |
+
|
| 958 |
+
(`deploy/render/main.py:247-291,383-414`; the full taxonomy is in
|
| 959 |
+
[08 — The API Contract](architecture/08-api-contract.md) §12.)
|
| 960 |
+
|
| 961 |
+
> **The two codes marked "patch only" do not exist in the deployed revision.** An operator on the
|
| 962 |
+
> deployed system will not see `forward_unavailable` or `upstream_timeout`; they will see `wake_timeout`
|
| 963 |
+
> after ≈249 s instead (§13).
|
| 964 |
+
|
| 965 |
+
---
|
| 966 |
+
|
| 967 |
+
# Part VIII — Routine maintenance
|
| 968 |
+
|
| 969 |
+
## 23. Rotating credentials (procedure only)
|
| 970 |
+
|
| 971 |
+
> **This document records no credential value, and no path to a credential file.** The repository's own
|
| 972 |
+
> evidence index records where credentials are held (`release/CURRENT_RELEASE_STATE.md` §7) by
|
| 973 |
+
> **location and kind only**; that index is not reproduced here. The procedure below is deliberately
|
| 974 |
+
> value-free.
|
| 975 |
+
|
| 976 |
+
Four credential purposes exist, and each is rotated by changing a value the operator holds — never a
|
| 977 |
+
value in this document:
|
| 978 |
+
|
| 979 |
+
| Purpose | Where it is consumed | What rotation changes |
|
| 980 |
+
|---|---|---|
|
| 981 |
+
| Codespace control (wake) | the Render env var `GITHUB_TOKEN` | the token the orchestrator uses to `GET`/`POST .../start` a Codespace (`deploy/render/codespaces.py:65-70`) |
|
| 982 |
+
| Repository writes | the operator's local tooling | the token used for the Git Data API deploy path (`release/repo/docs/DEPLOYMENT.md` §7.1) |
|
| 983 |
+
| Codespace account access | the GitHub account session | the account credential used to open the Codespace |
|
| 984 |
+
| Hugging Face model access | the Hub download path | the token used to resolve pinned model revisions |
|
| 985 |
+
|
| 986 |
+
**The procedure, in order:**
|
| 987 |
+
|
| 988 |
+
1. **Create the replacement credential** in the provider's UI, with the minimum scope the purpose needs.
|
| 989 |
+
For the wake path the scope is `codespace` (`deploy/render/codespaces.py:50`, `:90-92`).
|
| 990 |
+
2. **Update the consumer's configuration.** For the wake path this is the Render environment variable
|
| 991 |
+
`GITHUB_TOKEN` (`render.yaml:14-15`; `release/repo/docs/DEPLOYMENT.md` §6.1).
|
| 992 |
+
3. **Redeploy / restart the consumer** so it reads the new value. The orchestrator reads its environment
|
| 993 |
+
once at process start, so a value change without a restart is invisible.
|
| 994 |
+
4. **Verify.** Call `GET /api/health` and confirm `config.has_github_token: true`
|
| 995 |
+
(`deploy/render/main.py:454`). Then trigger one wake to confirm the token is accepted end to end.
|
| 996 |
+
5. **Revoke the old credential** at the provider, once step 4 passes.
|
| 997 |
+
|
| 998 |
+
**Two rules that must not be broken:**
|
| 999 |
+
|
| 1000 |
+
- **Never set a token variable to an empty string.** An empty `HF_TOKEN` produced
|
| 1001 |
+
`Authorization: Bearer `, which httpx rejects with a `LocalProtocolError` that is misreported as an
|
| 1002 |
+
upstream failure. **Omit the variable entirely rather than setting it empty**
|
| 1003 |
+
(`docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1). The same reasoning applies to `GITHUB_TOKEN`: the
|
| 1004 |
+
orchestrator raises a clear `orchestrator_config_error` when it is absent
|
| 1005 |
+
(`deploy/render/main.py:90-96`), which is better than a confusing transport failure.
|
| 1006 |
+
- **Never commit a credential.** The Render variables are `sync: false` in `render.yaml` and must be set
|
| 1007 |
+
in the dashboard (`deploy/render/README.md` §Environment variables).
|
| 1008 |
+
|
| 1009 |
+
> **`has_github_token` is a boolean, never the value.** The health payload reports only whether a token is
|
| 1010 |
+
> present (`deploy/render/main.py:454`), so a health probe is a safe way to confirm rotation without
|
| 1011 |
+
> exposing the credential.
|
| 1012 |
+
|
| 1013 |
+
## 24. Restarting the tunnel agent
|
| 1014 |
+
|
| 1015 |
+
The agent is supervised and should not need a manual restart, but the procedure is:
|
| 1016 |
+
|
| 1017 |
+
1. **Check whether it is running** — the Codespace-side check is
|
| 1018 |
+
`pgrep -f "deploy/codespace/tunnel_agent.py"` (`deploy/codespace/launch.sh:153,179`).
|
| 1019 |
+
2. **Re-run the launcher** — `bash deploy/codespace/launch.sh`. It is **guarded** on the process, so
|
| 1020 |
+
re-running is a no-op if the agent is already up (`deploy/codespace/launch.sh:24,153-155`).
|
| 1021 |
+
3. **Confirm the reconnect** — read the agent log for the announce line
|
| 1022 |
+
(`grep -q "announced to hub" "$TUNNEL_LOG"`, `deploy/codespace/launch.sh:185`) and then confirm
|
| 1023 |
+
`tunnel.agent_connected: true` from outside (§6 step 1).
|
| 1024 |
+
|
| 1025 |
+
> **The launcher restarts the serve process too, if it is stale** (`deploy/codespace/launch.sh:100-147`).
|
| 1026 |
+
> That is intentional: a stale server answering from old code is worse than a restart (§7).
|
| 1027 |
+
|
| 1028 |
+
> **Caveat.** The agent log path is `/tmp/satquery-tunnel.log` and the serve log is
|
| 1029 |
+
> `/tmp/satquery-serve.log` (`deploy/codespace/launch.sh:63-64`). Both are on the **ephemeral** Codespace
|
| 1030 |
+
> filesystem, so they vanish with the Codespace (`deploy/codespace/launch.sh:44-47`).
|
| 1031 |
+
|
| 1032 |
+
## 25. Re-warming the model cache
|
| 1033 |
+
|
| 1034 |
+
Re-warm after a Codespace rebuild, or whenever the first analysis is slower than expected:
|
| 1035 |
+
|
| 1036 |
+
```bash
|
| 1037 |
+
python deploy/codespace/warm_cache.py
|
| 1038 |
+
```
|
| 1039 |
+
|
| 1040 |
+
`post_create.sh` runs this automatically on container **creation** (`deploy/codespace/post_create.sh:5-6`),
|
| 1041 |
+
and it is wired as the devcontainer `postCreateCommand` (`.devcontainer/devcontainer.json:17`).
|
| 1042 |
+
|
| 1043 |
+
> **Creation vs. start.** `postCreateCommand` runs only when the container is **created**, whereas
|
| 1044 |
+
> `postStartCommand` runs on every start (`.devcontainer/devcontainer.json:17-18`). The `containerEnv`
|
| 1045 |
+
> block is likewise applied only at creation, which is why `launch.sh` re-exports the asset-upload
|
| 1046 |
+
> variables on every start — *"`containerEnv` is only applied when the container is CREATED"*
|
| 1047 |
+
> (`deploy/codespace/launch.sh:49-52`). An operator who changes an env var must restart, not just reload.
|
| 1048 |
+
|
| 1049 |
+
## 26. Deploying a change (the Git Data API path)
|
| 1050 |
+
|
| 1051 |
+
Deployment does **not** use `git push`. Every deployed file is uploaded as a **blob** whose sha256 is
|
| 1052 |
+
computed locally and verified against the uploaded blob, then assembled into a tree, committed, and the
|
| 1053 |
+
branch ref patched (`release/repo/docs/DEPLOYMENT.md` §7.1).
|
| 1054 |
+
|
| 1055 |
+
Why this matters operationally:
|
| 1056 |
+
|
| 1057 |
+
- each file is **content-verified** rather than trusted;
|
| 1058 |
+
- deletions are expressed explicitly as `sha: null` tree entries;
|
| 1059 |
+
- the deploy is **idempotent** — re-running it with identical content produces no change.
|
| 1060 |
+
|
| 1061 |
+
**Measured:** 9 deployed files were re-read from the API and found **sha256 byte-identical** to the local
|
| 1062 |
+
copies, with the deployed HEAD re-read independently (`verify_deployed_head.py`,
|
| 1063 |
+
`release/CURRENT_RELEASE_STATE.md` §5).
|
| 1064 |
+
|
| 1065 |
+
**Frontend deploy** (Cloudflare Pages) uses a staging step, not the Git Data API:
|
| 1066 |
+
|
| 1067 |
+
```bash
|
| 1068 |
+
node scripts/stage_pages.mjs \
|
| 1069 |
+
--out=.deploy/dist-final \
|
| 1070 |
+
--include=_headers \
|
| 1071 |
+
--include=robots.txt \
|
| 1072 |
+
--include=assets/img/eo/provenance.json \
|
| 1073 |
+
--include=assets/img/eo/CREDITS.md
|
| 1074 |
+
|
| 1075 |
+
npx wrangler pages deploy "C:/Users/anish/satquery-ai/.deploy/dist-final" --project-name <name>
|
| 1076 |
+
```
|
| 1077 |
+
|
| 1078 |
+
(`docs/DEPLOYMENT_DECISION.md` §7)
|
| 1079 |
+
|
| 1080 |
+
`_headers` and `robots.txt` must be **force-included** because no page references them; `provenance.json`
|
| 1081 |
+
and `CREDITS.md` likewise (`docs/DEPLOYMENT_DECISION.md` §7).
|
| 1082 |
+
|
| 1083 |
+
---
|
| 1084 |
+
|
| 1085 |
+
# Part IX — Known operational gaps
|
| 1086 |
+
|
| 1087 |
+
## 27. The explicit gaps list
|
| 1088 |
+
|
| 1089 |
+
Each row is something an operator might reasonably expect and that does **not** exist. None is a
|
| 1090 |
+
regression; each is a boundary of the current release.
|
| 1091 |
+
|
| 1092 |
+
| # | Gap | Consequence | Status |
|
| 1093 |
+
|---|---|---|---|
|
| 1094 |
+
| 1 | **No alerting** of any kind | an outage is discovered by trying to use the system | **not implemented** (§16) |
|
| 1095 |
+
| 2 | **No APM / metrics / distributed tracing** | no latency, error-rate or throughput trend exists | **not implemented** |
|
| 1096 |
+
| 3 | **No cost accounting** | consumption is unmeasured | **not implemented** (§19.2) |
|
| 1097 |
+
| 4 | **No dashboard / SLO / SLA** | no shared view of health; no target defined | **not implemented** |
|
| 1098 |
+
| 5 | **No structured logs / log shipping** | logs are per-host free text; the Codespace's are ephemeral | **not implemented** |
|
| 1099 |
+
| 6 | **B-07 is unfixed in production** | a request can hang ≈249 s then `504` | **`OPEN`** (§13) |
|
| 1100 |
+
| 7 | **B-02 trailing `\n`** | cosmetic; a wrong-looking field in health | **`OPEN` (cosmetic)** (§14) |
|
| 1101 |
+
| 8 | **No autoscaling, no replicas** | a restart is a full outage | **BY DESIGN** (plan §74) |
|
| 1102 |
+
| 9 | **No database, queue or persistence** | no run history survives a restart | **BY DESIGN** (`deploy/render/main.py:4-6`) |
|
| 1103 |
+
| 10 | **One Codespace** | capacity is bounded by one CPU host | **BY DESIGN** (§18) |
|
| 1104 |
+
| 11 | **No auth** | the contract documents *"no auth in v1"*; paths are scrubbed from client-visible fields as a partial mitigation | **BY DESIGN** |
|
| 1105 |
+
| 12 | **A system-level E2E benchmark does not exist** | no single system accuracy number can be quoted | **NOT RUN** |
|
| 1106 |
+
| 13 | **No measured cold-start distribution** | only *"tens of seconds"* is documented | **`UNKNOWN`** |
|
| 1107 |
+
| 14 | **The effective log level / destination per host** | not established | **`UNKNOWN`** |
|
| 1108 |
+
| 15 | **Whether `doctor.sh` / `tunnel_agent.py` exist in the deployed repo** | the monorepo copy is stale and lacks them | **`UNKNOWN`** (§7) |
|
| 1109 |
+
| 16 | **The B-07 patch is not deployed** | the fast-fail codes are absent in production | **`OPEN`** (§13) |
|
| 1110 |
+
|
| 1111 |
+
## 28. `NOT RUN` / `OPEN` / `BLOCKED` / `UNKNOWN` for operations
|
| 1112 |
+
|
| 1113 |
+
| # | Item | Status |
|
| 1114 |
+
|---|---|---|
|
| 1115 |
+
| 1 | B-07 — tunnel gaps; patch prepared, **not deployed** | **`OPEN`** |
|
| 1116 |
+
| 2 | B-02 — `/api/health` `codespace_name` trailing `\n` | **`OPEN` (cosmetic)** |
|
| 1117 |
+
| 3 | A deployed system-level load test | **`NOT RUN`** |
|
| 1118 |
+
| 4 | A measured cold-start distribution | **`NOT RUN`** — only "tens of seconds" is documented |
|
| 1119 |
+
| 5 | Multi-region / HA deployment | **`NOT RUN`** |
|
| 1120 |
+
| 6 | A system-level end-to-end benchmark | **`NOT RUN`** — none exists |
|
| 1121 |
+
| 7 | Any APM / metrics / distributed tracing / alerting | **not implemented** |
|
| 1122 |
+
| 8 | Cost accounting | **not implemented** — `GPU_DURATIONS` is declared, not metered |
|
| 1123 |
+
| 9 | The effective log level and destination on each host | **`UNKNOWN`** |
|
| 1124 |
+
| 10 | Whether the deployed `SatQuery-Inference` repo carries `doctor.sh` / `tunnel_agent.py` | **`UNKNOWN`** |
|
| 1125 |
+
| 11 | Log retention on Render | **`UNKNOWN`** — a platform property, not observable from the code |
|
| 1126 |
+
| 12 | The ZeroGPU/Gradio deployment target | **`REJECTED`** (superseded; frozen paperwork only) |
|
| 1127 |
+
| 13 | The five historical backend blockers | **closed by construction, not proven in production** |
|
| 1128 |
+
|
| 1129 |
+
> Rows 9 and 11 are copied from `docs/architecture/10-observability-and-ops.md` §8.1 so that the two
|
| 1130 |
+
> documents cannot drift. Row 13 is the honest framing from `release/repo/docs/DEPLOYMENT.md` §9: the
|
| 1131 |
+
> design closes the blockers, and the first live run is what would *verify* them.
|
| 1132 |
+
|
| 1133 |
+
---
|
| 1134 |
+
|
| 1135 |
+
# Part X — Evidence
|
| 1136 |
+
|
| 1137 |
+
## 29. Where the evidence lives
|
| 1138 |
+
|
| 1139 |
+
| What | Where |
|
| 1140 |
+
|---|---|
|
| 1141 |
+
| the live health payload | `release/CURRENT_RELEASE_STATE.md` §1; `release/repo/docs/DEPLOYMENT.md` §5 |
|
| 1142 |
+
| the live capability contract | `release/CURRENT_RELEASE_STATE.md` §1 |
|
| 1143 |
+
| the deployed revisions | `release/CURRENT_RELEASE_STATE.md` §1; `release/repo/docs/DEPLOYMENT.md` §1 |
|
| 1144 |
+
| the B-07 root shape and ≈249 s | `release/CURRENT_RELEASE_STATE.md` §6; `release/repo/docs/DEPLOYMENT.md` §8.1 |
|
| 1145 |
+
| the undeployed B-07 patch | session scratch: `fix-b07-forward-unavailable.patch` |
|
| 1146 |
+
| B-02's status and witness role | `release/repo/docs/DEPLOYMENT.md` §5; `docs/architecture/10-observability-and-ops.md` §7.4 |
|
| 1147 |
+
| the cold-start shape and the wake loop | `deploy/render/main.py:76-78,299-356,486-489` |
|
| 1148 |
+
| the Codespace launcher | `deploy/codespace/launch.sh`; `.devcontainer/devcontainer.json:18` |
|
| 1149 |
+
| the warm-up contract | `deploy/codespace/warm_cache.py` |
|
| 1150 |
+
| the asset-upload environment | `deploy/codespace/launch.sh:36-54` |
|
| 1151 |
+
| the deploy mechanics (Git Data API) | `release/repo/docs/DEPLOYMENT.md` §7 |
|
| 1152 |
+
| the platform traps | `release/repo/docs/DEPLOYMENT.md` §10; `release/repo/docs/REPRODUCIBILITY.md` §10 |
|
| 1153 |
+
| the historical backend blockers | `docs/DEPLOYMENT_DECISION.md` §8; `release/repo/docs/DEPLOYMENT.md` §9 |
|
| 1154 |
+
| the frozen config and hash | `configs/base.yaml`; `core/config.py:76-80` |
|
| 1155 |
+
| the observability inventory | `docs/architecture/10-observability-and-ops.md` §6 |
|
| 1156 |
+
| the triage tables | `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §5.2, §7 |
|
| 1157 |
+
| the live validation (3 passes, 24 runs) | `.workbuddy-ai/scratch/live_validation/` |
|
| 1158 |
+
| the rate-limit finding (F-5) | `docs/BACKEND_DEPLOYMENT_RUNBOOK.md` §4.1.2 |
|
| 1159 |
+
|
| 1160 |
+
### Cross-references
|
| 1161 |
+
|
| 1162 |
+
| For | See |
|
| 1163 |
+
|---|---|
|
| 1164 |
+
| the four tiers, the tunnel, the wake flow, `transport_mode` | [02 — Deployment Topology](architecture/02-deployment-topology.md) |
|
| 1165 |
+
| health, counters, traces, what is and is not observed | [10 — Observability and Operations](architecture/10-observability-and-ops.md) |
|
| 1166 |
+
| the four endpoints, the envelopes, the error taxonomy | [08 — The API Contract](architecture/08-api-contract.md) |
|
| 1167 |
+
| the request lifecycle and the nine-state spine in motion | [03 — Request Lifecycle](architecture/03-request-lifecycle.md) |
|
| 1168 |
+
| the frozen config and `Config.hash == 78f1e3700da15aa1` | [07 — Configuration and Freeze](architecture/07-configuration-freeze.md) |
|
| 1169 |
+
| live revisions, env vars, deploy mechanics, platform traps | [../DEPLOYMENT.md](DEPLOYMENT.md) |
|
| 1170 |
+
| what a third party can and cannot reproduce | [../REPRODUCIBILITY.md](REPRODUCIBILITY.md) |
|
| 1171 |
+
| how to build, test and extend the codebase | [../DEVELOPMENT.md](DEVELOPMENT.md) |
|
| 1172 |
+
|
| 1173 |
+
---
|
| 1174 |
+
|
| 1175 |
+
> **Chapter summary.** SatQuery AI runs four tiers — a static frontend, a thin Render orchestrator, one
|
| 1176 |
+
> CPU Codespace reached over an outbound tunnel, and a Hugging Face model tier — with exactly one
|
| 1177 |
+
> inference host and no replicas. The operator's single most important check is
|
| 1178 |
+
> `GET /api/health` → `tunnel.agent_connected: true`; the single most important diagnostic is the
|
| 1179 |
+
> ≈249 s-then-`504` signature of a B-07 tunnel gap. **B-07 is `OPEN` and the patch is not deployed; B-02
|
| 1180 |
+
> is `OPEN` and cosmetic.** The system observes three things — the transport, one run's trace, and the
|
| 1181 |
+
> server logs — and has **no** alerting, APM, distributed tracing or cost accounting. Capacity is one
|
| 1182 |
+
> free-tier Render service and one CPU Codespace; nothing is metered. Four things are genuinely
|
| 1183 |
+
> `UNKNOWN — not established from the available evidence`: the effective log level and destination per
|
| 1184 |
+
> host, log retention on Render, whether the deployed inference repository carries the tools the monorepo
|
| 1185 |
+
> launcher references, and any measured cold-start distribution.
|