Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
| # 02 — Deployment Topology | |
| **Parent:** [Architecture hub](README.md) · **Sibling:** [01 System overview](01-system-overview.md) · | |
| **Next:** [03 Request lifecycle](03-request-lifecycle.md) | |
| **Status tags used in this document:** `IMPLEMENTED` · `VERIFIED` · `MEASURED` · `ATTEMPTED` · | |
| `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` · `OPEN` · `RESOLVED` · `CLOSED`. | |
| > **One-paragraph summary.** SatQuery AI is deployed as **four tiers**: a static browser client on | |
| > **Cloudflare Pages** (`satquery.pages.dev`), a thin stateless **gateway/orchestrator on Render** | |
| > (`satquery-orchestrator` at `<backend-host>`), a **FastAPI inference service inside | |
| > a GitHub Codespace** (FastAPI on port `8000`), and **Hugging Face** as the project/model-presence tier. | |
| > The gateway does not dial into the Codespace. Instead the Codespace **dials out** to the gateway over a | |
| > long-poll **tunnel** (`POST /tunnel/agent`), because a *forwarded* Codespace port returns **HTTP 302** | |
| > for a private repository. That inversion is the single most consequential decision in the topology, and | |
| > it is why the deployment works at all with private repositories. | |
| --- | |
| ## 1. The four tiers | |
| ### 1.1 Tier map | |
| ``` | |
| USER / BROWSER | |
| │ HTTPS | |
| ▼ | |
| Cloudflare Pages — static frontend satquery.pages.dev | |
| │ HTTPS, JSON, /api/* | |
| ▼ | |
| Render — orchestrator / API gateway <backend-host> | |
| │ service: satquery-orchestrator | |
| │ outbound long-poll (POST /tunnel/agent) ← direction is INVERTED | |
| ▼ | |
| GitHub Codespace — FastAPI inference potential-space-trout-r4ppw969w45j2pvvw :8000 | |
| │ build_space_app(): /v1/health · /v1/capabilities · /v1/analyze · /v1/assets | |
| │ specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet | |
| ▼ | |
| Hugging Face — project card + pinned model references | |
| ``` | |
| Sources: `docs/DEPLOYMENT_TOPOLOGY.md` §1; `docs/FINAL_DELIVERY_TODO.md` §1.2 (the same ASCII topology | |
| reproduced in the delivery single-source-of-truth); `docs/FINAL_DELIVERY_REPORT.md` §2. | |
| ```mermaid | |
| flowchart LR | |
| U["Browser<br/>satquery.pages.dev"] -->|"HTTPS"| CF["Cloudflare Pages<br/>static frontend"] | |
| CF -->|"HTTPS JSON /api/*"| R["Render<br/>satquery-orchestrator"] | |
| R -->|"POST /tunnel/agent<br/>long-poll (outbound)"| A["Codespace tunnel agent"] | |
| A -->|"http://127.0.0.1:8000"| I["FastAPI<br/>build_space_app()"] | |
| I --> S[("SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet")] | |
| I -.->|"model refs"| HF["Hugging Face<br/>project + pinned models"] | |
| R -.->|"HF proxy path<br/>(not used in live config)"| HF | |
| ``` | |
| ### 1.2 Tier responsibilities, at a glance | |
| | Tier | Host / identity | Runs | Holds secrets? | Holds state? | | |
| |---|---|---|---|---| | |
| | Browser | the user's machine | `frontend/` static JS | **No** — never | no | | |
| | Static | Cloudflare Pages, `satquery.pages.dev` | HTML/CSS/JS only | **No** | no | | |
| | Gateway | Render, `satquery-orchestrator`, `<backend-host>` | `deploy/render/main.py` (`SatQuery-Backend` in production) | **Yes** — `GITHUB_TOKEN` (and `HF_TOKEN` if the proxy path is used) | no | | |
| | Inference | GitHub Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port `8000` | `app/space_app.py::build_space_app()` via `deploy/codespace/serve.py` | no gateway secrets | ephemeral asset store only | | |
| | Presence | Hugging Face (`hf/`) | project README / model cards | no | no | | |
| Sources: `docs/DEPLOYMENT_TOPOLOGY.md` §3.1–§3.4; `render.yaml`; `deploy/render/main.py`; | |
| `app/space_app.py`; `.devcontainer/devcontainer.json`. | |
| ### 1.3 Why exactly four tiers and not three | |
| The plan forbids "unnecessary microservices" (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1, §6, quoting plan | |
| §73/§74). A gateway is nevertheless present, and `docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 gives three | |
| concrete reasons rather than an architectural preference: | |
| 1. **The inference host cannot hold the security boundary.** It is a public ASGI app on third-party | |
| infrastructure. Rate limiting, size caps, CORS and secret custody belong outside it | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 1). | |
| 2. **A request that can be rejected on shape must never reach inference.** In the original design the | |
| scarce resource was the ZeroGPU `5 GPU-minutes/day` budget; in the active design it is inference wall | |
| time on a CPU Codespace. Either way the gateway is where a malformed request dies cheaply | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 2). | |
| 3. **The plan's §74 boundary excludes auth, multi-tenancy and queues.** So the gateway is a *proxy with | |
| validation*, and must not grow into a platform (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 item 3, §2.2). | |
| The conclusion recorded in the source document is **"two services, not three"** — a static client, a | |
| gateway, and one inference service (`docs/DEPLOYMENT_ARCHITECTURE.md` §1.1). Hugging Face is a presence | |
| tier, not a runtime tier, in the active design. | |
| --- | |
| ## 2. Why a gateway exists | |
| This section is the load-bearing one. A reader who understands only one part of the deployment should | |
| understand this: **the gateway is not there to compute anything. It is there to be the boundary.** | |
| ### 2.1 The responsibility table (authoritative) | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 is the authoritative statement. Reproduced with the active | |
| host name substituted (`Railway` → `Render`): | |
| | Responsibility | Detail | Why it must be here | | |
| |---|---|---| | |
| | **Schema validation** | reject malformed bodies with the §5 error envelope | avoids spending inference on a request that will fail | | |
| | **Size limits** | per-request body cap *and* per-file cap | the inference host cannot refuse a body it has already received | | |
| | **Rate limiting** | per-IP count + window | back-pressure against accidental loops; **fairness, not security** — see §2.4 | | |
| | **CORS** | explicit allowlist of the frontend origin | **never `*`** | | |
| | **Request IDs** | generate, inject, echo `X-Request-Id` | correlation across two services | | |
| | **Timeouts** | upstream timeout **shorter** than the inference host's own budget | prevents a hung proxy holding a connection | | |
| | **Secret custody** | `GITHUB_TOKEN` (and `HF_TOKEN` if used) live here only | the browser never sees them | | |
| | **Error translation** | inference errors → the documented envelope | the error contract is a gateway product | | |
| | **Body relaying for upload** | read and forward the raw body for `POST /v1/assets` | the upload path is not JSON-shaped, so JSON-oriented handling does not apply | | |
| ### 2.2 The CORS allowlist is explicit, and never a wildcard | |
| The orchestrator's CORS list is assembled by `_allowed_origins()` in `deploy/render/main.py`, in a | |
| documented order: | |
| 1. `SATQUERY_ALLOWED_ORIGINS` — the operator's comma-separated list. The authoritative source for any | |
| additional deployment origin. | |
| 2. `_PRODUCTION_ORIGINS` — `("https://satquery.pages.dev",)`, **always present**, so a deployment that | |
| forgets the environment variable still serves the real frontend. `deploy/render/main.py` records the | |
| reasoning: *"an empty allowlist would otherwise take the live site down, which is a worse failure than | |
| the one this guards."* | |
| 3. `_DEV_ORIGINS` — 20 enumerated `host:port` pairs (10 ports × `localhost`/`127.0.0.1`), added unless | |
| `SATQUERY_ALLOW_DEV_ORIGINS` is set to `0`/`false`/`no`/`""`. | |
| The dev-origin list is **enumerated, not a regex and not a suffix match** (`deploy/render/main.py`): | |
| ```python | |
| _DEV_ORIGINS: tuple[str, ...] = tuple( | |
| f"http://{host}:{port}" | |
| for host in ("localhost", "127.0.0.1") | |
| for port in ("3000", "5500", "5173", "8000", "8080") | |
| ) | |
| ``` | |
| A wildcard is refused in **two** places, deliberately: | |
| * `_allowed_origins()` raises `ValueError` if `"*"` appears in the assembled list, and its docstring | |
| records why the check exists there as well as in the config validator: *"this function cannot be the | |
| way a `*` reaches `CORSMiddleware`, which does not run that validator."* | |
| * `GatewayConfig.__post_init__` (`gateway/policy.py`) refuses a wildcard at construction, so a | |
| misconfiguration fails at startup rather than on the first request. | |
| `allow_credentials=False` is set explicitly in `create_app()` (`deploy/render/main.py`), matching the | |
| contract's "no auth, no cookies" position (`docs/API_CONTRACT.md` §7; plan §74). | |
| The gateway also **strips CORS headers coming back from upstream**, so the CORS answer is the gateway's | |
| alone. `gateway/app.py::_proxy` asserts this rather than trusting it: | |
| ```python | |
| assert not any(_is_cors_header(k) for k in out_headers) or decision.headers, ( | |
| "a CORS header reached the response without a policy decision; the " | |
| "upstream's headers are no longer filtered (see F-2)" | |
| ) | |
| ``` | |
| ### 2.3 Size limits: two caps, both enforced twice, on purpose | |
| Two independent caps exist, and they are different numbers with different jobs | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` §4): | |
| | Cap | Default | Scope | Where read | | |
| |---|---|---|---| | |
| | `SATQUERY_MAX_FILE_BYTES` | `4 * 1024 * 1024` = **4,194,304 bytes** | one uploaded file | **both** layers, from one variable | | |
| | `SATQUERY_MAX_BODY_BYTES` | `8 * 1024 * 1024` | the whole request body | gateway | | |
| `gateway/policy.py:221` declares `max_file_bytes: int = 4 * 1024 * 1024`; | |
| `app/space_app.py::_asset_max_file_bytes()` returns `4 * 1024 * 1024` when the variable is unset. The | |
| per-file cap is deliberately shared so the two layers cannot disagree about what "too large" means | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` §4). | |
| `SATQUERY_MAX_BODY_BYTES` is **enforced twice** — from the `Content-Length` header *and* while reading | |
| the bytes — because the header check is *declarative*: it measures what the client claims. The measured | |
| consequence is in `docs/DEPLOYMENT_ARCHITECTURE.md` §4 (F-6), with the cap at 8 MiB and a 12 MiB body: | |
| | Client behaviour | Result | Peak allocation | Bytes read | | |
| |---|---|---|---| | |
| | `Content-Length` declared, 12 MiB | `413 oversized_image` | **0.2 MiB** | **0** | | |
| | `Content-Length` omitted, 12 MiB | `502 model_unavailable` | **13.9 MiB** | **12 MiB** | | |
| and allocation tracked body size exactly with no ceiling: `1/8/16/32/64 MiB in → 3.0/8.1/16.0/32.0/64.0 MiB | |
| allocated`. The remedy was to make the cap unconditional by enforcing it **while reading**, in the single | |
| shared reader `gateway/assets.py::read_body_bounded`, called by both layers | |
| (`gateway/app.py::_read_body_bounded` is now a thin adapter over it; `app/space_app.py`'s `/v1/assets` | |
| handler calls the same function — that is F-9, which found the Space calling `await request.body()` and | |
| holding 64 MiB in → 128 MiB peak). | |
| > **The honest framing, quoted from the source:** *"Operators should not treat the header check as the | |
| > protection — it protects the gateway's memory against honest clients, not against hostile ones."* | |
| > (`docs/DEPLOYMENT_ARCHITECTURE.md` §4, F-6 note.) | |
| ### 2.4 Rate limiting is FAIRNESS, not security | |
| This is a ruling, not an implementation detail. `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2 carries the | |
| owner ruling of 2026-09-23: | |
| > *"✅ RULED 2026-09-23 (owner ruling): the limiter is RETAINED as a fairness / rate-control mechanism | |
| > only, and it is explicitly NOT a security or abuse-prevention boundary."* | |
| The measurement that forced the ruling is reproduced here because it is the whole argument. Limit set to | |
| **3 requests / 60 s**, **8 requests** sent in-process: | |
| | Case | Statuses | Throttled | | |
| |---|---|---| | |
| | One client, no `X-Forwarded-For` | `502 502 502 429 429 429 429 429` | **5 / 8** | | |
| | A fresh spoofed `X-Forwarded-For` per request | `502 502 502 502 502 502 502 502` | **0 / 8** | | |
| The mechanism is `gateway/app.py::_client_ip`, which derives the rate-limit key from the **first hop of | |
| `X-Forwarded-For`** — a client-supplied header. Its own docstring already said the value is | |
| attacker-controlled and is *"a rate-limit key, not an identity"*; what the measurement added is that the | |
| limiter **does not hold at all** against a caller willing to vary one header. | |
| Consequences a deployment must honour (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2): | |
| * **Do not size abuse protection on this limiter.** It is not that control. | |
| * **A `429` is a fairness signal, not a security signal**, and its **absence is not evidence** that no | |
| abuse occurred. | |
| * The gateway remains the request-side boundary for **shape, size and content type** — the things it can | |
| actually enforce. Rate is not one of them. | |
| The limit itself is two variables because the limit *is* the pair (`docs/DEPLOYMENT_ARCHITECTURE.md` §4): | |
| `SATQUERY_RATE_LIMIT_PER_IP` and `SATQUERY_RATE_LIMIT_WINDOW_S`; `10` and `60.0` mean "ten per minute". | |
| > **Why there is no code fix.** Correctly trusting `X-Forwarded-For` requires knowing how many proxy hops | |
| > the platform inserts — a deployment fact not verifiable from the build host. Hard-coding an assumption | |
| > would replace a *documented* weakness with an *undocumented* one | |
| > (`docs/DEPLOYMENT_ARCHITECTURE.md` §5.2). | |
| ### 2.5 Request IDs | |
| The gateway generates a request id, injects it on the upstream leg, and echoes it to the client | |
| (`gateway/app.py::_proxy`): | |
| ```python | |
| headers = policy.upstream_headers(dict(request.headers), token=token) | |
| headers["X-Request-Id"] = decision.request_id | |
| ``` | |
| Every non-2xx envelope the gateway owns carries the same id, including ones raised by the framework's own | |
| 404/405 handler, which is registered explicitly (`gateway/app.py`): | |
| ```python | |
| @app.exception_handler(StarletteHTTPException) | |
| async def _contract_envelope_for_transport_errors(request, exc): | |
| code = "routing_error" if exc.status_code < 500 else "satquery_error" | |
| status, body = translate_error(...) | |
| ``` | |
| The comment above that handler records the measurement that motivated it: before the fix, | |
| `GET /v1/whocares → 404 {"detail":"Not Found"}` and `GET /v1/assets → 405 {"detail":"Method Not | |
| Allowed"}`, while every handler-owned path answered with the contract envelope. A client written to the | |
| contract parses `error.code` and would get a `KeyError` **exactly when it is trying to explain a failure | |
| to a user**. `app/space_app.py` carries the same handler for the same reason (F-12/F-12b) — it was | |
| previously registered on the gateway only. | |
| ### 2.6 Timeouts, and the no-retry rule | |
| | Timeout | Default | Meaning | | |
| |---|---|---| | |
| | `SATQUERY_UPSTREAM_TIMEOUT_S` | `90.0` | gateway → inference request timeout; must sit inside the task budget | | |
| | `SATQUERY_WAKE_TIMEOUT_S` | `120` | how long the gateway polls for readiness before giving up | | |
| | `SATQUERY_TUNNEL_TIMEOUT_S` | `150` | how long a tunnel request parks before returning `tunnel_offline` | | |
| Defaults are declared in `deploy/render/main.py`: | |
| ```python | |
| def _wake_timeout_s() -> float: | |
| return float(os.environ.get("SATQUERY_WAKE_TIMEOUT_S", "120")) | |
| def _upstream_timeout_s() -> float: | |
| return float(os.environ.get("SATQUERY_UPSTREAM_TIMEOUT_S", "90")) | |
| ``` | |
| **The gateway never retries `POST /v1/analyze`.** `gateway/app.py::_proxy` states it inline: | |
| ```python | |
| except Exception as exc: # network-level failure | |
| # NO RETRY. A retry on /v1/analyze would spend GPU quota twice | |
| # (docs/DEPLOYMENT_ARCHITECTURE.md section 2.2). | |
| ``` | |
| and `docs/DEPLOYMENT_TOPOLOGY.md` §2 repeats it for the active design: *"Render must not retry | |
| `POST /api/infer` on its own — a retry would consume inference a second time. The client decides on | |
| retry."* The client-side consequence is a hard rule in the frontend contract: **never automatically retry | |
| `POST /v1/analyze`** (`docs/FRONTEND_INTEGRATION.md` §6.1). | |
| ### 2.7 Secret custody | |
| | Secret | Lives | Never | | |
| |---|---|---| | |
| | `GITHUB_TOKEN` | Render environment only | in the browser, in the repo, in a client bundle | | |
| | `HF_TOKEN` | Render environment only, *if* the HF proxy path is used | as above | | |
| The live Render configuration was measured on 2026-09-25 and **has no `SATQUERY_UPSTREAM_URL` and no | |
| `HF_TOKEN`** (`docs/DEPLOYMENT_TOPOLOGY.md` header note; `release/repo/docs/DEPLOYMENT.md` §3.1). The | |
| token that *is* present is `GITHUB_TOKEN` — needed only by the GitHub-API wake path, and reported in the | |
| health payload as a boolean, never a value: | |
| ```python | |
| "has_github_token": bool(os.environ.get("GITHUB_TOKEN")), | |
| ``` | |
| `docs/FRONTEND_INTEGRATION.md` §7 states the frontend requirement plainly: **no secrets in the browser**, | |
| talk only to the gateway, and never call the inference host directly — *"it is not the security boundary | |
| and its CORS will not welcome you."* | |
| > **This document contains no credential, token, key or password, and no path to a credential file.** | |
| > Every secret is described by *where it lives*, never by its value. | |
| ### 2.8 Error translation, and one rule about codes | |
| The gateway translates upstream failures into the documented envelope but **passes the `code` through | |
| unchanged** (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.3): | |
| > *"The `code` is passed through **unchanged**. The gateway must not invent codes: the taxonomy in | |
| > `core/errors.py` is the single source of truth, and a gateway that remapped it would make the | |
| > frontend's error handling unpredictable."* | |
| The envelope shape is fixed (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.3): | |
| ```json | |
| { | |
| "error": { | |
| "code": "pair_misaligned", | |
| "message": "The images are not sufficiently co-registered for spatial analysis.", | |
| "detail": "RMSE 4.21 px exceeds the 2.0 px budget", | |
| "recoverable": false, | |
| "request_id": "req_01H...", | |
| "run_id": "9f2c1c0e-..." | |
| } | |
| } | |
| ``` | |
| The orchestrator's own translation table is small and explicit (`deploy/render/main.py`): | |
| | Orchestrator error class | `code` | HTTP | `recoverable` | | |
| |---|---|---|---| | |
| | `WakeTimeout` | `wake_timeout` | `504` | `true` | | |
| | `OrchestratorConfigError` | `orchestrator_config_error` | `500` | `false` | | |
| | `OrchestratorUpstreamError` | `upstream_unreachable` | `502` | `true` | | |
| | connection/timeout to upstream (`_proxy`) | `upstream_unreachable` | `502` | `true` | | |
| | other transport error (`_proxy`) | `upstream_error` | `502` | `true` | | |
| | non-JSON upstream body (`_proxy`) | `schema_validation_error` | `502` | `true` | | |
| | non-JSON request body (`/api/infer`) | `invalid_request` | `400` | `false` | | |
| A **non-JSON upstream body is a defect**, not a pass-through. `gateway/app.py::_proxy` enforces this for | |
| **every** status, not only 2xx, and the comment records why: the guard originally read | |
| `upstream.status_code < 400`, so a non-JSON 4xx/5xx — a proxy error page, an HTML 502 from a load | |
| balancer, a plain-text stack trace — was forwarded verbatim. A sandbox egress proxy returned a 502 whose | |
| body disclosed `os error 10061`; that is how it was found. `/v1/health` is exempt because a liveness probe | |
| may legitimately answer non-JSON. | |
| **A transport failure's raw exception text is never published** (F-15c, owner ruling 2026-09-23). | |
| `gateway/app.py::_transport_failure_detail` maps the exception's MRO class names to a path-free | |
| classification: | |
| ```python | |
| _TRANSPORT_FAILURES: tuple[tuple[str, str], ...] = ( | |
| ("TimeoutException", "the upstream did not respond within the gateway timeout"), | |
| ("ConnectError", "the upstream could not be reached"), | |
| ("ProxyError", "the gateway's egress proxy refused the connection"), | |
| ) | |
| ``` | |
| The full exception still reaches the operator through `_log.error(..., exc_info=exc)`. It is **moved, not | |
| deleted**. | |
| ### 2.9 What the gateway must NOT do | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 is a closed list: | |
| * No persistence. No database, no Redis, no session store. | |
| * No model inference. | |
| * No auth system (plan §74). | |
| * No request queue (plan §73 forbids Redis-cluster/queue infrastructure). | |
| * No retries on `POST /v1/analyze`. | |
| * **No second copy of the capability table.** The gateway proxies `/v1/capabilities` and nothing else | |
| decides that question. The authoritative sources for asset counts are `core.planner.CAPABILITY_ASSETS` | |
| and `SpecialistSpec.requires_assets`; per `_indices_for`'s docstring, *"duplicating that logic here | |
| would give two places to disagree."* | |
| * **No asset storage.** The gateway relays upload bytes; it does not retain them. The store lives with the | |
| inference host, which is the only component that will read them back | |
| (`app/space_app.py::get_asset_store` docstring). | |
| `deploy/render/main.py`'s module docstring states the same three absences in one line: *"It holds no | |
| model, no state, no database, and performs **no auth** (per plan §73/§74)."* And it repeats the | |
| capability-table rule: *"There is deliberately **no second copy** of the capability table here; the | |
| gateway proxies `/v1/capabilities` and nothing else decides that question."* | |
| ### 2.10 The proxied route allowlist | |
| The gateway forwards an **allowlist, not a passthrough** (`gateway/app.py`): | |
| ```python | |
| PROXIED_ROUTES: tuple[str, ...] = ( | |
| "/v1/health", | |
| "/v1/capabilities", | |
| "/v1/analyze", | |
| "/v1/assets", | |
| ) | |
| COSTLY_ROUTES: tuple[str, ...] = ("/v1/analyze", "/v1/assets") | |
| ``` | |
| The two tuples answer different questions and are deliberately separate — *"may this reach the Space at | |
| all?"* versus *"does it cost a metered resource?"* — because collapsing them would make the rate | |
| limiter's coverage depend on the proxy allowlist (`gateway/app.py`). | |
| `BLOCKED_ROUTES` is **empty**, and the comment says it should stay that way: the tuple exists so a route | |
| the contract discusses but the server does not implement answers **501 with a reason** instead of a 404 | |
| a frontend developer would debug as a typo. | |
| On the orchestrator side the four routes are `/api/health`, `/api/infer`, `/api/capabilities`, | |
| `/api/assets`, each proxying to the matching `/v1/*` route (`docs/DEPLOYMENT_TOPOLOGY.md` §3.2; | |
| `deploy/render/main.py`). `/api/health` is the exception: it **never** answers for the inference host. | |
| Its docstring says so — *"Reports its own configuration; never answers for the Codespace (that is | |
| `/api/capabilities`)."* | |
| --- | |
| ## 3. Why the transport is an outbound tunnel | |
| ### 3.1 The forwarded-port failure | |
| A GitHub Codespace exposes a forwarded port publicly, but **for a private repository that forwarded URL | |
| returns HTTP 302** — a redirect to a sign-in page, not the service. `docs/DEPLOYMENT_TOPOLOGY.md` records | |
| this in its measured note: | |
| > *"Transport is an **outbound tunnel**, not a polled forwarded port: the Codespace runs | |
| > `deploy/codespace/tunnel_agent.py`, which dials out to `POST /tunnel/agent` (long-poll) and executes | |
| > against `http://127.0.0.1:8000` locally."* | |
| `release/repo/docs/DEPLOYMENT.md` §7 lists it among the platform traps: | |
| > *"A forwarded Codespace port returns `302` for a private repo — which is *why* the tunnel exists."* | |
| and `docs/DEPLOYMENT_DECISION.md`'s correction banner records the historical position and its reversal: | |
| > *"Codespaces were **not** dropped; the forwarded-port path is dead (HTTP 302 for a private repo) and an | |
| > outbound tunnel is used instead."* | |
| ### 3.2 What the inversion buys | |
| `deploy/codespace/launch.sh` states the property in its header comment, and it is worth quoting because it | |
| is the whole reason the design is robust to repository visibility: | |
| > *"The tunnel is why this works with a PRIVATE repository: the agent makes only outbound HTTPS calls, so | |
| > GitHub's port-forwarding relay, port visibility and the repository's visibility are all irrelevant. The | |
| > orchestrator never dials into this Codespace."* | |
| Consequences, each observable: | |
| | Property | Value under the tunnel | | |
| |---|---| | |
| | Repository visibility | irrelevant — only outbound HTTPS is used | | |
| | Port visibility setting | irrelevant | | |
| | Inbound firewall / NAT | no inbound connection is required at all | | |
| | Who initiates | the **Codespace**, to `SATQUERY_HUB_URL` | | |
| | What the hub needs | a long-poll endpoint and a way to match a response to a pending request | | |
| ### 3.3 Direction, restated as a diagram | |
| ```mermaid | |
| sequenceDiagram | |
| autonumber | |
| participant CF as "Cloudflare Pages" | |
| participant R as "Render hub" | |
| participant TA as "Codespace tunnel agent" | |
| participant API as "FastAPI :8000" | |
| Note over TA,R: startup — agent dials OUT | |
| TA->>R: POST /tunnel/agent (announce, long-poll) | |
| R-->>TA: (holds the poll open) | |
| CF->>R: POST /api/infer | |
| R->>TA: deliver request on the open poll | |
| TA->>API: POST http://127.0.0.1:8000/v1/analyze | |
| API-->>TA: ResultEnvelope | |
| TA-->>R: response | |
| R-->>CF: envelope + X-SatQuery-State | |
| ``` | |
| > **Honest note on the agent's internals.** `deploy/codespace/launch.sh` invokes | |
| > `python deploy/codespace/tunnel_agent.py` and greps its log for the string `announced to hub`. That | |
| > file is **not present in the monorepo working tree** and is **not tracked by git** (see §9.4), so its | |
| > function names, arguments and payload shapes are | |
| > `UNKNOWN — not established from the available evidence`. What *is* established is: the agent exists in | |
| > the production `SatQuery-Inference` repository (`docs/FINAL_DELIVERY_TODO.md` §1.3), it dials | |
| > `SATQUERY_HUB_URL`, it executes against `http://127.0.0.1:8000`, and it is supervised by | |
| > `deploy/codespace/launch.sh`. | |
| ### 3.4 The observable proof of transport | |
| The frontend treats a response header as the evidence that the hub forwarded to the Codespace rather than | |
| answering locally. `frontend/assets/js/live.js` reads it, and the unit suite pins the read: | |
| > *"`x-satquery-transport: tunnel` is the proof that Render forwarded to the Codespace rather than | |
| > answering locally. It is only readable before the response object is discarded."* | |
| > (`tests/unit/test_frontend_live_wiring.py`, `test_the_client_reads_the_transport_header_as_evidence`) | |
| The measured live value is `x-satquery-transport: tunnel` on `POST /api/infer` (`docs/FINAL_DELIVERY_TODO.md` | |
| §1.4, §6 E-03; `docs/FINAL_DELIVERY_REPORT.md` §3 P3). | |
| --- | |
| ## 4. Wake flow | |
| ### 4.1 The flow | |
| The inference Codespace is CPU-first and **may be stopped when idle**. Before a request can be served the | |
| hub starts it (if stopped) and polls health until it answers. The frontend shows *"Waking inference | |
| engine…"* while this happens (`docs/DEPLOYMENT_TOPOLOGY.md` §2). | |
| ```mermaid | |
| sequenceDiagram | |
| participant CF as "Cloudflare Pages" | |
| participant R as "Render hub" | |
| participant C as "GitHub Codespace" | |
| participant HF as "Hugging Face" | |
| CF->>R: GET /api/health (or POST /api/infer) | |
| R->>C: is the Codespace running? | |
| alt stopped | |
| R->>C: start Codespace | |
| R->>C: poll GET /v1/health | |
| C-->>R: 200 {status: ok|degraded} | |
| R-->>CF: "Waking inference engine…" | |
| end | |
| CF->>R: POST /api/infer (query + assets) | |
| R->>C: POST /v1/analyze | |
| C->>HF: resolve pinned model references | |
| C-->>R: ResultEnvelope | |
| R-->>CF: result (envelope + error translation) | |
| ``` | |
| Source: `docs/DEPLOYMENT_TOPOLOGY.md` §2 (verbatim structure). | |
| ### 4.2 The wake path in code | |
| `deploy/render/main.py::ensure_codespace_up()` is the wake implementation. Its contract is precise: | |
| ```python | |
| async def ensure_codespace_up() -> tuple[str, bool]: | |
| """Ensure the Codespace is running; return ``(base_url, woke)``. | |
| Steps: | |
| 1. ``GET`` the Codespace via the GitHub API. | |
| 2. If ``state != "available"``, ``POST .../start``. | |
| 3. Poll ``GET {base}/v1/health`` until 200 or until | |
| ``SATQUERY_WAKE_TIMEOUT_S`` elapses. | |
| """ | |
| ``` | |
| Its polling knobs are module constants: | |
| ```python | |
| _WAKE_POLL_INTERVAL_S = 2.0 | |
| _WAKE_HEALTH_TIMEOUT_S = 10.0 | |
| ``` | |
| and the failure mapping is explicit: a GitHub auth/transport failure becomes | |
| `OrchestratorUpstreamError` (`502`, recoverable), a missing Codespace name becomes | |
| `OrchestratorConfigError` (`500`, not recoverable), and an exhausted deadline raises `WakeTimeout` | |
| (`504`, recoverable) with the last probe error in the detail. | |
| ### 4.3 The response header the client reads | |
| `/api/infer` tags the proxied response so the frontend can tell whether the delay was a cold start | |
| (`deploy/render/main.py`): | |
| ```python | |
| out = await _proxy("POST", f"{base}/v1/analyze", json=body) | |
| out.headers["X-SatQuery-State"] = "waking" if woke else "ready" | |
| return out | |
| ``` | |
| The unit suite pins both headers on the client side: `assert "x-satquery-transport" in source` and | |
| `assert "X-SatQuery-State" in source` (`tests/unit/test_frontend_live_wiring.py`). | |
| ### 4.4 The wake path is a *fallback* in the tunnel design | |
| The measured note in `docs/DEPLOYMENT_TOPOLOGY.md` §2 is explicit that the tunnel design does not depend | |
| on the GitHub-API wake: | |
| > *"The GitHub-API wake path (`POST /user/codespaces/{name}/start`) still exists but the tunnel design | |
| > relies on the agent reconnecting on Codespace start via the devcontainer `postStartCommand`."* | |
| So there are two mechanisms and they are not equivalent: | |
| | Mechanism | Trigger | Effect when it works | Effect when it fails | | |
| |---|---|---|---| | |
| | Devcontainer `postStartCommand` → `launch.sh` → tunnel agent | every Codespace start | agent reconnects; `agent_connected: true` | `agent_connected: false`; `/api/infer` parks to `SATQUERY_TUNNEL_TIMEOUT_S` | | |
| | GitHub-API wake (`ensure_codespace_up`) | any `/api/*` request | starts a stopped Codespace, polls `/v1/health` | `wake_timeout` (`504`, recoverable) | | |
| --- | |
| ## 5. Cold start — documented, not hidden | |
| Render's free tier **sleeps when idle**, and the Codespace **may be stopped** (the live GitHub value | |
| recorded is `idle_timeout_minutes=30`, `docs/FINAL_DELIVERY_TODO.md` §6 E-04). The measured statement is: | |
| > *"Render's free tier also sleeps when idle. Cold start is therefore tens of seconds and is **documented, | |
| > not hidden**."* (`docs/DEPLOYMENT_TOPOLOGY.md` §2) | |
| The UI consequence is recorded in `docs/FRONTEND_INTEGRATION.md` §6: | |
| | Constraint | Value | UI consequence | | |
| |---|---|---| | |
| | Cold start | tens of seconds | *"A determinate-looking progress bar would lie. Use an indeterminate state with a 'this can take up to a minute' hint."* | | |
| and the operator consequence in `docs/FINAL_DELIVERY_REPORT.md` §8: | |
| > *"**Warm the demo stack** ~10 min before presenting: open the Codespace and confirm `GET /api/health` | |
| > shows `tunnel.agent_connected:true`. If the Codespace idle-stops, restart it (the tunnel agent | |
| > reconnects via the devcontainer `postStartCommand`)."* | |
| > **No latency characterisation exists.** `docs/FRONTEND_INTEGRATION.md` §9 states it plainly: | |
| > *"Latency is not characterized. No cold-start or throughput measurement has been taken against a live | |
| > Space."* The phrase "tens of seconds" is a documented expectation, not a measurement. A precise cold-start | |
| > distribution is `UNKNOWN — not established from the available evidence`. | |
| --- | |
| ## 6. `transport_mode: auto`, the fallthrough, and B-07 | |
| ### 6.1 The live transport configuration | |
| The live Render service reports its transport settings in the health payload. Measured | |
| 2026-09-25: | |
| | Setting | Live value | | |
| |---|---| | |
| | `transport_mode` | `auto` | | |
| | `tunnel_timeout_s` | `150.0` | | |
| | `wake_timeout_s` | `120.0` | | |
| | `upstream_timeout_s` | `90.0` | | |
| Source: `release/repo/docs/DEPLOYMENT.md` §2 (live payload) and `docs/DEPLOYMENT_TOPOLOGY.md` header note. | |
| ### 6.2 The fallthrough, exactly | |
| `docs/FINAL_DELIVERY_TODO.md` §5 (blocker register, row B-07) records the confirmed root shape: | |
| > *"Root shape confirmed 2026-09-25: in `auto` transport mode a tunnel timeout **falls through** to the | |
| > forward path (`SatQuery-Backend/main.py:546`), which then burns `wake_timeout_s=120` on a 302 → the | |
| > observed 504."* | |
| `DELIVERY_REPORT_2026-09-25.md` §4 gives the mechanism and the arithmetic: | |
| > *"in `auto` transport mode a tunnel timeout **falls through** to the forward path (`main.py:546` returns | |
| > early only when `mode == "tunnel"`); the forward path then burns `wake_timeout_s = 120` on a 302. | |
| > Measured timing ≈ 249 s ≈ `tunnel_timeout_s=150` + `wake_timeout_s=120`."* | |
| So the worst case is: | |
| ``` | |
| tunnel park 150 s (SATQUERY_TUNNEL_TIMEOUT_S) | |
| + wake poll 120 s (SATQUERY_WAKE_TIMEOUT_S) | |
| ------------------------- | |
| ≈ 249 s → a 504 the client waited four minutes for | |
| ``` | |
| ```mermaid | |
| flowchart TD | |
| A["POST /api/infer<br/>transport_mode = auto"] --> B{"tunnel agent<br/>connected?"} | |
| B -- yes --> C["execute via tunnel<br/>x-satquery-transport: tunnel"] | |
| B -- "no / timeout" --> D["tunnel park expires<br/>SATQUERY_TUNNEL_TIMEOUT_S = 150 s"] | |
| D --> E{"mode == tunnel?"} | |
| E -- yes --> F["return tunnel_offline<br/>503 recoverable"] | |
| E -- "no (auto) → FALLS THROUGH" --> G["forward path:<br/>forwarded port answers 302"] | |
| G --> H["burns wake_timeout_s = 120 s<br/>polling health"] | |
| H --> I["wake_timeout<br/>504 recoverable"] | |
| style I fill:#fde,stroke:#c33 | |
| style D fill:#ffe,stroke:#cc3 | |
| ``` | |
| ### 6.3 B-07 is OPEN | |
| **`B-07` — Transient tunnel-agent gaps — is `OPEN`.** Stated three times in the sources so it cannot be | |
| mistaken: | |
| > *"`B-07` | **Transient tunnel-agent gaps** | OPEN | A request can hang or return 504 (`tunnel_offline` | |
| > / wake timeout; the forwarded port returns 302). Observed once live. Mitigation: keep the Codespace | |
| > warm before the demo; the client shows an actionable retry message."* (`docs/FINAL_DELIVERY_REPORT.md` | |
| > §6) | |
| > *"`B-07` | Transient tunnel-agent gaps (agent briefly absent) → a request can hang or return 504 | |
| > (`tunnel_offline` / wake timeout, forward path 302) | … | **OPEN — patch prepared, not deployed.**"* | |
| > (`docs/FINAL_DELIVERY_TODO.md` §5) | |
| > *"B-07 backend patch **prepared, NOT deployed**."* (`docs/FINAL_DELIVERY_TODO.md` sprint-status note) | |
| ### 6.4 The prepared patch — prepared, NOT deployed | |
| The patch is `fix-b07-forward-unavailable.patch`, in the session workspace at | |
| `.workbuddy-ai/scratch/deployed-backend/fix-b07-forward-unavailable.patch` | |
| (`DELIVERY_REPORT_2026-09-25.md` §8). Its content and verification: | |
| | Item | Detail | | |
| |---|---| | |
| | Base | the **deployed** `SatQuery-Backend/main.py` @ `89d80eaddec5` (769 lines) | | |
| | Size | 9 hunks plus a 340-line test | | |
| | Change A | adds `forward_unavailable` (`503`, `recoverable: true`) for a **terminal** 302/401/403 on the forward path, instead of burning the wake timeout | | |
| | Change B | adds `upstream_timeout` (`504`) for "tunnel healthy but slow" | | |
| | Change C | fixes `/api/health` `codespace_name` trailing `\n` via `.strip()` | | |
| | Independent verification | `git apply --check` clean, `git apply` clean, `py_compile` OK | | |
| | Presence check | `forward_unavailable` @ `main.py:326`, `upstream_timeout` @ `:601`, `codespace_name` `.strip()` @ `:686` | | |
| | Deployment status | **NOT deployed** | | |
| Sources: `DELIVERY_REPORT_2026-09-25.md` §4; `docs/FINAL_DELIVERY_TODO.md` §6 E-12. | |
| > **A retracted claim, recorded because the honesty matters.** The report records that an earlier claim | |
| > that the patch *"would not have prevented"* the observed 504 *"was wrong and was retracted"*. The | |
| > corrected position: *"Change A is genuinely **on the failing path** — it converts a 504-after-249 s into | |
| > a 503-early with an actionable code."* (`DELIVERY_REPORT_2026-09-25.md` §4.) | |
| **Why it is not deployed:** *"the patch is not needed for the demo and touches the live backend. The | |
| residual is better mitigated operationally (keep the Codespace warm, raise the idle timeout)."* | |
| (`DELIVERY_REPORT_2026-09-25.md` §4.) | |
| ### 6.5 Operational trap recorded with the patch | |
| > *"the local `C:/Users/anish/SatQuery-Backend` (680 lines) is **STALE**. Always fetch the deployed | |
| > `main.py` before touching backend code."* (`DELIVERY_REPORT_2026-09-25.md` §4) | |
| This is the same class of trap as §9.4 below: **the working copy is not the deployed source.** | |
| --- | |
| ## 7. The full live health payload | |
| ### 7.1 The measured payload | |
| Probed live on 2026-09-25 against `https://<backend-host>/api/health` | |
| (`release/repo/docs/DEPLOYMENT.md` §2): | |
| ```json | |
| {"status":"ok","service":"satquery-orchestrator", | |
| "tunnel":{"agent_connected":true,"agent_id":"codespaces-fd1038","pending":0,"completed":97}, | |
| "config":{"codespace_name":"potential-space-trout-r4ppw969w45j2pvvw\n","codespace_port":8000, | |
| "transport_mode":"auto","tunnel_timeout_s":150.0,"wake_timeout_s":120.0, | |
| "upstream_timeout_s":90.0,"device":"cpu","has_github_token":true}} | |
| ``` | |
| Exact command used elsewhere in the project's evidence register: | |
| `curl --noproxy '*' https://<backend-host>/api/health` | |
| (`docs/FINAL_DELIVERY_REPORT.md` §4). | |
| ### 7.2 Field-by-field | |
| | Field | Type | Meaning | Live value | | |
| |---|---|---|---| | |
| | `status` | string | the hub's own liveness | `"ok"` | | |
| | `service` | string | the service identity | `"satquery-orchestrator"` | | |
| | `tunnel.agent_connected` | bool | is a tunnel agent currently polling? | `true` | | |
| | `tunnel.agent_id` | string | which agent identity holds the poll | `"codespaces-fd1038"` | | |
| | `tunnel.pending` | int | requests delivered but not yet answered | `0` | | |
| | `tunnel.completed` | int | requests completed since the agent connected | `97` | | |
| | `config.codespace_name` | string | the target Codespace | `"…pvvw\n"` — **carries a trailing `\n`** | | |
| | `config.codespace_port` | int | the inference port | `8000` | | |
| | `config.transport_mode` | string | transport selection | `"auto"` | | |
| | `config.tunnel_timeout_s` | float | tunnel park budget | `150.0` | | |
| | `config.wake_timeout_s` | float | wake poll budget | `120.0` | | |
| | `config.upstream_timeout_s` | float | proxy request timeout | `90.0` | | |
| | `config.device` | string | declared device | `"cpu"` | | |
| | `config.has_github_token` | bool | is a GitHub token configured? | `true` | | |
| ### 7.3 `completed` was observed at three different values — do not treat any as a constant | |
| The tunnel counter is a **monotonic runtime counter**, not a fixed fact. Three measured readings exist, | |
| each with its own provenance: | |
| | Reading | Where recorded | | |
| |---|---| | |
| | `completed: 97` | `release/repo/docs/DEPLOYMENT.md` §2 (the live health probe) | | |
| | `completed: 314` | `docs/FINAL_DELIVERY_TODO.md` §1.4 and §6 E-02; `docs/DEPLOYMENT_TOPOLOGY.md` §2 | | |
| | `completed: 338` | `docs/FINAL_DELIVERY_REPORT.md` §3 P2 | | |
| They are consistent with each other — the counter grows — and the honest statement is | |
| **"`completed` was measured at 97, 314 and 338 at three different times on 2026-09-25."** Quoting any one | |
| of them as *the* value would be wrong. | |
| ### 7.4 `codespace_name` carries a trailing newline — B-02, OPEN (cosmetic) | |
| `config.codespace_name` reports `…pvvw\n`. This is **B-02**, and its status is `OPEN` **but cosmetic**: | |
| > *"`P2-T03` `/api/health` `codespace_name` trailing `\n` | DEFERRED (cosmetic) | Wake path is safe | |
| > (`_codespace_name()` strips, `main.py:123,357`); only the health payload reports the raw value."* | |
| > (`docs/FINAL_DELIVERY_REPORT.md` §6) | |
| > *"`B-02` | `/api/health` reports `codespace_name` with a trailing `\n` | **Cosmetic** — reporting only; | |
| > the wake path strips via `_codespace_name()` (`main.py:123,357`) | P2-T03 | none needed | | |
| > DOWNGRADED"* (`docs/FINAL_DELIVERY_TODO.md` §5) | |
| The fix is known and one line — *"change line 619 to `_codespace_name()`, then Render redeploys"* | |
| (`docs/FINAL_DELIVERY_TODO.md` §4, P2-T03) — and the row's own reasoning for deferring is that *"a | |
| live-backend redeploy before the demo is not worth the risk."* | |
| > **Do not upgrade this.** `B-02` is `OPEN`. It is not `RESOLVED`, and it is not `CLOSED`. | |
| ### 7.5 The orchestrator's *own* `/api/health` shape in the repository | |
| `deploy/render/main.py` — the monorepo copy, which is **not** the deployed source (§9.4) — declares a | |
| different, simpler health payload. Reproduced because it documents the *contract* of the route even where | |
| the deployed implementation has grown: | |
| ```python | |
| @app.get("/api/health") | |
| async def health() -> dict[str, Any]: | |
| """Orchestrator liveness. Reports its own configuration; never answers | |
| for the Codespace (that is /api/capabilities).""" | |
| return { | |
| "status": "ok", | |
| "service": "satquery-orchestrator", | |
| "config": { | |
| "codespace_name": os.environ.get("CODESPACE_NAME", ""), | |
| "codespace_port": _codespace_port(), | |
| "has_github_token": bool(os.environ.get("GITHUB_TOKEN")), | |
| "allowed_origins": _allowed_origins(), | |
| "production_origins": list(_PRODUCTION_ORIGINS), | |
| "dev_origins_enabled": _dev_origins_enabled(), | |
| "wake_timeout_s": _wake_timeout_s(), | |
| "upstream_timeout_s": _upstream_timeout_s(), | |
| "device": os.environ.get("SATQUERY_DEVICE", ""), | |
| }, | |
| } | |
| ``` | |
| Note the design decision visible here: `allowed_origins` reports the **effective** list, so an operator can | |
| confirm **from outside** what the service will actually accept — not just what they set. The comment says | |
| the dev entries being visible *"is how a production deployment proves it turned them off."* | |
| > **The discrepancy is real and is stated rather than smoothed over.** The deployed payload carries a | |
| > `tunnel` block and `config.transport_mode` / `config.tunnel_timeout_s`, which the monorepo copy does | |
| > not. The monorepo copy is a **532-line** file with no tunnel code at all; the deployed | |
| > `SatQuery-Backend/main.py` is **768–769 lines** with it (`docs/FINAL_DELIVERY_TODO.md` §1.1; | |
| > `DELIVERY_REPORT_2026-09-25.md` §4). | |
| --- | |
| ## 8. Environment variables | |
| ### 8.1 Render (orchestrator) — measured live values | |
| | Variable | Live value | Purpose | | |
| |---|---|---| | |
| | `CODESPACE_NAME` | `potential-space-trout-r4ppw969w45j2pvvw` | which Codespace to target | | |
| | `CODESPACE_PORT` | `8000` | the inference port on that Codespace | | |
| | `SATQUERY_ALLOWED_ORIGINS` | `https://satquery.pages.dev` | CORS allowlist (the Pages origin) | | |
| | `SATQUERY_DEVICE` | `cpu` | declared device | | |
| | `SATQUERY_TRANSPORT` | `auto` | transport selection | | |
| | `SATQUERY_TUNNEL_TIMEOUT_S` | `150` | tunnel park budget | | |
| | `SATQUERY_WAKE_TIMEOUT_S` | `120` | wake poll budget | | |
| | `SATQUERY_UPSTREAM_TIMEOUT_S` | `90` | gateway → upstream request timeout | | |
| | `GITHUB_TOKEN` | present | GitHub API wake path; never sent to the browser | | |
| Source: `docs/DEPLOYMENT_TOPOLOGY.md` header note (measured against `GET /api/health`); | |
| `release/repo/docs/DEPLOYMENT.md` §3.1. | |
| > **Two absences are as important as the presences.** There is **no `SATQUERY_UPSTREAM_URL`** and **no | |
| > `HF_TOKEN`** in the live configuration (`docs/DEPLOYMENT_TOPOLOGY.md`; `release/repo/docs/DEPLOYMENT.md` | |
| > §3.1). `SATQUERY_UPSTREAM_URL` is absent because the transport is the outbound tunnel, not a forwarded | |
| > port; `HF_TOKEN` is absent because the HF proxy path is not used live. | |
| ### 8.2 Render — the blueprint's declared variables | |
| `render.yaml` (the blueprint) declares the same vocabulary as a service definition: | |
| ```yaml | |
| services: | |
| - type: web | |
| name: satquery-orchestrator | |
| runtime: python | |
| plan: free | |
| buildCommand: pip install -r deploy/render/requirements.txt | |
| startCommand: uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT | |
| healthCheckPath: /api/health | |
| envVars: | |
| - key: PORT | |
| sync: false | |
| - key: SATQUERY_ALLOWED_ORIGINS | |
| sync: false | |
| - key: GITHUB_TOKEN | |
| sync: false | |
| - key: CODESPACE_NAME | |
| sync: false | |
| - key: CODESPACE_PORT | |
| value: "8000" | |
| - key: SATQUERY_DEVICE | |
| value: "cpu" | |
| - key: SATQUERY_WAKE_TIMEOUT_S | |
| value: "120" | |
| - key: SATQUERY_UPSTREAM_TIMEOUT_S | |
| value: "90" | |
| ``` | |
| Three things this file establishes that are easy to miss: | |
| 1. `plan: free` — the free tier, which is *why* Render sleeps when idle (§5). | |
| 2. `healthCheckPath: /api/health` — the platform's own liveness probe points at the orchestrator's | |
| self-report route, which never touches the inference host. | |
| 3. `sync: false` on `SATQUERY_ALLOWED_ORIGINS`, `GITHUB_TOKEN`, `CODESPACE_NAME` and `PORT` means those are | |
| **operator-supplied**, not blueprint-committed. No secret value appears in the repository. | |
| > The blueprint does **not** declare `SATQUERY_TRANSPORT` or `SATQUERY_TUNNEL_TIMEOUT_S`, which the live | |
| > service reports. The blueprint and the live service have diverged. Whether the live service sets them | |
| > through the dashboard or through a newer blueprint is `UNKNOWN — not established from the available | |
| > evidence`; what is established is the live value set in §8.1. | |
| ### 8.3 Codespace (inference) — declared and effective | |
| | Variable | Where set | Purpose | | |
| |---|---|---| | |
| | `PORT` | `containerEnv` = `"8000"`, re-exported by `launch.sh` | platform-assigned; **must be read** (historical blocker #2) | | |
| | `SATQUERY_DEVICE` | `containerEnv` = `"cpu"`, re-exported by `launch.sh` | `cpu` \| `cuda` \| `mps` \| `null`; read **without importing torch** | | |
| | `SATQUERY_ASSET_ENABLED` | `containerEnv` = `"1"`, re-exported by `launch.sh` | enables `POST /v1/assets`; **both** this and the dir are required | | |
| | `SATQUERY_ASSET_DIR` | `containerEnv` = `"/tmp/satquery-assets"`, re-exported by `launch.sh` | where uploaded bytes are written | | |
| | `SATQUERY_MAX_FILE_BYTES` | not set live (default applies) | per-file cap, shared with Render | | |
| | `SATQUERY_ASSET_MAX_FILES` | not set live (default applies) | optional handle capacity, default `32` | | |
| | `SATQUERY_ASSET_TTL_S` | not set live (default applies) | optional handle lifetime, default `900.0` | | |
| | `SATQUERY_HUB_URL` | defaulted by `launch.sh` | the hub the agent dials | | |
| | `PYTHONPATH` | set by `launch.sh` | repo root, so `import app` resolves | | |
| Sources: `.devcontainer/devcontainer.json`; `deploy/codespace/launch.sh`; `app/space_app.py`. | |
| `.devcontainer/devcontainer.json` in full: | |
| ```json | |
| { | |
| "name": "SatQuery AI — Codespace Inference", | |
| "image": "mcr.microsoft.com/devcontainers/python:3.12", | |
| "forwardPorts": [8000], | |
| "portsAttributes": { | |
| "8000": { "label": "SatQuery inference", "visibility": "public" } | |
| }, | |
| "containerEnv": { | |
| "SATQUERY_DEVICE": "cpu", | |
| "PORT": "8000", | |
| "SATQUERY_ASSET_ENABLED": "1", | |
| "SATQUERY_ASSET_DIR": "/tmp/satquery-assets" | |
| }, | |
| "postCreateCommand": "bash deploy/codespace/post_create.sh", | |
| "postStartCommand": "bash deploy/codespace/launch.sh", | |
| "customizations": { "vscode": { "extensions": ["ms-python.python"] } } | |
| } | |
| ``` | |
| > **A trap worth recording, from `launch.sh`'s own comment:** *"`containerEnv` is only applied when the | |
| > container is CREATED, so setting it there alone would leave an already-running Codespace unconfigured | |
| > until a rebuild. This script runs on every start and is therefore the effective source of truth."* The | |
| > variables are therefore set **twice** — in `containerEnv` and in `launch.sh` — and `launch.sh` is the | |
| > one that governs a running container. | |
| ### 8.4 The historical vocabulary — still the contract | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §4 remains authoritative for the env-var *vocabulary*; only host names | |
| moved. Its full table, reproduced, with the active host substituted: | |
| | Variable | Where it lives (historical → active) | Purpose | | |
| |---|---|---| | |
| | `HF_TOKEN` | Railway only → **Render only, if used** | upstream credential; never sent to the browser | | |
| | `SATQUERY_SPACE_URL` | Railway → **`SATQUERY_UPSTREAM_URL`** | upstream URL | | |
| | `SATQUERY_ALLOWED_ORIGINS` | Railway → **Render** | CORS allowlist | | |
| | `PORT` | Railway → **Render** | supplied by the platform | | |
| | `SATQUERY_DEVICE` | Space → **Codespace** | `cpu` \| `cuda` \| `mps` \| `null`; read **without importing torch** | | |
| | `SATQUERY_ASSET_ENABLED` | Space → **Codespace** | enables `POST /v1/assets`; fails closed | | |
| | `SATQUERY_ASSET_DIR` | Space → **Codespace** | where uploaded bytes are written | | |
| | `SATQUERY_MAX_FILE_BYTES` | **both** | per-file size cap, read by both layers from one variable | | |
| | `SATQUERY_MAX_BODY_BYTES` | Railway → **Render** | whole-request body cap, above the per-file cap | | |
| | `SATQUERY_UPSTREAM_TIMEOUT_S` | Railway → **Render** | gateway → upstream timeout; default `90.0` | | |
| | `SATQUERY_RATE_LIMIT_PER_IP` / `_WINDOW_S` | Railway → **Render** | per-IP count + window | | |
| | `SATQUERY_ASSET_MAX_FILES` / `SATQUERY_ASSET_TTL_S` | Space → **Codespace** | **optional** handle capacity / lifetime | | |
| Three notes from that section are worth carrying forward because they explain *why* the vocabulary has | |
| this shape: | |
| 1. **`SATQUERY_MAX_FILE_BYTES` is applied while reading at both layers, not after** (F-9, F-6). Both layers | |
| call the single reader `gateway/assets.py::read_body_bounded`, so the two enforcement points cannot | |
| drift. | |
| 2. **Both layers refuse an unparsable or non-positive value and name the variable** (F-7). Reading one | |
| variable is not the same as agreeing on its value: the two parsers previously diverged in **opposite | |
| directions** — `'abc'` raised at the gateway but silently defaulted to 4 MiB on the inference host; | |
| `'0'` was accepted at the gateway but rejected on the inference host. A malformed cap now fails startup | |
| at both layers rather than running on a limit nobody chose. | |
| 3. **None of the asset variables is a config key**, and that is deliberate: adding a key to | |
| `configs/base.yaml` moves `Config.hash` off `78f1e3700da15aa1` and invalidates the frozen Phase-9 | |
| benchmark. Asset storage is deployment state, so it is read from the environment. | |
| The F-8 note on `SATQUERY_DEVICE` is also load-bearing and is reproduced in §8.5. | |
| ### 8.5 `SATQUERY_DEVICE`: four read sites, and the case bug | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §4 records that the served value is validated against the contract's | |
| closed set. Measured before the fix, `SATQUERY_DEVICE` had **four read sites** and only three normalised: | |
| | Site | Behaviour before the fix | | |
| |---|---| | |
| | `core/config.py:88` (`Config.device_preference`) | raw — no strip, no lower | | |
| | `app/deployment.py:570` (`gpu_available`) | `.strip().lower()` | | |
| | `app/deployment.py:946` (the served device resolver) | `.strip()`, **no lower** — the odd one | | |
| | `app/deployment.py:970` (`_cuda_detected`) | `.strip().lower()` | | |
| Two defects followed, both measured: | |
| * **Case changed the answer.** `'cuda'` → `'cpu'` but `'CUDA'` → `'CUDA'`, so one payload could announce | |
| `gpu_available: true` alongside `device: "CUDA"` — a GPU is claimed and the device name is not a device. | |
| * **An unparsable value was echoed.** `'garbage'` → `device: "garbage"`, against a field the contract | |
| publishes as a closed set. | |
| The served resolver now normalises and **validates**, returning `None` for anything outside | |
| `{"cpu", "cuda", "mps"}`. `None` is chosen over raising or over a silent `"cpu"`, because it is already a | |
| legal value for the field, it is honest, and defaulting to `"cpu"` *"would mean a typo silently changes | |
| which device the process is believed to use, which is the `_asset_max_file_bytes` mistake from F-7 in a | |
| different variable."* The guard is kept as a literal, not derived from the implementation, so it encodes | |
| the **contract's** set and cannot drift with the code: | |
| ```python | |
| _LEGAL_DEVICES: frozenset[str] = frozenset({"cpu", "cuda", "mps"}) | |
| ``` | |
| > `Config.device_preference` still returns the raw override, deliberately: it is a general-purpose property | |
| > whose other callers may legitimately want the operator's literal text, and narrowing it would be a wider | |
| > change than the defect warrants. **The served path is the one the contract constrains, so it is the one | |
| > that validates.** | |
| --- | |
| ## 9. The tunnel agent and the Codespace launcher | |
| ### 9.1 `deploy/codespace/serve.py` — the entrypoint | |
| The file is 25 lines and its whole job is to bind `build_space_app()` to `$PORT`: | |
| ```python | |
| import os | |
| from app.space_app import build_space_app | |
| import uvicorn | |
| app = build_space_app() | |
| if __name__ == "__main__": | |
| port = int(os.environ.get("PORT", "8000")) | |
| uvicorn.run(app, host="0.0.0.0", port=port) | |
| ``` | |
| Its docstring records the properties that make it import-safe on a CPU host with no GPU and no weights: | |
| > *"`build_space_app()` is cheap to import: FastAPI is imported inside it and no model is loaded at module | |
| > scope, so this file stays import-safe on a CPU host with no GPU and no weights present."* | |
| and it names the device resolution path: *"The serving controller (via | |
| `app.serving.build_serving_controller`) resolves `device` from the `SATQUERY_DEVICE` env var; set it to | |
| `cpu` for the CPU-first adaptation."* | |
| ### 9.2 `deploy/codespace/launch.sh` — what runs on every Codespace start | |
| The script is the devcontainer's `postStartCommand` target. It runs three stages plus a preflight. | |
| **Stage 0 — preflight, refusing to start half-configured.** The header comment states why this exists: | |
| > *"A silently-broken environment is the single worst failure mode here: the server dies, nothing listens | |
| > on the port, and the only external symptom is a bare 401/302 from GitHub's relay — which looks like a | |
| > visibility problem."* | |
| It therefore checks the Python dependencies **including `httpx` explicitly**, and the comment records the | |
| incident: | |
| > *"NOTE: httpx is checked explicitly. `tunnel_agent.py` imports it directly, and it was previously absent | |
| > from `requirements.txt` — so the agent died instantly and the supervised restart loop hid the error in a | |
| > log file."* | |
| ```bash | |
| if ! python -c "import yaml, pydantic, fastapi, uvicorn, httpx" 2>/dev/null; then | |
| echo "ERROR: Python deps are missing (need yaml, pydantic, fastapi, uvicorn, httpx)." >&2 | |
| ... | |
| exit 1 | |
| fi | |
| if ! python -c "import app.space_app" 2>/dev/null; then | |
| echo "ERROR: cannot import the 'app' package even with PYTHONPATH=$REPO_ROOT" >&2 | |
| ... | |
| exit 1 | |
| fi | |
| ``` | |
| **Stage 1 — the inference server, with a staleness guard.** The script records a **stamp** of the git | |
| revision and the asset-upload environment, because `_port_open` alone cannot tell you what the running | |
| process was started from: | |
| ```bash | |
| _current_stamp() { | |
| printf 'rev=%s asset_enabled=%s asset_dir=%s\n' \ | |
| "$(git rev-parse HEAD 2>/dev/null || echo nogit)" \ | |
| "${SATQUERY_ASSET_ENABLED:-}" \ | |
| "${SATQUERY_ASSET_DIR:-}" | |
| } | |
| ``` | |
| and the comment explains the failure a stale process causes: | |
| > *"A stale serve process is worse than no process: it answers `/v1/health` and `/v1/capabilities` from OLD | |
| > code, so the deployment looks alive while reporting the previous revision's capabilities."* | |
| The restart uses the real invocation, not the file path — `pkill -f "python deploy/codespace/serve.py"`, | |
| because *"`pgrep -f serve.py` would also match an editor or this script's own argv."* If SIGTERM is not | |
| enough it escalates to `pkill -9`, and the process is launched detached: | |
| ```bash | |
| setsid nohup python deploy/codespace/serve.py > "$SERVE_LOG" 2>&1 < /dev/null & | |
| ``` | |
| **Stage 2 — the outbound tunnel agent, supervised.** The header comment records the production incident | |
| that shaped the launch: | |
| > *"`setsid` alone is NOT enough in Codespaces. The lifecycle shell that runs `postStartCommand` can still | |
| > reap the process group, which showed up in production as 'the agent announced once, then vanished' — the | |
| > hub then reported `agent_connected=false` and `/api/infer` fell back to the dead forwarded-port path | |
| > (401 -> wake_timeout)."* | |
| The remedy is `setsid + nohup + </dev/null` **plus a supervising wrapper** that relaunches the agent if it | |
| ever exits: | |
| ```bash | |
| setsid nohup bash -c ' | |
| while true; do | |
| echo "[supervisor $(date +%H:%M:%S)] starting tunnel agent" >> "'"$TUNNEL_LOG"'" | |
| python deploy/codespace/tunnel_agent.py >> "'"$TUNNEL_LOG"'" 2>&1 | |
| rc=$? | |
| echo "[supervisor $(date +%H:%M:%S)] tunnel agent exited rc=$rc — restarting in 5s" >> "'"$TUNNEL_LOG"'" | |
| sleep 5 | |
| done | |
| ' > /dev/null 2>&1 < /dev/null & | |
| ``` | |
| The guard is on the **process, not a port** — *"the agent listens on nothing"* — and the supervisor itself | |
| is what gets detached, so *"the agent is effectively immortal for the life of the Codespace."* | |
| **Stage 3 — verify the agent actually connected.** This stage exists because backgrounding with all output | |
| discarded makes a crashing agent invisible: | |
| > *"Backgrounding with all output discarded means a crashing agent is completely invisible — that is | |
| > exactly how a missing `httpx` hid itself. So we wait, then check: the process is alive, and the log shows | |
| > a successful announce."* | |
| ```bash | |
| sleep 4 | |
| if ! pgrep -f "deploy/codespace/tunnel_agent.py" > /dev/null 2>&1; then | |
| echo "WARNING: the tunnel agent is not running. Last log lines:" >&2 | |
| tail -n 20 "$TUNNEL_LOG" 2>/dev/null >&2 || echo " (no log at $TUNNEL_LOG)" >&2 | |
| ... | |
| else | |
| echo "tunnel agent process is up (pid $(pgrep -f 'deploy/codespace/tunnel_agent.py' | head -1))" | |
| if grep -q "announced to hub" "$TUNNEL_LOG" 2>/dev/null; then | |
| echo "tunnel agent announced to the hub successfully" | |
| ... | |
| ``` | |
| The hub URL is a defaulted variable, so a renamed Render service can be overridden in the Codespace: | |
| ```bash | |
| export SATQUERY_HUB_URL="${SATQUERY_HUB_URL:-https://<backend-host>}" | |
| ``` | |
| ### 9.3 The asset-upload environment, and why `/tmp` is correct | |
| `launch.sh` sets the asset variables on every start, and its comment argues the choice rather than | |
| asserting it: | |
| > *"`/tmp` is correct here and not a compromise: the Codespace filesystem is ephemeral, handles are TTL'd | |
| > (900s), and `cache_max_models: 1` means an uploaded asset is consumed within one analysis, so nothing | |
| > needs to outlive the process. The store creates the directory if absent."* | |
| ```bash | |
| export SATQUERY_ASSET_ENABLED="${SATQUERY_ASSET_ENABLED:-1}" | |
| export SATQUERY_ASSET_DIR="${SATQUERY_ASSET_DIR:-/tmp/satquery-assets}" | |
| ``` | |
| The comment also cross-references the exact fallback path in code — *"the directory is intentionally the | |
| same path the code falls back to (`space_app.py:310`)"* — which is | |
| `Path(tempfile.gettempdir()) / "satquery-assets"` in `app/space_app.py::_asset_root()`. That is a | |
| deliberate alignment: *"a deployment that set only the flag — or neither — cannot silently start writing | |
| to a barely-chosen location."* | |
| ### 9.4 The tunnel agent file itself | |
| | Question | Answer | Status | | |
| |---|---|---| | |
| | Is `deploy/codespace/tunnel_agent.py` in the monorepo working tree? | **No** — `Glob **/tunnel_agent*` finds nothing | `MEASURED` | | |
| | Is it tracked by git? | **No** — `git ls-files deploy/` is empty; the whole `deploy/` tree is untracked | `MEASURED` | | |
| | Where does it exist? | `Anish-lab-blip/SatQuery-Inference` (private) — *"Codespace FastAPI + `deploy/codespace/tunnel_agent.py`"* (`docs/FINAL_DELIVERY_TODO.md` §1.3) | `VERIFIED` | | |
| | What are its function names, arguments, payload shapes? | `UNKNOWN — not established from the available evidence` | `OPEN` | | |
| | What *is* established about it? | it imports `httpx`; it dials `SATQUERY_HUB_URL`; it executes against `http://127.0.0.1:8000`; it logs `announced to hub`; it is supervised by `launch.sh` | `VERIFIED` (from `launch.sh` comments and greps) | | |
| > **This is the single largest evidence gap in this chapter**, and it is recorded rather than filled in | |
| > with a plausible guess. A reader who needs the agent's protocol should read | |
| > `SatQuery-Inference/deploy/codespace/tunnel_agent.py`. | |
| ### 9.5 The "stale working copy" trap — B-03 | |
| **`B-03` is `KNOWN`**, and it is the reason §9.4 has a gap at all: | |
| > *"`B-03` | Local `deploy/` stale + untracked | Edits there do not deploy | all deploy tasks | edit the 3 | |
| > real repos instead | KNOWN"* (`docs/FINAL_DELIVERY_TODO.md` §5) | |
| > *"Local `deploy/render/main.py` (532 lines, no tunnel) is superseded by `SatQuery-Backend/main.py` (768 | |
| > lines, tunnel)."* (`docs/FINAL_DELIVERY_TODO.md` §1.1) | |
| > *"**Critical:** the deployed backend is **not** this working copy."* (`docs/FINAL_DELIVERY_TODO.md` §1.1) | |
| > *"Local `deploy/` | stale/untracked | Edit the 3 real repos, not this copy."* | |
| > (`docs/FINAL_DELIVERY_REPORT.md` §6) | |
| ### 9.6 Repositories of record | |
| `docs/FINAL_DELIVERY_TODO.md` §1.3: | |
| | Repo | Role | Deployed from | | |
| |---|---|---| | |
| | `Anish-lab-blip/SatQuery-Frontend` (private) | Cloudflare Pages (static) | root = local `frontend/` contents | | |
| | `Anish-lab-blip/SatQuery-Backend` (private) | Render hub + `tunnel.py` + `codespaces.py` | Render `satquery-orchestrator` | | |
| | `Anish-lab-blip/SatQuery-Inference` (private) | Codespace FastAPI + `deploy/codespace/tunnel_agent.py` | Codespace | | |
| | `Anish-lab-blip/SatQuery-AI` (**public**) | umbrella / monorepo mirror | — | | |
| ### 9.7 Deployed revisions | |
| | Component | Repository | Branch | Revision | Host | | |
| |---|---|---|---|---| | |
| | Frontend | `SatQuery-Frontend` | `main` | **`2d7ae53b482d`** | Cloudflare Pages → `satquery.pages.dev` | | |
| | Backend / orchestrator | `SatQuery-Backend` | `main` | **`89d80eaddec5`** | Render → `<backend-host>` | | |
| | Inference | `SatQuery-Inference` | `main` | **`5a0936ace491`** | Codespace `potential-space-trout-r4ppw969w45j2pvvw`, port 8000 | | |
| | Public umbrella | `SatQuery-AI` | `main` | `3dcabd32da41` | the release home | | |
| | Monorepo working copy | `C:/Users/anish/satquery-ai` | `master` | `9d57aed` | local only, **no remote**, 334 dirty entries | | |
| Source: `release/repo/docs/DEPLOYMENT.md` §1. This is the correct place to look up a deployed revision; | |
| **the monorepo HEAD is not the deployed revision.** | |
| --- | |
| ## 10. Deployment mechanics | |
| ### 10.1 Frontend → Cloudflare Pages | |
| Staged by `scripts/stage_pages.mjs`, deployed with `npx wrangler pages deploy`. The staging run measured on | |
| 2026-09-25 (`docs/DEPLOYMENT_DECISION.md` §7): | |
| ``` | |
| files staged : 60 | |
| total bytes : 39,173,936 (37.36 MiB) | |
| largest file : assets/video/satquery-launch-50s.mp4 22,710,313 B (21.66 MiB) | |
| 25 MiB headroom left : 3,504,087 B on the largest file | |
| missing refs in staged : 0 | |
| external network deps : 0 (HERMETIC) | |
| exit : 0 | |
| ``` | |
| `_headers` and `robots.txt` must be **force-included** because no page references them; `provenance.json` | |
| and `CREDITS.md` likewise, because they are provenance records rather than assets | |
| (`docs/DEPLOYMENT_DECISION.md` §7). | |
| ### 10.2 Backend → Render | |
| `render.yaml` is the blueprint (§8.2); `main.py` exposes `app` | |
| (`uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT`). | |
| ### 10.3 Inference → Codespace | |
| `deploy/codespace/serve.py` serves `build_space_app()` on `$PORT`; `.devcontainer/` forwards port `8000` | |
| and runs the tunnel agent on start via `postStartCommand` (`release/repo/docs/DEPLOYMENT.md` §4). | |
| ### 10.4 Repository writes use the GitHub Git Data API, not `git push` | |
| > *"Repository writes are performed through the **GitHub Git Data API** (blob → tree → commit → `PATCH` | |
| > ref) with **sha256 byte-verification** of every uploaded blob. Deletions are expressed as `sha: null` | |
| > tree entries. This is used instead of `git push` so each deployed file is verified by content hash."* | |
| > (`release/repo/docs/DEPLOYMENT.md` §4) | |
| The integrity check is recorded: *"Deployed files were re-read from the GitHub API and compared | |
| byte-for-byte against the local copies: **9 files sha256 byte-identical**, and the deployed HEAD re-read | |
| from the API."* (`release/repo/docs/DEPLOYMENT.md` §4.1; `docs/FINAL_DELIVERY_TODO.md` §6 E-10.) | |
| --- | |
| ## 11. The five historical backend blockers, and how the design closes them | |
| `docs/DEPLOYMENT_DECISION.md` §8 enumerated **five** verified backend blockers that had to be closed | |
| before any backend could boot. `docs/DEPLOYMENT_TOPOLOGY.md` §4 carries them forward with the active | |
| design's response. | |
| | # | Blocker (verified, old doc) | How the new topology addresses it | | |
| |---|---|---| | |
| | 1 | `requirements.txt` declared no `fastapi` / `uvicorn` / `httpx` / `starlette` | the Codespace/Render runtime installs the ASGI stack so `build_space_app()` and the gateway `app` can import | | |
| | 2 | No code read `$PORT` — a platform port would be ignored | `deploy/codespace/serve.py` binds `build_space_app()` to `$PORT`; Render reads its own `$PORT` | | |
| | 3 | Hand-rolled CORS; `OPTIONS` raised `405`, so browser preflight failed | the gateway registers `OPTIONS` explicitly (or relies on Starlette's CORS middleware) so preflight succeeds | | |
| | 4 | Module-level `app = create_app()` swallowed config errors into `app = None` | construction errors propagate (fail-fast) instead of silently leaving a dead `app` | | |
| | 5 | Adapter integrity unverified on load (`_adapter_sha256` computed but never compared) | the load path compares the computed digest against an expected value, or fails startup | | |
| ### 11.1 The blockers, with their original verification | |
| `docs/DEPLOYMENT_DECISION.md` §8 is the primary record, and it is more specific than the summary table: | |
| | # | Blocker | State (as recorded) | | |
| |---|---|---| | |
| | 1 | `requirements.txt` declares no fastapi / uvicorn / httpx / starlette | **VERIFIED** | | |
| | 2 | No code reads `$PORT` — a platform-assigned port would be ignored | **VERIFIED** | | |
| | 3 | CORS is hand-rolled (`gateway/policy.py:429-451`); `policy.py:591` admits OPTIONS but routes register only GET/HEAD/POST (`gateway/app.py:334-346`), so Starlette raises **405** and browser preflight fails | **VERIFIED** | | |
| | 4 | Module-level `app = create_app()` swallows config errors into `app = None` (`gateway/app.py:683-690`) | **VERIFIED** | | |
| | 5 | Adapter integrity unverified on load — `_adapter_sha256` is computed and stored (`:258`, `:273`) but never compared against an expected digest | **VERIFIED** | | |
| Blockers 3 and 4 are visible in the code this document cites. `gateway/app.py`'s own tail is blocker 4 | |
| exactly: | |
| ```python | |
| try: # pragma: no cover - depends on FastAPI being importable | |
| app = create_app() | |
| except Exception: # pragma: no cover - the sandbox path | |
| app = None # type: ignore[assignment] | |
| ``` | |
| and `deploy/render/main.py` is the fail-fast counterpart — its `create_app()` is called at module scope | |
| with no `try`, so a misconfiguration raises at import: | |
| ```python | |
| # The ASGI object uvicorn imports: `uvicorn deploy.render.main:app`. | |
| app = create_app() | |
| ``` | |
| The CORS half of blocker 3 is closed in `deploy/render/main.py` by registering | |
| `CORSMiddleware`, whose comment names the defect it fixes: | |
| > *"CORS fix: `CORSMiddleware` answers OPTIONS preflight itself, which resolves the earlier 405 on | |
| > preflight."* | |
| ### 11.2 Status: closed by construction, not proven in production | |
| `docs/DEPLOYMENT_TOPOLOGY.md` §4 is careful about the claim, and this document keeps that caution: | |
| > *"They are recorded honestly here — the new infra (`deploy/render/`, `deploy/codespace/`) is **in | |
| > progress**, so treat these as *closed by construction / to be verified on first live run*, not as | |
| > already proven in production."* | |
| **However**, the live deployment has since been exercised end-to-end: `docs/FINAL_DELIVERY_REPORT.md` §3 | |
| records `/api/health` 200, `/api/capabilities` 200 with 6× `available:true`, `/api/infer {}` → 422 | |
| `invalid_request` with `x-satquery-transport: tunnel`, and real inference for all six tasks. So the honest | |
| composite statement is: **the five blockers are closed in the deployed system as evidenced by the live | |
| behaviour recorded in the delivery documents, while `docs/DEPLOYMENT_TOPOLOGY.md` §4's own text still | |
| carries the earlier "in progress" framing.** Where the two disagree, the dated measurement is the stronger | |
| evidence, and it is cited here rather than substituted for the source's own words. | |
| --- | |
| ## 12. The superseded design, and exactly what did NOT change | |
| ### 12.1 The historical topology | |
| The superseded design ran inference on an **HF Space with ZeroGPU** (5 GPU-min/day, | |
| `@spaces.GPU(duration=…)` decoration) behind a **Railway** gateway | |
| (`docs/DEPLOYMENT_TOPOLOGY.md` §5; `docs/DEPLOYMENT_ARCHITECTURE.md` §1, §3). | |
| | Old (superseded) | New (active) | | |
| |---|---| | |
| | Railway (gateway/API) | **Render** (orchestrator / API gateway) | | |
| | Hugging Face Space (inference) | **GitHub Codespace** (FastAPI inference) | | |
| | Cloudflare Pages | Cloudflare Pages (**unchanged**) | | |
| | Hugging Face (project/models) | Hugging Face (project card + pinned model references) | | |
| Source: `docs/DEPLOYMENT_TOPOLOGY.md` §1. | |
| ### 12.2 The three things that changed | |
| `docs/DEPLOYMENT_TOPOLOGY.md` §5 enumerates them: | |
| 1. **CPU-first instead of ZeroGPU.** No code change was required — `device_preference` honours | |
| `SATQUERY_DEVICE` and defaults to CPU, every specialist defaults to `device="cpu"`, and all placement | |
| is `.to(device)` (never `.cuda()`). ZeroGPU's GPU-minute quota and `@spaces.GPU` decoration are no | |
| longer on the critical path. | |
| 2. **A real, always-buildable inference environment.** A GitHub Codespace gives a reproducible container | |
| that builds and runs `build_space_app()` without a GPU quota or a Space's ephemeral-cold-start | |
| constraint. The wake flow (§4) replaces ZeroGPU lazy-loading as the cold-start story. | |
| 3. **No GPU quota to protect at the gateway.** Because there is no ZeroGPU budget, the gateway's | |
| rate/size limits remain as *fairness* controls, but the "never spend GPU quota on a shape-rejected | |
| request" rationale no longer dominates the design. | |
| The CPU adaptation is independently verified in `docs/DEPLOYMENT_DECISION.md` §5, which lists the | |
| specific sites: `core/config.py:87-91` (`device_preference`), `specialists/vqa/model.py:234` | |
| (`float16` on cuda, **`float32` on cpu**), the per-specialist `device: str = "cpu"` defaults | |
| (`change/specialist.py:153`, `change/stanet.py:641`, `change/vqa_specialist.py:116`, | |
| `grounding/remoteclip.py:111`, `grounding/specialist.py:682`), `configs/base.yaml:293` | |
| (`cpu_mode_required: true`), and the fact that **no `.cuda()` call exists anywhere** — all placement is | |
| `.to(device)`. | |
| ### 12.3 What did NOT change | |
| `docs/DEPLOYMENT_TOPOLOGY.md` §5 closes with the list, and it is the most important part of this section: | |
| > *"**What did NOT change:** the 4-endpoint contract, the gateway responsibility table, the env-var | |
| > vocabulary (only host names moved: `SATQUERY_SPACE_URL` → `SATQUERY_UPSTREAM_URL`), and the | |
| > `Config.hash == 78f1e3700da15aa1` freeze. The backend contract in `DEPLOYMENT_ARCHITECTURE.md` §1.1, | |
| > §2, §3.3, §4, §5 remains authoritative."* | |
| Expanded: | |
| | Unchanged artefact | Where it lives | Why it survived the host change | | |
| |---|---|---| | |
| | **The 4-endpoint contract** | `docs/API_CONTRACT.md`; `app/space_app.py`; `gateway/app.py::PROXIED_ROUTES` | it is a *client-facing* contract; hosts are an implementation detail | | |
| | **The gateway responsibility table** | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 | the responsibilities are the same regardless of who hosts the upstream | | |
| | **The env-var vocabulary** | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 | only `SATQUERY_SPACE_URL` → `SATQUERY_UPSTREAM_URL` moved | | |
| | **The config freeze `78f1e3700da15aa1`** | `core/config.py::Config.hash`; `configs/base.yaml` | the deployment was changed *around* the config, never inside it | | |
| | **The gateway failure-mode table** | `docs/DEPLOYMENT_ARCHITECTURE.md` §5 | still governs, host names aside | | |
| | **The entrypoint requirements** | `docs/DEPLOYMENT_ARCHITECTURE.md` §3.3 | import cheaply without torch; reuse `app.serving`; degrade don't crash; never load a model for a metadata request; honour the config hash | | |
| ### 12.4 The frozen paperwork | |
| `configs/deploy.yaml` still describes an **HF Space + Gradio + ZeroGPU** target, and it is left | |
| **undisturbed** (`docs/DEPLOYMENT_TOPOLOGY.md` §3.4; `docs/DEPLOYMENT_DECISION.md` §4). The reasoning is | |
| structural, not sentimental, and it is a good example of why the config freeze matters: | |
| 1. **`Config.hash` cannot move.** `core/config.py` reads only `configs/base.yaml`. `configs/deploy.yaml` | |
| carries `registry: false` and is never loaded — **but** `scripts/validate_deploy_config.py` hard-fails | |
| if the `deployment:` block in `deploy.yaml` differs key-for-key from `base.yaml`'s (assertions at | |
| `:114-131`). So changing `zerogpu: true` → `false` in `deploy.yaml` alone fails the validator, and | |
| moving `base.yaml` to match moves the frozen hash. **Both paths are closed.** | |
| 2. **There is no Gradio runtime to conflict with.** No `import gradio`, no `gr.Blocks`, no `gr.Interface` | |
| and no Gradio entrypoint exists anywhere. Gradio appears only as `requirements.txt:36` and the manifest | |
| value `sdk: gradio` (`configs/base.yaml:284`). The one ZeroGPU code path — | |
| `spaces.GPU(duration=duration)` at `app/space_app.py:165` — sits inside `decorate_gpu()`, which **is | |
| never applied to any route**; routes use plain `@api.get`/`@api.post` at `:521/:549/:555/:661`. The real | |
| entrypoint is FastAPI: `build_space_app()` at `app/space_app.py:409`. | |
| > *"Conclusion: the frozen contract describes a Gradio Space that does not exist in code. It is frozen | |
| > paperwork, not a competing deployment."* (`docs/DEPLOYMENT_DECISION.md` §4) | |
| `configs/base.yaml` still carries the frozen ZeroGPU declarations, and `app/space_app.py` transcribes the | |
| durations into `GPU_DURATIONS`: | |
| ```python | |
| GPU_DURATIONS: dict[str, int] = { | |
| "vqa": 20, | |
| "caption": 20, | |
| "grounding": 45, | |
| "change": 30, | |
| "optical_sar": 45, | |
| "change_vqa": 30, | |
| } | |
| ``` | |
| with `change_vqa` reusing the `change` budget **because adding a key of its own would move `Config.hash`** | |
| (`docs/DEPLOYMENT_ARCHITECTURE.md` §3.4; `app/space_app.py`). And the decoration is applied conditionally, | |
| because `spaces` is not installed on a CPU host: | |
| ```python | |
| def decorate_gpu(task: str) -> Callable[[Callable[..., Any]], Callable[..., Any]]: | |
| ... | |
| spaces = _spaces_module() | |
| if spaces is None or not hasattr(spaces, "GPU"): | |
| def _identity(fn): return fn | |
| return _identity | |
| return spaces.GPU(duration=duration) | |
| ``` | |
| > **The honest status of the ZeroGPU path:** *"the ZeroGPU decoration has **never executed** here. It is | |
| > specified from finding C-8 and the frozen `gpu_duration_*` values, and that is all it is."* | |
| > (`app/space_app.py` docstring; `docs/PHASE19_FINAL_HARDENING.md`.) | |
| --- | |
| ## 13. Failure modes and their handling | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §5 is the authoritative table. Reproduced, with the active host names: | |
| | Failure | Detected by | Surface | Recovery | | |
| |---|---|---|---| | |
| | Upstream cold start | gateway upstream timeout | `504` with `recoverable: true` | client retries once, manually | | |
| | Model absent | `capabilities[].available: false` | `503 model_unavailable` | capability disabled in the UI | | |
| | Model corrupt | `ModelLoadError` | `503 model_load_error` | **defect** — report it | | |
| | GPU quota exhausted | allocation error | `503` | wait for the daily reset | | |
| | Request too large | gateway size check | `413` | client re-encodes | | |
| | Upload content type absent or refused | content-type allowlist | `415` | client sends a supported type; **the server does not guess** | | |
| | Uploaded handle expired or unknown | store lookup on read | `400 input_error` | re-upload; handles are ephemeral by design | | |
| | Asset store not configured or full | store construction / capacity check | `503` | ⚠️ **distinguishing these two needs an instrument the deployment does not expose** | | |
| | Asset root configured but unusable | **nothing** — `get_asset_store()` raises outside the route's `try` | **`500 text/plain`** on the upstream directly; the gateway masks it as `502` | ⚠️ bounded defect (F-12) | | |
| | Framework error upstream (`404`/`405`) | **nothing** on the upstream | `{"detail": …}` upstream; the gateway masks it as an envelope | ⚠️ bounded defect (F-12b) — **now fixed upstream too** (`app/space_app.py` registers the handler) | | |
| | Malformed body | gateway schema validation | `422` | client bug | | |
| | `trace.inputs` echoing a path | **nothing** | `200` with a server-side path | ✅ fixed (F-13) — `core/controller.py::_asset_label` | | |
| | `trace.steps[PARSE].detail["inputs"]` echoing the same path | **nothing** | `200` with a path in the `PARSE` step record | ✅ fixed (F-14) | | |
| | A construction failure's exception string reaching the client | **nothing** | `200` with a path in `result.warnings[]`, `evidence[].payload["message"]`, and the registry block **twice** | ✅ fixed (F-15) — path-scrubbed to a **basename**, raw detail logged server-side; **four** live carriers, not three | | |
| | `artifact_ref` / `result.change_map` carrying a path | **nothing**, and only when `artifact_dir` is configured | `200` with a path where the contract documents an `artifact://` URI | ✅ fixed (F-16) — refs are `null`, no `artifact://` fabricated, explicit non-retrievable warning | | |
| | `change_vqa.artifact_dir` configured but never read | **nothing** — the key is accepted and silently ignored | **no surface at all** | ⚠️ documented, not patched (F-17) | | |
| | Upstream unreachable | gateway connection error | `502` | report; do not silently retry analyze | | |
| | Analysis exceeds budget | `SpecialistTimeoutError` | `504`, `recoverable: true` | offer a retry | | |
| | Non-JSON response upstream | gateway parse check | `502` with the upstream body logged | **defect** | | |
| | Per-IP rate limit bypassed | **not detected** | no `429` is produced | fairness only; **not** a protection control (§2.4) | | |
| ### 13.1 Two failure modes that the gateway and the upstream now agree on | |
| The `404`/`405` and unhandled-exception rows were originally *upstream* holes that the gateway masked. | |
| Both are now closed **on the upstream as well**, so a client following the runbook to the upstream's own | |
| URL gets the same envelope as a client going through the gateway. `app/space_app.py` registers both | |
| handlers, and its comment records the measurement that forced it: | |
| > *"Measured, direct to the Space, before this fix: `GET /v1/whocares -> 404 {"detail":"Not Found"}`, | |
| > `GET /v1/assets -> 405 {"detail":"Method Not Allowed"}`, an unwrapped failure -> `500 text/plain`, no | |
| > envelope at all."* | |
| ### 13.2 A saturated asset store is indistinguishable from a misconfigured one | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 records finding F-11 and its resolution. `POST /v1/assets` answers | |
| `503` in two unrelated situations — the store is **not configured**, or the store is **full** — with the | |
| same status and the same envelope shape, so *"a client and an operator cannot tell them apart from a | |
| response."* | |
| The one value that would have separated them (`capacity_refusals` from `AssetStore.stats()`) was computed | |
| on every request and read by nothing. **RESOLVED 2026-09-23 by owner ruling — the unused computation was | |
| REMOVED, not given a consumer.** The owner's reasoning: *"a metrics surface with no reader is a cost paid | |
| on every request for an instrument nobody holds."* The ambiguity itself **remains**, and the document says | |
| so: | |
| > *"Removing the counter did NOT remove the ambiguity. The two `503` causes remain indistinguishable from a | |
| > response, and the deployment still **does not expose** an instrument that tells them apart."* | |
| The remedy is unchanged: `SATQUERY_ASSET_MAX_FILES` / `SATQUERY_ASSET_TTL_S` if the store is saturating, | |
| and those two variables if it is unconfigured — but **confirming which requires inspecting the | |
| deployment**, because the response will not say. | |
| ### 13.3 A deployment precondition list, carried forward | |
| `docs/DEPLOYMENT_TOPOLOGY.md` §6, with host names updated: | |
| 1. Cloudflare Pages project name / domain (needed for the deploy command and `robots.txt` sitemap). | |
| 2. Artifacts present, or capabilities shipped `available: false` (change head, change_vqa head, | |
| calibration JSON) — degrades honestly, not broken. | |
| 3. `HF_TOKEN` set on Render **if** the HF proxy path is used (not used in the live config). | |
| 4. Codespace `.devcontainer/` forwarding `:8000` **and** starting the tunnel agent. | |
| 5. The five blockers in §11 closed and verified on the first live run. | |
| --- | |
| ## 14. What is deliberately absent from the deployment | |
| `docs/DEPLOYMENT_ARCHITECTURE.md` §6 records the exclusions so that omission is not mistaken for | |
| oversight. From plan §73/§74: | |
| * **No Kubernetes, no Docker swarm.** Render plus one Codespace is the whole fleet. | |
| * **No Kafka, no Redis cluster, no queue.** Requests are synchronous. | |
| * **No autoscaling.** The free tier has a fixed quota; autoscaling cannot raise it. | |
| * **No multi-tenant isolation, no auth, no user accounts.** | |
| * **No second VLM and no foundation-model retraining.** | |
| * **No vector database.** The retriever-free RAG decision is separate and upstream. | |
| * **No database, no session store** at the gateway (`docs/DEPLOYMENT_ARCHITECTURE.md` §2.2). | |
| `docs/FRONTEND_INTEGRATION.md` §7 adds the client-side counterpart: **no login screen**, because *"There | |
| is none to build (plan §74)."* | |
| --- | |
| ## 15. What is `NOT RUN` / `OPEN` / `BLOCKED` for this topic | |
| | Item | Status | Note | | |
| |---|---|---| | |
| | **B-07** transient tunnel-agent gaps | **OPEN** | patch prepared, **not deployed**; worst case ≈ 249 s (§6) | | |
| | **B-02** `codespace_name` trailing `\n` | **OPEN (cosmetic)** | reporting only; the wake path strips (§7.4) | | |
| | **B-03** local `deploy/` stale + untracked | **KNOWN** | the working copy is not the deployed source (§9.5) | | |
| | **B-06** Render free-tier sleep / Codespace idle 30 min | **KNOWN** | cold start delay; documented, not hidden (§5) | | |
| | Tunnel agent source (`tunnel_agent.py`) | **UNKNOWN** | not in the monorepo; `UNKNOWN — not established from the available evidence` (§9.4) | | |
| | Cold-start latency distribution | **NOT MEASURED** | "tens of seconds" is a documented expectation; no distribution exists (§5) | | |
| | Throughput / concurrency characterisation | **NOT RUN** | `docs/FRONTEND_INTEGRATION.md` §9: *"Latency is not characterized."* | | |
| | ZeroGPU decoration execution | **NOT RUN** | never executed anywhere; CPU path only (`app/space_app.py`; `docs/PHASE19_FINAL_HARDENING.md`) | | |
| | Sequential-request test under `cache_max_models=1` | **NOT DONE — environment-blocked** | requires a reachable upstream (`docs/DEPLOYMENT_ARCHITECTURE.md` §7) | | |
| | Gateway's rate limiter as an abuse control | **REJECTED** | ruled fairness-only, 2026-09-23 (§2.4) | | |
| | `HF_TOKEN` proxy path | **not used live** | absent from the live Render config (§8.1) | | |
| | `SATQUERY_UPSTREAM_URL` | **not used live** | absent from the live Render config (§8.1) | | |
| | A second copy of the capability table at the gateway | **REJECTED** | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 | | |
| | End-to-end benchmark of the deployed stack | **does not exist** | no system-level accuracy is claimed anywhere | | |
| | Asset-store 503 disambiguation instrument | **absent** | removed by ruling; ambiguity remains (§13.2) | | |
| --- | |
| ## 16. Where the evidence lives | |
| | Claim | Source | | |
| |---|---| | |
| | four tiers, host names, tunnel direction | `docs/DEPLOYMENT_TOPOLOGY.md` §1, §2; `docs/FINAL_DELIVERY_TODO.md` §1.2 | | |
| | gateway rationale (three reasons) | `docs/DEPLOYMENT_ARCHITECTURE.md` §1.1 | | |
| | gateway responsibility table | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.1 | | |
| | what the gateway must NOT do | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.2 | | |
| | CORS assembly + wildcard refusal | `deploy/render/main.py::_allowed_origins`, `_DEV_ORIGINS`, `_PRODUCTION_ORIGINS` | | |
| | CORS header filtering assertion | `gateway/app.py::_proxy` (F-2) | | |
| | per-file cap 4,194,304 B | `gateway/policy.py:221`; `app/space_app.py::_asset_max_file_bytes` | | |
| | body cap 8 MiB + F-6 measurement | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 | | |
| | rate-limiter ruling + measurement | `docs/DEPLOYMENT_ARCHITECTURE.md` §5.2 | | |
| | no-retry rule | `gateway/app.py::_proxy`; `docs/DEPLOYMENT_TOPOLOGY.md` §2 | | |
| | request-id injection | `gateway/app.py::_proxy` | | |
| | 404/405 envelope handler | `gateway/app.py`; `app/space_app.py` | | |
| | error-envelope shape | `docs/DEPLOYMENT_ARCHITECTURE.md` §2.3 | | |
| | orchestrator error classes + statuses | `deploy/render/main.py` | | |
| | transport-failure classification | `gateway/app.py::_transport_failure_detail` (F-15c) | | |
| | proxied / costly / blocked route tuples | `gateway/app.py::PROXIED_ROUTES`, `COSTLY_ROUTES`, `BLOCKED_ROUTES` | | |
| | forwarded port returns 302 | `docs/DEPLOYMENT_TOPOLOGY.md` measured note; `release/repo/docs/DEPLOYMENT.md` §7 | | |
| | tunnel rationale for private repos | `deploy/codespace/launch.sh` header comment | | |
| | `x-satquery-transport` as proof | `tests/unit/test_frontend_live_wiring.py`; `docs/FINAL_DELIVERY_TODO.md` §6 E-03 | | |
| | wake flow | `docs/DEPLOYMENT_TOPOLOGY.md` §2; `deploy/render/main.py::ensure_codespace_up` | | |
| | `X-SatQuery-State` header | `deploy/render/main.py::infer` | | |
| | cold start, documented not hidden | `docs/DEPLOYMENT_TOPOLOGY.md` §2; `docs/FRONTEND_INTEGRATION.md` §6 | | |
| | Codespace idle 30 min | `docs/FINAL_DELIVERY_TODO.md` §6 E-04 | | |
| | B-07 root shape + ≈249 s | `docs/FINAL_DELIVERY_TODO.md` §5; `DELIVERY_REPORT_2026-09-25.md` §4 | | |
| | B-07 patch contents + verification | `DELIVERY_REPORT_2026-09-25.md` §4; `docs/FINAL_DELIVERY_TODO.md` §6 E-12 | | |
| | live health payload | `release/repo/docs/DEPLOYMENT.md` §2 | | |
| | `completed` readings 97 / 314 / 338 | `release/repo/docs/DEPLOYMENT.md` §2; `docs/FINAL_DELIVERY_TODO.md` §1.4; `docs/FINAL_DELIVERY_REPORT.md` §3 | | |
| | B-02 cosmetic | `docs/FINAL_DELIVERY_REPORT.md` §6; `docs/FINAL_DELIVERY_TODO.md` §5 | | |
| | live Render env vars | `docs/DEPLOYMENT_TOPOLOGY.md` header note; `release/repo/docs/DEPLOYMENT.md` §3.1 | | |
| | blueprint env vars | `render.yaml` | | |
| | Codespace env vars | `.devcontainer/devcontainer.json`; `deploy/codespace/launch.sh` | | |
| | env-var vocabulary + F-7/F-8/F-9 | `docs/DEPLOYMENT_ARCHITECTURE.md` §4 | | |
| | `serve.py` entrypoint | `deploy/codespace/serve.py` | | |
| | launcher stages 0–3 | `deploy/codespace/launch.sh` | | |
| | tunnel agent existence + gap | `docs/FINAL_DELIVERY_TODO.md` §1.3; `Glob`/`git ls-files` on the monorepo | | |
| | repos of record | `docs/FINAL_DELIVERY_TODO.md` §1.3 | | |
| | deployed revisions | `release/repo/docs/DEPLOYMENT.md` §1 | | |
| | staging measurement | `docs/DEPLOYMENT_DECISION.md` §7 | | |
| | Git Data API + sha256 verification | `release/repo/docs/DEPLOYMENT.md` §4, §4.1 | | |
| | five blockers + original verification | `docs/DEPLOYMENT_DECISION.md` §8; `docs/DEPLOYMENT_TOPOLOGY.md` §4 | | |
| | superseded design + what did not change | `docs/DEPLOYMENT_TOPOLOGY.md` §5 | | |
| | frozen paperwork | `docs/DEPLOYMENT_DECISION.md` §4; `docs/DEPLOYMENT_TOPOLOGY.md` §3.4 | | |
| | `GPU_DURATIONS` | `app/space_app.py` | | |
| | failure modes | `docs/DEPLOYMENT_ARCHITECTURE.md` §5 | | |
| | asset-store ambiguity (F-11) | `docs/DEPLOYMENT_ARCHITECTURE.md` §5.1 | | |
| | deliberate exclusions | `docs/DEPLOYMENT_ARCHITECTURE.md` §6; `docs/FRONTEND_INTEGRATION.md` §7 | | |
| --- | |
| *Continue to [03 — Request lifecycle](03-request-lifecycle.md).* | |