# Architecture **Status vocabulary used throughout these docs:** `IMPLEMENTED` (code exists and runs) · `VERIFIED` (checked against evidence) · `MEASURED` (a number was produced) · `ATTEMPTED` (tried, outcome recorded) · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` (tried and explicitly not accepted). Every architectural claim below carries one of these tags. --- ## 1. What the system is SatQuery AI answers natural-language questions about satellite imagery. It is a **router + specialists** system: a small intent router reads the question, dispatches it to one of six task specialists, and assembles the specialist's output into a single structured `ResultEnvelope` that the browser renders as an answer, an evidence list, and a confidence. It is not a single end-to-end vision-language model. The only learned components are: - a **50,822-parameter router adapter** over a frozen `all-MiniLM-L6-v2` sentence encoder; - four trained **task heads/adapters** (change, change_vqa, optical-SAR fusion, grounding); - one **LoRA adapter** over a frozen `SmolVLM-500M-Instruct` for caption/VQA. Everything else is frozen, publicly-pinned backbone weights resolved at run time. ## 2. Deployment topology (IMPLEMENTED, VERIFIED) The shipped system is a four-tier chain. **Status: VERIFIED live on 2026-09-25** — the health, capabilities and infer endpoints were probed against the running deployment. ``` Browser │ HTTPS ▼ Cloudflare Pages — satquery.pages.dev (static frontend, 11 pages) │ HTTPS/JSON ▼ Render — satquery-backend-m4yv.onrender.com (orchestrator / gateway) │ outbound long-poll tunnel POST /tunnel/agent ▼ GitHub Codespace — FastAPI inference, port 8000 (CPU) │ build_space_app() ▼ Specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet ``` ```mermaid flowchart LR U[Browser] -->|HTTPS| CF["Cloudflare Pages
static frontend"] CF -->|"HTTPS JSON
/api/health · /api/capabilities · /api/infer"| R["Render
orchestrator / gateway"] R -->|"outbound long-poll
POST /tunnel/agent"| C["GitHub Codespace
FastAPI inference :8000"] C --> S[(SmolVLM · RemoteCLIP
MiniLM · CROMA · STANet)] C -->|ResultEnvelope| R R -->|"envelope + error translation"| CF ``` ### 2.1 Why a gateway and a tunnel - **The gateway (Render)** is deliberately thin and stateless: schema validation, size limits, per-IP rate limiting, a CORS allowlist (**never** `*`), request IDs, timeouts, secret custody, and translation of upstream failures into the documented error envelope. It is **not** a model host and has no database, auth, or queue. - **The tunnel** exists because the inference host (a Codespace) is not publicly addressable for this private repository — a forwarded port returns `302`. Instead the Codespace runs `tunnel_agent.py`, which **dials out** to `POST /tunnel/agent` and long-polls. Work is executed against `http://127.0.0.1:8000` locally. This inverts the usual direction: the inference host needs no inbound firewall hole. > **Status:** IMPLEMENTED and VERIFIED. `GET /api/health` reported `tunnel.agent_connected:true` > with a non-zero `completed` counter on 2026-09-25. ### 2.2 Cold start and the `transport_mode: auto` fallthrough (KNOWN ISSUE, OPEN) The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore tens of seconds and is **documented, not hidden**. A timeout on the tunnel path is surfaced to the client as a recoverable error. `SATQUERY_TRANSPORT=auto` means: try the tunnel; if it does not answer within `SATQUERY_TUNNEL_TIMEOUT_S` (150 s), **fall through to the forward path**. The forward path to a private repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S` (120 s) first — so a worst-case failed request can take ≈ 249 s. This is the root of the observed "transient tunnel gap" and is recorded as **OPEN** (not fixed) in [`LIMITATIONS.md`](LIMITATIONS.md). A deployed fix (`codespace_name` newline handling on the wake path) was authored and pushed separately; the fallthrough itself is by design. ## 3. Request lifecycle (IMPLEMENTED, VERIFIED) The controller is a nine-state machine, declared in `configs/base.yaml`: ``` RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND ``` 1. **RECEIVE / PARSE** — the query and 1–2 image assets are decoded. 2. **VALIDATE** — modality inference from band count, dimension equality for pairs, byte cap (`4,194,304` bytes/file), GeoTIFF acceptance for the optical-SAR path. 3. **PLAN** — the router embeds the query with the frozen MiniLM encoder and the trained adapter predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change task. 4. **PREPROCESS** — percentile normalisation for optical (2/98), dB clip for SAR (−30…+5 dB), tiling for large scenes. 5. **EXECUTE** — the chosen specialist runs. 6. **AGGREGATE** — the specialist's raw output becomes `Evidence` items (`Region` / `ChangeRegion`). 7. **VERIFY** — confidence is computed; temperature scaling is applied from `calibration_v001.json`. 8. **RESPOND** — a `ResultEnvelope` is assembled and returned. ### 3.1 The eight execution events The frontend renders a live trace from eight named events (source of truth: `frontend/assets/js/core.js`, `SQ.EVENT_NAMES`): | # | Event | Emitted when | |---|---|---| | 1 | `QUERY_RECEIVED` | the request is accepted | | 2 | `QUERY_UNDERSTOOD` | the router produces task + modality | | 3 | `ROUTE_SELECTED` | the specialist is chosen | | 4 | `SPECIALIST_STARTED` | specialist execution begins | | 5 | `SPECIALIST_COMPLETED` | specialist returns | | 6 | `EVIDENCE_GENERATED` | evidence items are built | | 7 | `CONFIDENCE_COMPUTED` | confidence (and calibration) is computed | | 8 | `RESULT_ASSEMBLED` | the envelope is finalised | **Live runs emit all eight; the offline preview path emits none of the specialist events and is labelled as a preview in the UI.** Measured trace fill on every live run: **94.4444 %**. ## 4. Two reading paths: `interpret()` vs `chooseTask()` (IMPLEMENTED, VERIFIED) There are two separate functions, and they see different information: - **`interpret()`** — produces the *human-readable* interpretation shown in the console. It is **asset-count-blind**: it only sees the query text. - **`chooseTask()`** — performs the *dispatch*. It is **asset-count-aware**: it knows how many assets are attached and will not route a single-image query to a two-image task. This asymmetry is real and explains a documented behaviour: for the query *"What changed between the earlier and later image?"* with a single asset attached, the console can *read* `change` while dispatch correctly falls back to `change_vqa`. This is intentional, not a bug, and is covered in [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md). ## 5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED) The Codespace exposes exactly four routes; the gateway mirrors each under `/api/*`: | Codespace | Gateway | Cost | Purpose | |---|---|---|---| | `GET /v1/health` | `GET /api/health` | cheap | liveness, tunnel state, counters | | `GET /v1/capabilities` | `GET /api/capabilities` | cheap | per-task availability | | `POST /v1/analyze` | `POST /api/infer` | **costly** | run a query | | `POST /v1/assets` | `POST /api/assets` | **costly** | stage an asset handle | The gateway must **not** retry `POST /api/infer` on its own — a retry would consume inference a second time. The client decides on retry. Errors are translated into a stable envelope with a machine code (`invalid_request`, `tunnel_offline`, …) and a `recoverable` flag. **Verified:** `POST /api/infer {}` returns `422 invalid_request` with header `x-satquery-transport: tunnel`. ## 6. The frozen configuration (VERIFIED) All tunables live in `configs/base.yaml`; **no magic numbers in Python**. The loader validates invariants and computes a hash. The frozen hash is: ``` 78f1e3700da15aa1 ``` The loader **refuses to run** a config that violates a recorded invariant — for example `fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels` (2318) and `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (2048). A mismatch in the latter is a *silent* shape error otherwise (torch raises only later, after features are cached), so it is enforced at load time. Editing `configs/base.yaml` moves the hash, which invalidates every artifact keyed to it. This is why the stale `configs/deploy.yaml` (which still describes an HF-Space/ZeroGPU target) is left **undisturbed as frozen paperwork** — it is never read at run time. ## 7. Component map | Layer | Path (in the source repo) | Notes | |---|---|---| | Frontend | `frontend/` | static, staged by `scripts/stage_pages.mjs` | | Gateway | `deploy/render/` | `main.py` exposes `app`; `render.yaml` blueprint | | Inference entry | `deploy/codespace/serve.py` | serves `build_space_app()` on `$PORT` | | Inference app | `app/space_app.py` | `build_space_app()`; 4-endpoint contract | | Composition root | `app/serving.py` | reuses the local serving path | | Specialists | `specialists/` | vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA) | | Router | `router/` | frozen MiniLM + trained adapter | | Config | `configs/base.yaml` | authoritative registry; hashed | | Artifacts | `artifacts/` | weights, metrics, calibration, reports | > **Caveat.** `deploy/` inside the monorepo is **stale and untracked**; the deployed sources live in > three separate private repositories. See [`DEPLOYMENT.md`](DEPLOYMENT.md). ## 8. What is deliberately absent - **No database, no auth, no queue.** The gateway is stateless by design. - **No GPU requirement.** Device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`, never `.cuda()`. The shipped deployment runs CPU-only. - **No system-level end-to-end benchmark.** There is no measured end-to-end accuracy number, and none is claimed. Per-specialist metrics exist and are reported in [`BENCHMARKS.md`](BENCHMARKS.md).