# Architecture
**Status vocabulary used throughout these docs:** `IMPLEMENTED` (code exists and runs) ·
`VERIFIED` (checked against evidence) · `MEASURED` (a number was produced) · `ATTEMPTED`
(tried, outcome recorded) · `NOT RUN` · `BLOCKED` · `DEFERRED` · `REJECTED` (tried and
explicitly not accepted). Every architectural claim below carries one of these tags.
---
## 1. What the system is
SatQuery AI answers natural-language questions about satellite imagery. It is a **router +
specialists** system: a small intent router reads the question, dispatches it to one of six
task specialists, and assembles the specialist's output into a single structured
`ResultEnvelope` that the browser renders as an answer, an evidence list, and a confidence.
It is not a single end-to-end vision-language model. The only learned components are:
- a **50,822-parameter router adapter** over a frozen `all-MiniLM-L6-v2` sentence encoder;
- four trained **task heads/adapters** (change, change_vqa, optical-SAR fusion, grounding);
- one **LoRA adapter** over a frozen `SmolVLM-500M-Instruct` for caption/VQA.
Everything else is frozen, publicly-pinned backbone weights resolved at run time.
## 2. Deployment topology (IMPLEMENTED, VERIFIED)
The shipped system is a four-tier chain. **Status: VERIFIED live on 2026-09-25** — the health,
capabilities and infer endpoints were probed against the running deployment.
```
Browser
│ HTTPS
▼
Cloudflare Pages — satquery.pages.dev (static frontend, 11 pages)
│ HTTPS/JSON
▼
Render — satquery-backend-m4yv.onrender.com (orchestrator / gateway)
│ outbound long-poll tunnel POST /tunnel/agent
▼
GitHub Codespace — FastAPI inference, port 8000 (CPU)
│ build_space_app()
▼
Specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
```
```mermaid
flowchart LR
U[Browser] -->|HTTPS| CF["Cloudflare Pages
static frontend"]
CF -->|"HTTPS JSON
/api/health · /api/capabilities · /api/infer"| R["Render
orchestrator / gateway"]
R -->|"outbound long-poll
POST /tunnel/agent"| C["GitHub Codespace
FastAPI inference :8000"]
C --> S[(SmolVLM · RemoteCLIP
MiniLM · CROMA · STANet)]
C -->|ResultEnvelope| R
R -->|"envelope + error translation"| CF
```
### 2.1 Why a gateway and a tunnel
- **The gateway (Render)** is deliberately thin and stateless: schema validation, size limits,
per-IP rate limiting, a CORS allowlist (**never** `*`), request IDs, timeouts, secret custody,
and translation of upstream failures into the documented error envelope. It is **not** a model
host and has no database, auth, or queue.
- **The tunnel** exists because the inference host (a Codespace) is not publicly addressable for
this private repository — a forwarded port returns `302`. Instead the Codespace runs
`tunnel_agent.py`, which **dials out** to `POST /tunnel/agent` and long-polls. Work is executed
against `http://127.0.0.1:8000` locally. This inverts the usual direction: the inference host
needs no inbound firewall hole.
> **Status:** IMPLEMENTED and VERIFIED. `GET /api/health` reported `tunnel.agent_connected:true`
> with a non-zero `completed` counter on 2026-09-25.
### 2.2 Cold start and the `transport_mode: auto` fallthrough (KNOWN ISSUE, OPEN)
The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore
tens of seconds and is **documented, not hidden**. A timeout on the tunnel path is surfaced to the
client as a recoverable error.
`SATQUERY_TRANSPORT=auto` means: try the tunnel; if it does not answer within
`SATQUERY_TUNNEL_TIMEOUT_S` (150 s), **fall through to the forward path**. The forward path to a
private repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S`
(120 s) first — so a worst-case failed request can take ≈ 249 s. This is the root of the observed
"transient tunnel gap" and is recorded as **OPEN** (not fixed) in
[`LIMITATIONS.md`](LIMITATIONS.md). A deployed fix (`codespace_name` newline handling on the
wake path) was authored and pushed separately; the fallthrough itself is by design.
## 3. Request lifecycle (IMPLEMENTED, VERIFIED)
The controller is a nine-state machine, declared in `configs/base.yaml`:
```
RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND
```
1. **RECEIVE / PARSE** — the query and 1–2 image assets are decoded.
2. **VALIDATE** — modality inference from band count, dimension equality for pairs, byte cap
(`4,194,304` bytes/file), GeoTIFF acceptance for the optical-SAR path.
3. **PLAN** — the router embeds the query with the frozen MiniLM encoder and the trained adapter
predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change
task.
4. **PREPROCESS** — percentile normalisation for optical (2/98), dB clip for SAR (−30…+5 dB),
tiling for large scenes.
5. **EXECUTE** — the chosen specialist runs.
6. **AGGREGATE** — the specialist's raw output becomes `Evidence` items
(`Region` / `ChangeRegion`).
7. **VERIFY** — confidence is computed; temperature scaling is applied from
`calibration_v001.json`.
8. **RESPOND** — a `ResultEnvelope` is assembled and returned.
### 3.1 The eight execution events
The frontend renders a live trace from eight named events (source of truth:
`frontend/assets/js/core.js`, `SQ.EVENT_NAMES`):
| # | Event | Emitted when |
|---|---|---|
| 1 | `QUERY_RECEIVED` | the request is accepted |
| 2 | `QUERY_UNDERSTOOD` | the router produces task + modality |
| 3 | `ROUTE_SELECTED` | the specialist is chosen |
| 4 | `SPECIALIST_STARTED` | specialist execution begins |
| 5 | `SPECIALIST_COMPLETED` | specialist returns |
| 6 | `EVIDENCE_GENERATED` | evidence items are built |
| 7 | `CONFIDENCE_COMPUTED` | confidence (and calibration) is computed |
| 8 | `RESULT_ASSEMBLED` | the envelope is finalised |
**Live runs emit all eight; the offline preview path emits none of the specialist events and is
labelled as a preview in the UI.** Measured trace fill on every live run: **94.4444 %**.
## 4. Two reading paths: `interpret()` vs `chooseTask()` (IMPLEMENTED, VERIFIED)
There are two separate functions, and they see different information:
- **`interpret()`** — produces the *human-readable* interpretation shown in the console. It is
**asset-count-blind**: it only sees the query text.
- **`chooseTask()`** — performs the *dispatch*. It is **asset-count-aware**: it knows how many
assets are attached and will not route a single-image query to a two-image task.
This asymmetry is real and explains a documented behaviour: for the query *"What changed between
the earlier and later image?"* with a single asset attached, the console can *read* `change` while
dispatch correctly falls back to `change_vqa`. This is intentional, not a bug, and is covered in
[`RESEARCH_NOTES.md`](RESEARCH_NOTES.md).
## 5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED)
The Codespace exposes exactly four routes; the gateway mirrors each under `/api/*`:
| Codespace | Gateway | Cost | Purpose |
|---|---|---|---|
| `GET /v1/health` | `GET /api/health` | cheap | liveness, tunnel state, counters |
| `GET /v1/capabilities` | `GET /api/capabilities` | cheap | per-task availability |
| `POST /v1/analyze` | `POST /api/infer` | **costly** | run a query |
| `POST /v1/assets` | `POST /api/assets` | **costly** | stage an asset handle |
The gateway must **not** retry `POST /api/infer` on its own — a retry would consume inference a
second time. The client decides on retry. Errors are translated into a stable envelope with a
machine code (`invalid_request`, `tunnel_offline`, …) and a `recoverable` flag.
**Verified:** `POST /api/infer {}` returns `422 invalid_request` with header
`x-satquery-transport: tunnel`.
## 6. The frozen configuration (VERIFIED)
All tunables live in `configs/base.yaml`; **no magic numbers in Python**. The loader validates
invariants and computes a hash. The frozen hash is:
```
78f1e3700da15aa1
```
The loader **refuses to run** a config that violates a recorded invariant — for example
`fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels`
(2318) and `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (2048). A mismatch
in the latter is a *silent* shape error otherwise (torch raises only later, after features are
cached), so it is enforced at load time.
Editing `configs/base.yaml` moves the hash, which invalidates every artifact keyed to it. This is
why the stale `configs/deploy.yaml` (which still describes an HF-Space/ZeroGPU target) is left
**undisturbed as frozen paperwork** — it is never read at run time.
## 7. Component map
| Layer | Path (in the source repo) | Notes |
|---|---|---|
| Frontend | `frontend/` | static, staged by `scripts/stage_pages.mjs` |
| Gateway | `deploy/render/` | `main.py` exposes `app`; `render.yaml` blueprint |
| Inference entry | `deploy/codespace/serve.py` | serves `build_space_app()` on `$PORT` |
| Inference app | `app/space_app.py` | `build_space_app()`; 4-endpoint contract |
| Composition root | `app/serving.py` | reuses the local serving path |
| Specialists | `specialists/` | vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA) |
| Router | `router/` | frozen MiniLM + trained adapter |
| Config | `configs/base.yaml` | authoritative registry; hashed |
| Artifacts | `artifacts/` | weights, metrics, calibration, reports |
> **Caveat.** `deploy/` inside the monorepo is **stale and untracked**; the deployed sources live in
> three separate private repositories. See [`DEPLOYMENT.md`](DEPLOYMENT.md).
## 8. What is deliberately absent
- **No database, no auth, no queue.** The gateway is stateless by design.
- **No GPU requirement.** Device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`,
never `.cuda()`. The shipped deployment runs CPU-only.
- **No system-level end-to-end benchmark.** There is no measured end-to-end accuracy number, and
none is claimed. Per-specialist metrics exist and are reported in [`BENCHMARKS.md`](BENCHMARKS.md).