SatQuery / docs /ARCHITECTURE.md
thundercode's picture
release: add docs/ARCHITECTURE.md
724b09e verified
|
Raw History Blame
10.2 kB

Architecture

Status vocabulary used throughout these docs: IMPLEMENTED (code exists and runs) · VERIFIED (checked against evidence) · MEASURED (a number was produced) · ATTEMPTED (tried, outcome recorded) · NOT RUN · BLOCKED · DEFERRED · REJECTED (tried and explicitly not accepted). Every architectural claim below carries one of these tags.


1. What the system is

SatQuery AI answers natural-language questions about satellite imagery. It is a router + specialists system: a small intent router reads the question, dispatches it to one of six task specialists, and assembles the specialist's output into a single structured ResultEnvelope that the browser renders as an answer, an evidence list, and a confidence.

It is not a single end-to-end vision-language model. The only learned components are:

  • a 50,822-parameter router adapter over a frozen all-MiniLM-L6-v2 sentence encoder;
  • four trained task heads/adapters (change, change_vqa, optical-SAR fusion, grounding);
  • one LoRA adapter over a frozen SmolVLM-500M-Instruct for caption/VQA.

Everything else is frozen, publicly-pinned backbone weights resolved at run time.

2. Deployment topology (IMPLEMENTED, VERIFIED)

The shipped system is a four-tier chain. Status: VERIFIED live on 2026-09-25 — the health, capabilities and infer endpoints were probed against the running deployment.

Browser
  │ HTTPS
  ▼
Cloudflare Pages  —  satquery.pages.dev          (static frontend, 11 pages)
  │ HTTPS/JSON
  ▼
Render            —  satquery-backend-m4yv.onrender.com   (orchestrator / gateway)
  │ outbound long-poll tunnel  POST /tunnel/agent
  ▼
GitHub Codespace  —  FastAPI inference, port 8000 (CPU)
  │ build_space_app()
  ▼
Specialists: SmolVLM · RemoteCLIP · MiniLM · CROMA · STANet
flowchart LR
  U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"]
  CF -->|"HTTPS JSON<br/>/api/health · /api/capabilities · /api/infer"| R["Render<br/>orchestrator / gateway"]
  R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"]
  C --> S[(SmolVLM · RemoteCLIP<br/>MiniLM · CROMA · STANet)]
  C -->|ResultEnvelope| R
  R -->|"envelope + error translation"| CF

2.1 Why a gateway and a tunnel

  • The gateway (Render) is deliberately thin and stateless: schema validation, size limits, per-IP rate limiting, a CORS allowlist (never *), request IDs, timeouts, secret custody, and translation of upstream failures into the documented error envelope. It is not a model host and has no database, auth, or queue.
  • The tunnel exists because the inference host (a Codespace) is not publicly addressable for this private repository — a forwarded port returns 302. Instead the Codespace runs tunnel_agent.py, which dials out to POST /tunnel/agent and long-polls. Work is executed against http://127.0.0.1:8000 locally. This inverts the usual direction: the inference host needs no inbound firewall hole.

Status: IMPLEMENTED and VERIFIED. GET /api/health reported tunnel.agent_connected:true with a non-zero completed counter on 2026-09-25.

2.2 Cold start and the transport_mode: auto fallthrough (KNOWN ISSUE, OPEN)

The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore tens of seconds and is documented, not hidden. A timeout on the tunnel path is surfaced to the client as a recoverable error.

SATQUERY_TRANSPORT=auto means: try the tunnel; if it does not answer within SATQUERY_TUNNEL_TIMEOUT_S (150 s), fall through to the forward path. The forward path to a private repo returns 302 quickly, but the wake step still consumes SATQUERY_WAKE_TIMEOUT_S (120 s) first — so a worst-case failed request can take ≈ 249 s. This is the root of the observed "transient tunnel gap" and is recorded as OPEN (not fixed) in LIMITATIONS.md. A deployed fix (codespace_name newline handling on the wake path) was authored and pushed separately; the fallthrough itself is by design.

3. Request lifecycle (IMPLEMENTED, VERIFIED)

The controller is a nine-state machine, declared in configs/base.yaml:

RECEIVE → PARSE → VALIDATE → PLAN → PREPROCESS → EXECUTE → AGGREGATE → VERIFY → RESPOND
  1. RECEIVE / PARSE — the query and 1–2 image assets are decoded.
  2. VALIDATE — modality inference from band count, dimension equality for pairs, byte cap (4,194,304 bytes/file), GeoTIFF acceptance for the optical-SAR path.
  3. PLAN — the router embeds the query with the frozen MiniLM encoder and the trained adapter predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change task.
  4. PREPROCESS — percentile normalisation for optical (2/98), dB clip for SAR (−30…+5 dB), tiling for large scenes.
  5. EXECUTE — the chosen specialist runs.
  6. AGGREGATE — the specialist's raw output becomes Evidence items (Region / ChangeRegion).
  7. VERIFY — confidence is computed; temperature scaling is applied from calibration_v001.json.
  8. RESPOND — a ResultEnvelope is assembled and returned.

3.1 The eight execution events

The frontend renders a live trace from eight named events (source of truth: frontend/assets/js/core.js, SQ.EVENT_NAMES):

# Event Emitted when
1 QUERY_RECEIVED the request is accepted
2 QUERY_UNDERSTOOD the router produces task + modality
3 ROUTE_SELECTED the specialist is chosen
4 SPECIALIST_STARTED specialist execution begins
5 SPECIALIST_COMPLETED specialist returns
6 EVIDENCE_GENERATED evidence items are built
7 CONFIDENCE_COMPUTED confidence (and calibration) is computed
8 RESULT_ASSEMBLED the envelope is finalised

Live runs emit all eight; the offline preview path emits none of the specialist events and is labelled as a preview in the UI. Measured trace fill on every live run: 94.4444 %.

4. Two reading paths: interpret() vs chooseTask() (IMPLEMENTED, VERIFIED)

There are two separate functions, and they see different information:

  • interpret() — produces the human-readable interpretation shown in the console. It is asset-count-blind: it only sees the query text.
  • chooseTask() — performs the dispatch. It is asset-count-aware: it knows how many assets are attached and will not route a single-image query to a two-image task.

This asymmetry is real and explains a documented behaviour: for the query "What changed between the earlier and later image?" with a single asset attached, the console can read change while dispatch correctly falls back to change_vqa. This is intentional, not a bug, and is covered in RESEARCH_NOTES.md.

5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED)

The Codespace exposes exactly four routes; the gateway mirrors each under /api/*:

Codespace Gateway Cost Purpose
GET /v1/health GET /api/health cheap liveness, tunnel state, counters
GET /v1/capabilities GET /api/capabilities cheap per-task availability
POST /v1/analyze POST /api/infer costly run a query
POST /v1/assets POST /api/assets costly stage an asset handle

The gateway must not retry POST /api/infer on its own — a retry would consume inference a second time. The client decides on retry. Errors are translated into a stable envelope with a machine code (invalid_request, tunnel_offline, …) and a recoverable flag.

Verified: POST /api/infer {} returns 422 invalid_request with header x-satquery-transport: tunnel.

6. The frozen configuration (VERIFIED)

All tunables live in configs/base.yaml; no magic numbers in Python. The loader validates invariants and computes a hash. The frozen hash is:

78f1e3700da15aa1

The loader refuses to run a config that violates a recorded invariant — for example fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels (2318) and grounding_head.feature_dim == 4 * grounding.encoder_projected_dim (2048). A mismatch in the latter is a silent shape error otherwise (torch raises only later, after features are cached), so it is enforced at load time.

Editing configs/base.yaml moves the hash, which invalidates every artifact keyed to it. This is why the stale configs/deploy.yaml (which still describes an HF-Space/ZeroGPU target) is left undisturbed as frozen paperwork — it is never read at run time.

7. Component map

Layer Path (in the source repo) Notes
Frontend frontend/ static, staged by scripts/stage_pages.mjs
Gateway deploy/render/ main.py exposes app; render.yaml blueprint
Inference entry deploy/codespace/serve.py serves build_space_app() on $PORT
Inference app app/space_app.py build_space_app(); 4-endpoint contract
Composition root app/serving.py reuses the local serving path
Specialists specialists/ vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA)
Router router/ frozen MiniLM + trained adapter
Config configs/base.yaml authoritative registry; hashed
Artifacts artifacts/ weights, metrics, calibration, reports

Caveat. deploy/ inside the monorepo is stale and untracked; the deployed sources live in three separate private repositories. See DEPLOYMENT.md.

8. What is deliberately absent

  • No database, no auth, no queue. The gateway is stateless by design.
  • No GPU requirement. Device is selected via SATQUERY_DEVICE; all placement is .to(device), never .cuda(). The shipped deployment runs CPU-only.
  • No system-level end-to-end benchmark. There is no measured end-to-end accuracy number, and none is claimed. Per-specialist metrics exist and are reported in BENCHMARKS.md.