Image-to-Text
PyTorch
Safetensors
PEFT
English
remote-sensing
satellite-imagery
earth-observation
change-detection
visual-grounding
image-captioning
visual-question-answering
optical-sar-fusion
sar
multimodal
lora
Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download docs/ARCHITECTURE.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 10.2 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/ARCHITECTURE.md
- Command line
-
hf download hf://thundercode/SatQuery@00a146ce419c6c7c109650e26cf45353fffb594e/docs/ARCHITECTURE.md
-
curl -L -o ARCHITECTURE.md https://huggingface.co/thundercode/SatQuery/resolve/00a146ce419c6c7c109650e26cf45353fffb594e/docs/ARCHITECTURE.md
10.2 kB
| # Architecture | |
| **Status vocabulary used throughout these docs:** `IMPLEMENTED` (code exists and runs) Β· | |
| `VERIFIED` (checked against evidence) Β· `MEASURED` (a number was produced) Β· `ATTEMPTED` | |
| (tried, outcome recorded) Β· `NOT RUN` Β· `BLOCKED` Β· `DEFERRED` Β· `REJECTED` (tried and | |
| explicitly not accepted). Every architectural claim below carries one of these tags. | |
| --- | |
| ## 1. What the system is | |
| SatQuery AI answers natural-language questions about satellite imagery. It is a **router + | |
| specialists** system: a small intent router reads the question, dispatches it to one of six | |
| task specialists, and assembles the specialist's output into a single structured | |
| `ResultEnvelope` that the browser renders as an answer, an evidence list, and a confidence. | |
| It is not a single end-to-end vision-language model. The only learned components are: | |
| - a **50,822-parameter router adapter** over a frozen `all-MiniLM-L6-v2` sentence encoder; | |
| - four trained **task heads/adapters** (change, change_vqa, optical-SAR fusion, grounding); | |
| - one **LoRA adapter** over a frozen `SmolVLM-500M-Instruct` for caption/VQA. | |
| Everything else is frozen, publicly-pinned backbone weights resolved at run time. | |
| ## 2. Deployment topology (IMPLEMENTED, VERIFIED) | |
| The shipped system is a four-tier chain. **Status: VERIFIED live on 2026-09-25** β the health, | |
| capabilities and infer endpoints were probed against the running deployment. | |
| ``` | |
| Browser | |
| β HTTPS | |
| βΌ | |
| Cloudflare Pages β satquery.pages.dev (static frontend, 11 pages) | |
| β HTTPS/JSON | |
| βΌ | |
| Render β satquery-backend-m4yv.onrender.com (orchestrator / gateway) | |
| β outbound long-poll tunnel POST /tunnel/agent | |
| βΌ | |
| GitHub Codespace β FastAPI inference, port 8000 (CPU) | |
| β build_space_app() | |
| βΌ | |
| Specialists: SmolVLM Β· RemoteCLIP Β· MiniLM Β· CROMA Β· STANet | |
| ``` | |
| ```mermaid | |
| flowchart LR | |
| U[Browser] -->|HTTPS| CF["Cloudflare Pages<br/>static frontend"] | |
| CF -->|"HTTPS JSON<br/>/api/health Β· /api/capabilities Β· /api/infer"| R["Render<br/>orchestrator / gateway"] | |
| R -->|"outbound long-poll<br/>POST /tunnel/agent"| C["GitHub Codespace<br/>FastAPI inference :8000"] | |
| C --> S[(SmolVLM Β· RemoteCLIP<br/>MiniLM Β· CROMA Β· STANet)] | |
| C -->|ResultEnvelope| R | |
| R -->|"envelope + error translation"| CF | |
| ``` | |
| ### 2.1 Why a gateway and a tunnel | |
| - **The gateway (Render)** is deliberately thin and stateless: schema validation, size limits, | |
| per-IP rate limiting, a CORS allowlist (**never** `*`), request IDs, timeouts, secret custody, | |
| and translation of upstream failures into the documented error envelope. It is **not** a model | |
| host and has no database, auth, or queue. | |
| - **The tunnel** exists because the inference host (a Codespace) is not publicly addressable for | |
| this private repository β a forwarded port returns `302`. Instead the Codespace runs | |
| `tunnel_agent.py`, which **dials out** to `POST /tunnel/agent` and long-polls. Work is executed | |
| against `http://127.0.0.1:8000` locally. This inverts the usual direction: the inference host | |
| needs no inbound firewall hole. | |
| > **Status:** IMPLEMENTED and VERIFIED. `GET /api/health` reported `tunnel.agent_connected:true` | |
| > with a non-zero `completed` counter on 2026-09-25. | |
| ### 2.2 Cold start and the `transport_mode: auto` fallthrough (KNOWN ISSUE, OPEN) | |
| The Render free tier sleeps when idle and the Codespace may be stopped. Cold start is therefore | |
| tens of seconds and is **documented, not hidden**. A timeout on the tunnel path is surfaced to the | |
| client as a recoverable error. | |
| `SATQUERY_TRANSPORT=auto` means: try the tunnel; if it does not answer within | |
| `SATQUERY_TUNNEL_TIMEOUT_S` (150 s), **fall through to the forward path**. The forward path to a | |
| private repo returns `302` quickly, but the wake step still consumes `SATQUERY_WAKE_TIMEOUT_S` | |
| (120 s) first β so a worst-case failed request can take β 249 s. This is the root of the observed | |
| "transient tunnel gap" and is recorded as **OPEN** (not fixed) in | |
| [`LIMITATIONS.md`](LIMITATIONS.md). A deployed fix (`codespace_name` newline handling on the | |
| wake path) was authored and pushed separately; the fallthrough itself is by design. | |
| ## 3. Request lifecycle (IMPLEMENTED, VERIFIED) | |
| The controller is a nine-state machine, declared in `configs/base.yaml`: | |
| ``` | |
| RECEIVE β PARSE β VALIDATE β PLAN β PREPROCESS β EXECUTE β AGGREGATE β VERIFY β RESPOND | |
| ``` | |
| 1. **RECEIVE / PARSE** β the query and 1β2 image assets are decoded. | |
| 2. **VALIDATE** β modality inference from band count, dimension equality for pairs, byte cap | |
| (`4,194,304` bytes/file), GeoTIFF acceptance for the optical-SAR path. | |
| 3. **PLAN** β the router embeds the query with the frozen MiniLM encoder and the trained adapter | |
| predicts a task + modality. Dispatch is asset-count-aware: a single image cannot be a change | |
| task. | |
| 4. **PREPROCESS** β percentile normalisation for optical (2/98), dB clip for SAR (β30β¦+5 dB), | |
| tiling for large scenes. | |
| 5. **EXECUTE** β the chosen specialist runs. | |
| 6. **AGGREGATE** β the specialist's raw output becomes `Evidence` items | |
| (`Region` / `ChangeRegion`). | |
| 7. **VERIFY** β confidence is computed; temperature scaling is applied from | |
| `calibration_v001.json`. | |
| 8. **RESPOND** β a `ResultEnvelope` is assembled and returned. | |
| ### 3.1 The eight execution events | |
| The frontend renders a live trace from eight named events (source of truth: | |
| `frontend/assets/js/core.js`, `SQ.EVENT_NAMES`): | |
| | # | Event | Emitted when | | |
| |---|---|---| | |
| | 1 | `QUERY_RECEIVED` | the request is accepted | | |
| | 2 | `QUERY_UNDERSTOOD` | the router produces task + modality | | |
| | 3 | `ROUTE_SELECTED` | the specialist is chosen | | |
| | 4 | `SPECIALIST_STARTED` | specialist execution begins | | |
| | 5 | `SPECIALIST_COMPLETED` | specialist returns | | |
| | 6 | `EVIDENCE_GENERATED` | evidence items are built | | |
| | 7 | `CONFIDENCE_COMPUTED` | confidence (and calibration) is computed | | |
| | 8 | `RESULT_ASSEMBLED` | the envelope is finalised | | |
| **Live runs emit all eight; the offline preview path emits none of the specialist events and is | |
| labelled as a preview in the UI.** Measured trace fill on every live run: **94.4444 %**. | |
| ## 4. Two reading paths: `interpret()` vs `chooseTask()` (IMPLEMENTED, VERIFIED) | |
| There are two separate functions, and they see different information: | |
| - **`interpret()`** β produces the *human-readable* interpretation shown in the console. It is | |
| **asset-count-blind**: it only sees the query text. | |
| - **`chooseTask()`** β performs the *dispatch*. It is **asset-count-aware**: it knows how many | |
| assets are attached and will not route a single-image query to a two-image task. | |
| This asymmetry is real and explains a documented behaviour: for the query *"What changed between | |
| the earlier and later image?"* with a single asset attached, the console can *read* `change` while | |
| dispatch correctly falls back to `change_vqa`. This is intentional, not a bug, and is covered in | |
| [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md). | |
| ## 5. The four-endpoint inference contract (IMPLEMENTED, VERIFIED) | |
| The Codespace exposes exactly four routes; the gateway mirrors each under `/api/*`: | |
| | Codespace | Gateway | Cost | Purpose | | |
| |---|---|---|---| | |
| | `GET /v1/health` | `GET /api/health` | cheap | liveness, tunnel state, counters | | |
| | `GET /v1/capabilities` | `GET /api/capabilities` | cheap | per-task availability | | |
| | `POST /v1/analyze` | `POST /api/infer` | **costly** | run a query | | |
| | `POST /v1/assets` | `POST /api/assets` | **costly** | stage an asset handle | | |
| The gateway must **not** retry `POST /api/infer` on its own β a retry would consume inference a | |
| second time. The client decides on retry. Errors are translated into a stable envelope with a | |
| machine code (`invalid_request`, `tunnel_offline`, β¦) and a `recoverable` flag. | |
| **Verified:** `POST /api/infer {}` returns `422 invalid_request` with header | |
| `x-satquery-transport: tunnel`. | |
| ## 6. The frozen configuration (VERIFIED) | |
| All tunables live in `configs/base.yaml`; **no magic numbers in Python**. The loader validates | |
| invariants and computes a hash. The frozen hash is: | |
| ``` | |
| 78f1e3700da15aa1 | |
| ``` | |
| The loader **refuses to run** a config that violates a recorded invariant β for example | |
| `fusion.input_dim == 3 * croma.encoder_dim + croma.optical_channels + croma.sar_channels` | |
| (2318) and `grounding_head.feature_dim == 4 * grounding.encoder_projected_dim` (2048). A mismatch | |
| in the latter is a *silent* shape error otherwise (torch raises only later, after features are | |
| cached), so it is enforced at load time. | |
| Editing `configs/base.yaml` moves the hash, which invalidates every artifact keyed to it. This is | |
| why the stale `configs/deploy.yaml` (which still describes an HF-Space/ZeroGPU target) is left | |
| **undisturbed as frozen paperwork** β it is never read at run time. | |
| ## 7. Component map | |
| | Layer | Path (in the source repo) | Notes | | |
| |---|---|---| | |
| | Frontend | `frontend/` | static, staged by `scripts/stage_pages.mjs` | | |
| | Gateway | `deploy/render/` | `main.py` exposes `app`; `render.yaml` blueprint | | |
| | Inference entry | `deploy/codespace/serve.py` | serves `build_space_app()` on `$PORT` | | |
| | Inference app | `app/space_app.py` | `build_space_app()`; 4-endpoint contract | | |
| | Composition root | `app/serving.py` | reuses the local serving path | | |
| | Specialists | `specialists/` | vqa/caption (SmolVLM), grounding (RemoteCLIP), change (STANet), optical_sar (CROMA) | | |
| | Router | `router/` | frozen MiniLM + trained adapter | | |
| | Config | `configs/base.yaml` | authoritative registry; hashed | | |
| | Artifacts | `artifacts/` | weights, metrics, calibration, reports | | |
| > **Caveat.** `deploy/` inside the monorepo is **stale and untracked**; the deployed sources live in | |
| > three separate private repositories. See [`DEPLOYMENT.md`](DEPLOYMENT.md). | |
| ## 8. What is deliberately absent | |
| - **No database, no auth, no queue.** The gateway is stateless by design. | |
| - **No GPU requirement.** Device is selected via `SATQUERY_DEVICE`; all placement is `.to(device)`, | |
| never `.cuda()`. The shipped deployment runs CPU-only. | |
| - **No system-level end-to-end benchmark.** There is no measured end-to-end accuracy number, and | |
| none is claimed. Per-specialist metrics exist and are reported in [`BENCHMARKS.md`](BENCHMARKS.md). | |