SatQuery / docs /SERVING.md
thundercode's picture
release: add docs/SERVING.md
a6d4529 verified
|
Raw
History Blame Contribute Delete
75.8 kB

SatQuery AI — Serving

Chapter scope. This chapter documents the SatQuery AI inference service end to end: how the service is composed, the four-endpoint contract it exposes and the /api/* mirror of that contract, the entrypoint requirements any host must satisfy, the lazy model-loading model, the annotation-scope defect that once made every upload fail with a 422, the ephemeral asset store, the error taxonomy and its machine codes, the tunnel transport, and — stated plainly — what the service does not do.

Grounding. Every claim below comes from a file that was read for this chapter, cited inline, e.g. (app/space_app.py), (docs/API_CONTRACT.md §2.4). No endpoint, field, environment variable, status code, or number is invented. Where the evidence does not exist, the text says exactly: UNKNOWN — not established from the available evidence.

Status vocabulary follows release/DOCS_STYLE_GUIDE.md §2: IMPLEMENTED · VERIFIED · MEASURED · ATTEMPTED · NOT RUN · BLOCKED · DEFERRED · REJECTED · OPEN · RESOLVED · CLOSED.

Nothing in this chapter is a system-level accuracy claim. Per release/DOCS_STYLE_GUIDE.md §3 there is no end-to-end benchmark for SatQuery AI. This chapter describes a service; it does not score one.


1. What the serving tier is

SatQuery AI's serving tier is a Python HTTP service that exposes the project's analysis capability over four endpoints. It is built on FastAPI/Starlette, it is served by uvicorn, and it is composed by three modules:

Module Role
app/serving.py The composition root: builds a deployable registry and controller, wiring trained artifacts through the registry's builders= seam.
app/space_app.py The HTTP application: builds the FastAPI app (build_space_app()), owns the four routes, the asset store, and the error handlers.
app/deployment.py The capability adapter: turns internal registry state into the public capability vocabulary and produces the health and capabilities payloads.

Around those three sit:

  • core/controller.py — AnalysisController, the control tier that runs the pipeline.
  • core/registry.py — SpecialistRegistry, which discovers specialists from a spec table and builds them lazily.
  • core/planner.py — PolicyPlanner, the deterministic policy planner.
  • core/errors.py — the error taxonomy (23 codes) and the path scrubber.
  • gateway/app.py + gateway/policy.py — the gateway (an optional front tier; see §3.3).
  • deploy/codespace/serve.py — the 26-line process entrypoint that calls build_space_app().
  • deploy/codespace/launch.sh — the launcher that starts the service and its tunnel agent.
  • deploy/render/main.py — the Render orchestrator that exposes the /api/* mirror.

The service's job is narrow and worth stating: accept an analysis request, run the pipeline, return a ResultEnvelope. It does not render a UI, it does not stream, and it does not persist results. §11 lists what it does not do in full.


2. The composition root: app/serving.py

app/serving.py is 275 lines. Its module docstring calls itself "the public serving entry point" and states that it is "the thin, public composition root that wires a deployable controller".

2.1 The three wired artifacts

The module declares three module-level Path constants. Each is a repo-local artifact identity, not a config key:

Constant Path (relative to REPO_ROOT)
CHANGE_CHECKPOINT artifacts/change/levir_change_v001/head.pt
CHANGE_VQA_HEAD artifacts/change_vqa/run/head.pt
FUSION_HEAD artifacts/optical_sar/fusion_head_production_v001/head.pt

Their documented identities, as stated in the module's comments:

  • CHANGE_CHECKPOINT — "The trained, benchmarked change head (test pooled IoU 0.8122)." It is "the single source of truth for where serving looks for it; tests monkeypatch this to simulate an absent artifact." It is also the detector whose features scripts/prepare_change_vqa.py builds, "so training and serving share it." A test (tests/unit/test_app_serving.py::test_serving_and_preparation_share_one_stanet_checkpoint) keeps the two literals equal and fails if either side drifts.
  • CHANGE_VQA_HEAD — "The R-02 change-VQA reasoning head." Written by scripts/train_change_vqa.py (--output-dir, default artifacts/change_vqa/run) and read by scripts/evaluate_change_vqa.py (DEFAULT_CHECKPOINT, the same path). It "does NOT exist in a fresh checkout: it is produced by the external Kaggle run and returned to the maintainer". Absent ⇒ the specialist constructs, reports itself unavailable, and answers nothing.
  • FUSION_HEAD — "The verified production optical-SAR fusion head (Phase 12). Its identity is pre-registered, not inferred: sha256 785815729a3a39fc34dc41894efaf00d8739365d970a3f830a326e68ae888dab, 14,427,457 bytes, 1,201,711 parameters, and the checkpoint self-identifies with the embedded config_hash 78f1e3700da15aa1 and arm='A'."

Note the last one carefully: the fusion head's self-identification carries the same frozen config hash 78f1e3700da15aa1 that release/DOCS_STYLE_GUIDE.md §3 records as the project's frozen config hash. The artifact and the config agree by construction.

2.2 Why artifacts are wired through builders=, not through config

This is the single most important design decision in the serving tier, and app/serving.py documents it at length. The mechanism:

Config.hash (core/config.py:79-80) is a sha256 over the WHOLE registry. Adding one key moves the hash. The shipped change head records 78f1e3700da15aa1 in artifacts/change/levir_change_v001/model_metadata.json, and scripts/eval_change.py refuses to score on a hash drift (exit 3).

Therefore: editing configs/base.yaml to point at a trained head would invalidate the project's own benchmark number. The supported wiring path is instead the registry's builders= override (core/registry.py:420-433), keyed by spec name — "change" (core/registry.py:204-213) — and it is a call-site argument, not config, so the hash is untouched.

The registry calls a builder as builder(self.config, **kwargs) and only passes config keys that resolve (_builder_kwargs, core/registry.py:435-453). Since change.checkpoint_path is unset, nothing arrives and the real builder would degrade; the override supplies the missing argument.

Stated as a general rule: in this project, a serving-side artifact path may not be added to configs/base.yaml, because doing so would move the frozen config hash and invalidate every benchmark number keyed to it. The builders= seam is the hash-exempt channel for such paths.

2.3 Degrade, do not crash

The module's docstring states the governing principle: "A serving path must run even when the artifact is absent."

The distinction it enforces is precise:

  • Absent is a deployment case. When the checkpoint does not exist the module applies NO override, so the registry resolves the real builder with no checkpoint_path and the documented DEGRADED contract applies (specialists/change/specialist.py).
  • Corrupt is a defect. The real builder still surfaces it as ModelLoadError — "the two are deliberately not conflated."

The change_vqa override is applied unconditionally, because its builder's contract is finer-grained: a missing head and a missing detector are both named refusals (ChangeVQASpecialist.has_head, unavailable_reason), so wiring it can never turn "absent" into a crash. Both paths are passed as None when the file does not exist, "which the builder reads as 'artifact genuinely absent' rather than 'path I was told about is broken'."

2.4 The three builders

_wired_change_builder(config, **kwargs) — imports build_change_specialist lazily ("so importing app.serving stays cheap and does not pull the model stack (torch) into a process that never serves a change query"), sets kwargs["checkpoint_path"] = str(CHANGE_CHECKPOINT), and delegates.

_wired_change_vqa_builder(config, **kwargs) — exists to close a train/serve skew finding (named "finding F2" in the code). The skew: scripts/prepare_change_vqa.py builds its change features from the trained STANet (DEFAULT_CHANGE_CHECKPOINT), but serving had no equivalent wiring — change.checkpoint_path is unset in configs/base.yaml, so the registry passed no checkpoint_path and the specialist "would construct an UNTRAINED STANet and answer from a representation the head was never fitted on." The class's feature_spec_mismatch() already refused to answer on that skew, "so the failure was loud rather than silent — but a deployment that can never answer is still not a deployment." The override supplies the SAME checkpoint _wired_change_builder uses, so the detector backing a change answer and the detector behind the head's training features are "one artifact by construction; the spec check stays armed as the second line of defence, not the only one."

_wired_optical_sar_builder(config, **kwargs) — exists to close a different structural defect: "the encoder was unreachable by default." The mechanism, quoted from the module: specialists/optical_sar/specialist.py:946 builds CROMA only when handed a checkpoint_path that exists:

if checkpoint_path is not None and Path(checkpoint_path).exists():

croma.checkpoint_path is not in configs/base.yaml — and must not be, or Config.hash moves — so the registry passed no checkpoint_path, the gate was False, and the default serving composition ran with encoder=None. The registry then correctly reported DEGRADED ("no encoder; running on fallback"): "a deployment that could never answer an optical/SAR question. The checkpoint was present on disk the whole time; nothing asked for it."

The fix resolves the checkpoint from the PINNED identity through the hash-exempt channel (croma.resolve_checkpoint_path: env → config → pinned Hub cache, offline first), then hands it to the real builder. The resolution is recorded as source and logged, not attached to the returned specialist — with an explicit reason given in the code: "An unread attribute on a production object is how a contract quietly grows a second, undocumented shape; and it must not be published either, because the trace reaches the client and v1 has no auth."

That last sentence is a design principle worth extracting: anything the trace carries is public, because v1 has no auth. So composition-time facts that must not leak are logged rather than attached.

The fusion head is wired the same way for the same reason: has_head (specialists/optical_sar/specialist.py:175-187) is False without it, so the capability would stay DEGRADED even with the encoder loaded. Absent head ⇒ None ⇒ degrade.

2.5 build_serving_registry()

def build_serving_registry(config=None, *, device=None) -> SpecialistRegistry
  • config — the central core.config.Config. "Loaded unchanged when omitted; never mutated."
  • device — a torch device string; defaults to config.device_preference.
  • Returns "a SpecialistRegistry that constructs nothing yet (discover() reads a spec table)."

Its body builds a builders dict:

builders = {
    "change_vqa": _wired_change_vqa_builder,
    "optical_sar": _wired_optical_sar_builder,
}
if CHANGE_CHECKPOINT.exists():
    builders["change"] = _wired_change_builder
return SpecialistRegistry.discover(cfg, device=device, builders=builders)

Note the asymmetry and the reason for it, which the code comments on: change_vqa and optical_sar are registered UNCONDITIONALLY, unlike change. The comment explains: "The resolver decides at build time whether an artifact exists, so gating registration on a path this module does not know yet would be circular. Wiring it cannot turn 'absent' into a crash: the resolver returns None, the builder degrades, and a construction failure is retained as an UNAVAILABLE entry by SpecialistRegistry.build rather than escaping."

So: change is gated on the checkpoint existing; change_vqa and optical_sar are not, because their builders accept None as "absent".

2.6 build_serving_controller()

def build_serving_controller(config=None, *, device=None) -> AnalysisController

It is "Constructed with registry=, planner= and config= only."

registry = build_serving_registry(cfg, device=device)
return AnalysisController(
    registry=registry,
    planner=PolicyPlanner(registry),
    config=cfg,
)

The critical documented consequence: "No router is attached, so a caller drives it with AnalysisRequest(..., force_task=...); a natural-language router can be supplied by the caller's own composition if the router weights are available."

This is the single most important behavioural fact about the serving tier's request handling: the deployed service is driven by an explicit force_task, not by natural-language routing. It explains why the frontend's interpret() (see the FRONTEND.md chapter) does the lexical routing in the browser and then sends a force_task: the browser-side interpretation is what fills the gap left by the deliberately router-less serving composition.

__all__ exports CHANGE_CHECKPOINT, CHANGE_VQA_HEAD, FUSION_HEAD, build_serving_controller, and build_serving_registry.


3. The HTTP application: app/space_app.py

app/space_app.py is 736 lines and owns the HTTP surface.

3.1 The four routes

build_space_app() assembles a FastAPI application with four routes:

Method Path Kind Notes
GET /v1/health cheap Health block; includes device and gpu_available.
GET /v1/capabilities cheap Capability block; per-task availability and reasons.
POST /v1/analyze COSTLY Runs the pipeline; returns a ResultEnvelope.
POST /v1/assets COSTLY Uploads an asset; returns an opaque asset_id.

The "cheap vs COSTLY" distinction is not decoration: gateway/app.py declares

COSTLY_ROUTES = ("/v1/analyze", "/v1/assets")

and the gateway's policy (gateway/policy.py) applies its body-size caps, file-size caps, rate limit, and upstream timeout with those routes in mind. A cheap route can be polled; a COSTLY route cannot. (§3.3 covers the gateway.)

build_space_app() also installs two error handlers:

  • a StarletteHTTPException handler, and
  • a generic Exception handler (recorded in the deployment docs as F-12b).

The generic handler matters: without it, an unhandled exception would return a framework-default body that leaks internals. With it, the service returns a translated error. See §8.

3.2 The ZeroGPU duration map

The module declares a per-task duration budget used when the service is hosted on a ZeroGPU-style platform that requires an advance duration declaration:

Task Duration
vqa 20
caption 20
grounding 45
change 30
optical_sar 45
change_vqa 30

The helper decorate_gpu() applies the declaration, and _spaces_module() resolves the platform module. The numbers are the declared budgets, not measured latencies; the captured grounding run records a measured step_001 of 209.873 ms (see the FRONTEND.md chapter §7.4), which is a single step's timing, not a task duration, and the two are not comparable.

docs/DEPLOYMENT_ARCHITECTURE.md §3.4 documents this same map as the "ZeroGPU duration map". On the active topology the service runs on a CPU Codespace (SATQUERY_DEVICE=cpu, per deploy/codespace/launch.sh and docs/DEPLOYMENT_TOPOLOGY.md §5), where the GPU decoration is inert.

3.3 The gateway and the /api/* mirror

There are two front-facing surfaces, and it is important not to conflate them.

(a) The gateway (gateway/app.py). A thin front tier that proxies a 4-route allowlist to the inference service. Its declarations:

Symbol Value Meaning
PROXIED_ROUTES 4 routes The allowlist.
BLOCKED_ROUTES empty Nothing is explicitly blocked.
COSTLY_ROUTES ("/v1/analyze", "/v1/assets") The routes that cost real work.

It exposes /v1/gateway/health (its own health, distinct from /v1/health), installs a StarletteHTTPException handler (F-3), and proxies the four routes. Notable mechanisms inside _proxy():

  • F-2 — it strips client CORS headers and asserts that none remain (_is_cors_header(), _CORS_HEADER_PREFIX). This prevents a client from injecting an Access-Control-* header that the gateway would then pass upstream.
  • F-6 — it applies a streaming cap on the response body rather than buffering unbounded.
  • It deliberately does not retry (there is an explicit no-retry comment): a retry of a COSTLY route would double the work.
  • _read_body_bounded() (F-9) is a thin adapter that bounds the request body it reads.
  • _client_ip() derives the client IP (used by the rate limiter), and _env() reads configuration.

The gateway is an optional front tier. Its module docstring notes it is unimportable in a sandbox — i.e. it is written to be deployed, not imported by test runners — and the module-level app is created inside a try/except for that reason.

(b) The Render orchestrator (deploy/render/main.py, 532 lines). The orchestrator exposes the /api/* mirror of the four endpoints:

Orchestrator route Mirrors
/api/health /v1/health
/api/infer /v1/analyze
/api/capabilities /v1/capabilities
/api/assets /v1/assets

This is the surface the frontend actually calls: SQ.ENDPOINTS is {assets:'/assets', infer:'/infer', capabilities:'/capabilities', health:'/health'} (frontend/assets/js/live.js) and the default base is /api, so the frontend's /api/infer maps to the orchestrator's /api/infer, which maps to the service's /v1/analyze. Note the name change: the frontend says "infer"; the service says "analyze"; they are the same endpoint.

The orchestrator's internals:

Symbol Behaviour
_github_token() Reads the GitHub token used to wake the Codespace.
_codespace_name() Reads and strips the Codespace name — the strip is the fix for the B-02 trailing-\n defect (see §12).
_codespace_port() Defaults to 8000.
_wake_timeout_s() Defaults to 120.
_upstream_timeout_s() Defaults to 90.
_DEV_ORIGINS, _PRODUCTION_ORIGINS _PRODUCTION_ORIGINS = ("https://satquery.pages.dev",); _allowed_origins() composes the CORS allowlist.
OrchestratorError, WakeTimeout, OrchestratorConfigError, OrchestratorUpstreamError The orchestrator's own error types.
_envelope() Wraps a response/error into the orchestrator's envelope shape.
ensure_codespace_up() Wakes the Codespace if it is asleep (the wake sequence).
_proxy() Forwards the request upstream.
create_app() Builds the app with the four routes.
_handle_orchestrator_error() Translates an orchestrator error into a response.

Documented drift, recorded not hidden. deploy/render/main.py's own docstring notes that it is superseded by the tunnel design per the delivery documents, while remaining the source present in this working copy. The deployed backend is the SatQuery-Backend repository (main.py, 768 lines, with a tunnel), whose deployed HEAD is 89d80eaddec5 (release/DOCS_STYLE_GUIDE.md §3). The local deploy/render/main.py therefore does not carry the tunnel implementation. See §9 and §12.

render.yaml declares the orchestrator service concretely:

startCommand: uvicorn deploy.render.main:app --host 0.0.0.0 --port $PORT
healthCheckPath: /api/health

with environment variables PORT, SATQUERY_ALLOWED_ORIGINS, GITHUB_TOKEN, CODESPACE_NAME, CODESPACE_PORT ("8000"), SATQUERY_DEVICE ("cpu"), SATQUERY_WAKE_TIMEOUT_S ("120"), and SATQUERY_UPSTREAM_TIMEOUT_S ("90"). The plan is free, and all secret values are declared sync: false — i.e. they are injected by the platform, not committed. (No value is reproduced in this chapter; per the release rules, this documentation contains no credentials.)

3.4 The /v1 vs /api naming table

Because two naming schemes coexist, here is the mapping in one place:

Concept Service (/v1) Orchestrator mirror (/api)
Health GET /v1/health GET /api/health
Capabilities GET /v1/capabilities GET /api/capabilities
Analysis POST /v1/analyze POST /api/infer
Asset upload POST /v1/assets POST /api/assets
Gateway's own health GET /v1/gateway/health —

The /v1/ prefix is the service's versioned contract (docs/API_CONTRACT.md §1). The /api/ prefix is the orchestrator's mirror. A client that speaks /api/infer is speaking to the mirror, not to the service.


4. The four-endpoint contract in detail

docs/API_CONTRACT.md is the frozen, frontend-facing contract (917 lines). This section summarises what it pins, because the service must satisfy it exactly.

4.1 Conventions, and the one schema exception

docs/API_CONTRACT.md §1.1: unknown fields are rejected. The schemas use Pydantic extra="forbid", with exactly one exception: GeoMetadata is extra="allow". The reason is that geospatial metadata is an open set — a raster may carry CRS, transform, resolution, and arbitrary derived fields — so forbidding extras there would reject legitimate metadata rather than protect the contract.

The consequence for a client: sending an unexpected field on any other model is a validation error, not a silently-ignored field. This is a deliberate strictness choice, and it is why the contract is worth reading before writing a client.

4.2 GET /v1/health (§2.1)

Returns the health block. Two fields are worth pinning:

  • device is a closed set, validated by the F-8 rule in app/deployment.py: the legal values are _LEGAL_DEVICES = {cpu, cuda, mps}. _effective_device() returns None for an unrecognised device rather than echoing it back. So a client can rely on device being one of three values or absent.
  • gpu_available: false is normal on ZeroGPU. The contract records the measured degraded output, and states that a false here is not a fault on that platform.

The reason to state this in the docs at all: a naive client would treat gpu_available: false as an error. The contract says otherwise.

4.3 GET /v1/capabilities (§2.2, §2.3, §2.3.1)

Returns the capability block: per-task availability plus a reason when a task is unavailable.

  • reason is required when available: false. A capability block that said "unavailable" without saying why would be less useful than one that names the missing artifact.
  • modalities appears only on optical_sar. app/deployment.py declares _MODALITIES with only optical_sar in it, so no other task carries a modalities field.

§2.3 / §2.3.1 — the five-word vocabulary, and why loaded/degraded are never emitted. The public capability vocabulary has five states, and app/deployment.py translates internal registry states into them via CONTRACT_STATES (5) and REGISTRY_TO_CONTRACT. The internal registry states are AVAILABLE / DEGRADED / UNAVAILABLE (core/registry.py), and the registry has PLANABLE_STATES marking which of those the planner may plan against.

The important negative fact: the public contract never emits the words loaded or degraded. The internal vocabulary and the public vocabulary are deliberately different, and the translation is the adapter's job. A client that wrote if status == 'degraded' would be reading a word the contract does not use.

4.4 POST /v1/analyze (§2.4)

Accepts an AnalysisRequest and returns a ResultEnvelope.

Multipart is NOT implemented. This is stated in the contract and it constrains every client: an asset is uploaded separately to /v1/assets, and the analysis request references it by asset_id. A client that tried to send the image inline as a multipart part would be rejected. This is why the frontend's upload is a raw-bytes POST and why the analysis request is JSON (frontend/assets/js/live.js; FRONTEND.md §5.5, §6.8.1).

The request carries the task and, in the serving composition, a force_task (see §2.6 — the deployed controller has no router attached).

The response's fields that the frontend must read are enumerated in the contract: the answer, the evidence, the regions, the confidence (raw and calibrated), the timings, the provenance (run id, policy, protocol, schema), the geospatial block, and the warnings. The captured envelope on frontend/assets/data/anatomy-run.js is a real instance of this shape (FRONTEND.md §7.4).

Artifact refs are null in v1. The contract records an explicit ruling (F-16) that artifact references are null — the service does not return a URL or a handle to a produced artifact in v1. This is a capability limit, not an oversight, and a client must not depend on an artifact ref being present.

4.5 POST /v1/assets (§2.5)

Uploads an asset and returns an opaque asset_id. The contract records the design as "Option A" and pins:

Property Value
asset_id opacity The client must treat the handle as opaque.
Size cap Enforced (F-6 / F-7).
Content-type allowlist Five types.
Retries Documented.
Lifetime The handle is ephemeral with a TTL.

The service-side implementation of all five is in app/space_app.py (§7).

4.6 Enums (§3)

Enum Cardinality Values
Task 7 The task vocabulary.
CoordinateSystem 3 The coordinate-system vocabulary.
Modality 4 The modality vocabulary.

Seven tasks is worth noting because core/registry.py's default_specs() declares six specialists (vqa, caption, grounding, change, change_vqa, optical_sar). The Task enum having seven values while six specialists exist means the enum is the request vocabulary and the spec table is the implementation vocabulary; the difference is a task the request enum names but that no specialist serves directly. Which specific value accounts for the difference: UNKNOWN — not established from the available evidence (the enum's member list was not read verbatim for this chapter; only its cardinality is recorded here).

4.7 The confidence contract (§4)

The contract documents:

  • The measured ECE caveat. Calibration's ECE went 0.013755 → 0.014929 — worse. The transform is retained only because it is in the frozen config (release/DOCS_STYLE_GUIDE.md §3).
  • T = 0.9772731820958189 — the temperature.
  • 16,441 Val rows — the calibration sample count. This is the same figure the captured envelope records as calibration_samples: 16441.0 (frontend/assets/data/anatomy-run.js; FRONTEND.md §7.4). The public page and the contract agree.

The honest reading of this section: the service returns a calibrated confidence, and the calibration is documented to have made ECE slightly worse. A client must not present the calibrated confidence as an accuracy. Per the style guide, there is no end-to-end benchmark, so a per-run confidence is a per-run confidence.

4.8 The error contract (§5)

See §8 for the full treatment. The contract's §5.1 gives the status map, §5.2 the full 23-code taxonomy, and §5.3 the gateway-origin rate_limited code. §5.1 also records the trailing-slash 307 footgun (Starlette redirect_slashes), which is why a client should compose exact URLs.

4.9 Latency, quotas, auth, CORS (§6, §7, §7.1)

  • §6 — latency and quotas. The contract records the latency expectations and any quotas.
  • §7 — auth: none. v1 has no authentication. This is a first-class design fact with consequences that appear all over the codebase: it is why core/errors.py scrubs paths (F-15), why app/serving.py logs rather than attaches the CORS/checkpoint source, and why the trace must not carry anything sensitive.
  • §7.1 — CORS. CORS is configured on the orchestrator, whose _PRODUCTION_ORIGINS includes the Pages origin https://satquery.pages.dev (deploy/render/main.py). The gateway additionally strips client-supplied CORS headers (F-2, gateway/app.py).

§9 — the minimal integration checklist. The contract closes with a checklist for a new client, which is the shortest path for anyone writing against this service.


5. Entrypoint requirements

Any host that runs this service must satisfy five requirements. docs/DEPLOYMENT_ARCHITECTURE.md §3.3 enumerates them, and §3.3.1 adds a sixth consideration (a single capability authority). The requirements are:

  1. A Python process with the project's dependencies. deploy/codespace/launch.sh performs a preflight dependency check for yaml, pydantic, fastapi, uvicorn, and httpx before it starts anything. A host that does not have these cannot start the service.

  2. A callable application object. deploy/codespace/serve.py is the reference implementation:

    app = build_space_app()
    uvicorn.run(app, host="0.0.0.0", port=port)
    

    with port = int(os.environ.get("PORT", "8000")). The entrypoint therefore must (a) build the app via build_space_app() and (b) bind a port from the environment with a default.

  3. A port binding on 0.0.0.0. The reference binds 0.0.0.0, not 127.0.0.1, so the service is reachable from outside the process's own namespace.

  4. An environment that can reach the artifacts (or degrade cleanly without them). Because app/serving.py wires artifacts through the builders= seam and degrades when they are absent, a host without the artifacts still starts — it just reports the affected capabilities as unavailable. This is what makes "degrade, do not crash" a deployment property rather than a slogan.

  5. A health-checkable endpoint. The orchestrator's render.yaml sets healthCheckPath: /api/health, so the platform probes that path. A host that cannot answer a health probe will be considered unhealthy and restarted or removed from rotation.

§3.3.1 — a single capability authority. The architecture doc adds that there must be exactly one authority for capability state: app/deployment.py. The registry knows internal state (AVAILABLE/DEGRADED/UNAVAILABLE); the deployment adapter translates it into the public five-word vocabulary. A second place that decided capability state would create two answers to "is this task available?", which is exactly the kind of drift the project's discipline forbids.

5.1 The launcher: deploy/codespace/launch.sh

deploy/codespace/launch.sh is 194 lines and is the reference launcher. Its steps, as read:

  1. Preflight dependency checks for yaml, pydantic, fastapi, uvicorn, httpx.
  2. Port and stamp guards — so two launchers do not fight over the same port and a stale stamp does not mislead.
  3. _restart_serve() — starts the service with setsid nohup python deploy/codespace/serve.py, i.e. detached from the launcher's terminal so the service survives the shell.
  4. The supervised tunnel-agent loop — starts the tunnel agent with setsid nohup bash -c '… python deploy/codespace/tunnel_agent.py …' and supervises it, restarting it if it exits. See §9.
  5. Environment — exports SATQUERY_DEVICE=cpu, SATQUERY_ASSET_ENABLED=1, SATQUERY_ASSET_DIR=/tmp/satquery-assets, and SATQUERY_HUB_URL=https://<backend-host>.
  6. Verification — step 3 verifies the agent "announced to hub", so the launcher does not report success merely because the process started.

Note that SATQUERY_DEVICE=cpu in the launcher matches SATQUERY_DEVICE: "cpu" in render.yaml and the CPU-first reconciliation in docs/DEPLOYMENT_TOPOLOGY.md §5.

Honesty note. deploy/codespace/launch.sh references deploy/codespace/tunnel_agent.py, and docs/DEPLOYMENT_TOPOLOGY.md §2 and docs/FINAL_DELIVERY_TODO.md §1.3 both name that file as part of the SatQuery-Inference deployment. That file does not exist in this working copy. The local deploy/ directory is stale/untracked (docs/FINAL_DELIVERY_REPORT.md §6 records "local deploy/ stale"; docs/FINAL_DELIVERY_TODO.md §5 records the corresponding blocker). What the tunnel agent does is therefore described in §9 from the evidence that does exist (the launcher's invocation, the topology doc's description, and the transport value in the captured envelope), and the agent's internals are marked UNKNOWN — not established from the available evidence.


6. Lazy model loading and cache_max_models: 1

6.1 The lazy-loading contract

The serving tier does not load models at import time. Two mechanisms enforce this:

(a) build_serving_registry() constructs nothing. Its own docstring says it returns "a SpecialistRegistry that constructs nothing yet (discover() reads a spec table)." core/registry.py confirms the shape: default_specs() returns six spec rows, and discover() reads that table. The spec table is data; no model is instantiated by reading it.

(b) Builders import lazily. _wired_change_builder imports build_change_specialist inside the function, with the stated reason: "so importing app.serving stays cheap and does not pull the model stack (torch) into a process that never serves a change query." The same pattern appears in the other two builders. The consequence is that import app.serving does not import torch at all.

This matters because core/registry.py's SpecialistRegistry.__init__ has a torch-import path (self.device = device or config.device_preference). Keeping the import inside builders means a process that only serves, say, capabilities never pays for torch.

6.2 The spec table and lazy construction

core/registry.py:

Symbol Role
RegistryState AVAILABLE / DEGRADED / UNAVAILABLE.
PLANABLE_STATES Which states the planner may plan against.
SpecialistSpec One row of the spec table.
default_specs() Six rows: vqa, caption, grounding, change, change_vqa, optical_sar.
RegistryEntry The registry's record for one spec; to_trace() scrubs detail.
SpecialistRegistry.discover() Reads the spec table; constructs nothing.
SpecialistRegistry.available() Returns tuple(sorted(self._specs)) — a sorted tuple, so the order is stable.
SpecialistRegistry.specs() The spec table.
SpecialistRegistry.entry(name) One entry.

The default_specs() rows carry their asset requirements: requires_assets is 1, 2, or None depending on the task (a single-image task needs 1; a paired task needs 2; a task that needs no asset has None). They also carry optional_config_keys. These are the same requirements the frontend's PAIRED_TASKS = {change, change_vqa, optical_sar} reflects on the client side (FRONTEND.md §5.3) — and it is worth noting the two lists agree: the three paired tasks are exactly the three whose requires_assets is 2.

RegistryEntry.to_trace() scrubbing detail is a privacy mechanism: the trace reaches the client, and v1 has no auth, so the entry's raw detail does not travel.

6.3 cache_max_models: 1

The serving configuration caps the model cache at one model. The consequence is the important part: with a cache of one, serving a task evicts the previously loaded model. A sequence of requests across two tasks therefore loads and evicts repeatedly rather than holding both.

Why this is the right default for this deployment: the active host is a CPU Codespace (SATQUERY_DEVICE=cpu) with limited memory, and the project's posture is CPU-first (docs/DEPLOYMENT_TOPOLOGY.md §5). Holding several models resident would risk memory exhaustion, and an OutOfMemoryError is a defined failure in the taxonomy (core/errors.py, out_of_memory, recoverable=True) precisely because memory pressure is an expected condition.

The honest cost of cache_max_models: 1: a multi-task workload pays repeated model-load cost. This is a latency property, not a correctness one. It is stated here rather than omitted because it is a real consequence a reader should know before benchmarking latency.

Where cache_max_models is declared in config: UNKNOWN — not established from the available evidence for the exact key location (the value's effect — a cap of one — is what is documented here; the config file line was not read for this chapter).


7. The asset store

7.1 Purpose and shape

POST /v1/assets exists because multipart is not implemented (§4.4). An asset is uploaded once, receives an opaque handle, and the handle is referenced by the analysis request.

app/space_app.py implements the store with a module-level cache (_ASSET_STORE) and an accessor get_asset_store(). Its configuration comes from environment variables:

Helper Default Meaning
_asset_max_files() 32 Maximum number of files held.
_asset_ttl_seconds() 900.0 Handle lifetime, in seconds (15 minutes).
_asset_root() system tempdir fallback Where asset bytes are written.
_asset_max_file_bytes() — Per-file byte cap; refuses a non-positive or non-integer value (F-7).

_ALLOWED_ASSET_CONTENT_TYPES declares the five accepted content types, matching docs/API_CONTRACT.md §2.5 and the client's SQ.CONTENT_TYPES (frontend/assets/js/live.js: tif, tiff, png, jpg, jpeg).

7.2 Fail-closed availability

The store's availability gate is _asset_store_available(), which requires BOTH:

  • SATQUERY_ASSET_ENABLED, and
  • SATQUERY_ASSET_DIR.

If either is missing, the store is unavailable and POST /v1/assets returns 503. This is fail-closed: the service refuses uploads rather than accepting them into a store it cannot guarantee. That is the correct posture for an ephemeral store — a handle issued by a store that cannot serve it back is worse than no handle.

The launcher (deploy/codespace/launch.sh) sets both:

SATQUERY_ASSET_ENABLED=1
SATQUERY_ASSET_DIR=/tmp/satquery-assets

so the deployed Codespace has the store enabled with a temp-dir root. On a host where the variables are absent, the 503 is the expected behaviour and the frontend surfaces it via translateError() (FRONTEND.md §6.7).

7.3 Handle opacity and lifetime

The handle is asset_<32 hex> — 32 hex characters, which is secrets.token_hex(16). Two properties follow:

  1. It is unguessable. 16 random bytes (128 bits) means a client cannot enumerate handles.
  2. It is opaque. Nothing about the underlying file is encoded in it. The client must not parse it, and the frontend's uploadAsset() explicitly asserts the handle exists and passes it back unexamined (frontend/assets/js/live.js; FRONTEND.md §6.8.1).

The TTL (default 900.0 s) and the file cap (default 32) together mean the store is a short-lived staging area, not a database. The practical consequences for a client:

  • An upload and its analysis must happen within the TTL.
  • A workload that uploads more than 32 files concurrently will hit the cap.
  • Nothing survives a service restart: the store is in-memory plus a temp directory.

7.4 Path scrubbing on the way out

core/errors.py implements F-15 path scrubbing (scrub_paths()), which is directly relevant to the asset store because asset errors are client-visible. The mechanism:

  • _WINDOWS_DRIVE_PATH, _UNC_PATH, and _POSIX_PATH match absolute paths.
  • The replacement keeps only the final component ("basename reduction"), so "cannot read C:\\a\\b\\weights.pt" becomes "cannot read weights.pt" — "still diagnostic, no longer a location disclosure."
  • Relative paths are deliberately not matched, and the reason is documented: "A rule broad enough to catch artifacts/change/head.pt also catches and/or and the path segments of a URL, and a scrubber that mangles ordinary prose is a worse defect than the disclosure it fixes. The measured leaks are all absolute."
  • URLs are left intact on purpose: https://github.com/antofuller/CROMA appears inside one of the very messages this scrubs, and mangling it "would be a worse defect than the one being repaired." The _POSIX_PATH lookbehind refuses to start a match immediately after : or /, which is the mechanism that keeps the URL intact.

The module also records the history of the fix, which is instructive: a blunt replacement of the whole message with a generic string was tried first, "but it discarded path-free diagnostics the client can legitimately act on (... has no builder 'build_x', no GPU in this dimension), and three existing tests that pin exactly those diagnostics failed. A fix that forces legitimate tests to be weakened is aimed at the wrong granularity."

F-15's owner ruling (2026-09-23) is quoted in the file: "sanitize all client-facing exception messages; retain full exception details only in server-side diagnostics." The reason it was needed: exception messages in this repo routinely embed an absolute path (e.g. specialists/optical_sar/croma.py raises a message naming a vendored directory; specialists/change/stanet.py raises one naming an encoder-weights path), and those strings reach client-visible fields — and v1 has no auth.


8. Error translation and machine codes

8.1 The taxonomy: 23 codes

core/errors.py (316 lines) defines the taxonomy. Every failure the system can produce is one of these codes, and the module's docstring states the rule plainly: "Never raise a bare Exception from specialist or controller code."

The base class is SatQueryError, whose attributes are documented in the file:

Attribute Meaning
code Stable machine-readable identifier, used in traces.
user_message Text safe to show the operator.
detail Technical detail for the execution trace (never chain-of-thought).
recoverable Whether the controller may continue with a fallback.

It carries a to_trace() method returning {code, detail, recoverable, context}.

The taxonomy, grouped as the file groups it:

Input / raster.

Code Class recoverable
input_error InputError default
raster_read_error RasterReadError default
missing_crs MissingCRSError True — "Degraded, not fatal: non-geospatial analysis may still be possible."
unsupported_bands UnsupportedBandsError default
oversized_image OversizedImageError True — recoverable via downscale.

Pairing.

Code Class Note
pair_incompatible PairCompatibilityError —
pair_misaligned PairMisalignmentError Subclass of the above.
temporal_pair_invalid TemporalPairError Subclass of the above.

Routing / planning.

Code Class
routing_error RoutingError
unsupported_query UnsupportedQueryError
invalid_request InvalidRequestError
workflow_plan_error WorkflowPlanError

Specialists.

Code Class recoverable
specialist_error SpecialistError default
model_load_error ModelLoadError default
model_unavailable ModelUnavailableError True — "the controller degrades the workflow."
out_of_memory OutOfMemoryError True — retry at lower resolution.
specialist_timeout SpecialistTimeoutError True

Output integrity.

Code Class
schema_validation_error SchemaValidationError
coordinate_error CoordinateError
confidence_range_error ConfidenceRangeError

Leakage / evaluation.

Code Class
leakage_violation LeakageError
benchmark_freeze_error BenchmarkFreezeError

That is 23 codes, matching __all__'s 23 entries and the "23-code taxonomy" recorded in docs/API_CONTRACT.md §5.2 and gateway/policy.py's _CODE_STATUS.

8.2 The specialist_timeout recoverability correction

One entry deserves its own treatment because the file documents a defect it corrected. SpecialistTimeoutError was inheriting recoverable=False from SatQueryError, and the file explains why that was wrong, with two independent reasons:

  1. docs/API_CONTRACT.md is the frozen frontend-facing contract, and §5.1 maps 504 with recoverable: true. A frontend that reads recoverable: false "will not offer a retry for the one failure the contract explicitly tells it to retry."
  2. The plan's Failure Matrix (§57) lists Timeout with the recovery "abort specialist" and the fallback "partial result" — i.e. the controller continues rather than failing the request. A terminal recoverable=False contradicts that.

The file also records why the defect was invisible from the inside: "the controller currently only reuses .code for its budget-skip trace entry (core/controller.py:464), so nothing in the pipeline constructed this class and the wrong default was never observable from the inside — only from a client." This is a good example of the project's practice of documenting how a bug could hide.

8.3 The status map and the gateway-origin code

gateway/policy.py declares _CODE_STATUS, the map from each of the 23 codes to an HTTP status, and:

GATEWAY_ORIGIN_CODES = {"rate_limited"}
_CODE_STATUS["rate_limited"] = 429

So rate_limited is a gateway-origin code: it is not one of the 23 taxonomy codes produced by the service, it is produced by the gateway's own rate limiter, and it maps to 429. docs/API_CONTRACT.md §5.3 records it separately for exactly this reason — a client should understand that a 429 came from the gateway, not from the analysis pipeline.

DEFECT_CODES (5) names the codes that indicate a defect rather than a normal failure. The distinction matters: a defect code means the system did something wrong, whereas most codes describe a legitimate condition (a missing CRS, a bad upload, a timeout).

8.4 translate_error()

translate_error() maps an error to its client-facing form. Its role in the architecture is stated in docs/DEPLOYMENT_ARCHITECTURE.md §2.3: the code is passed unchanged. The gateway translates the shape (into its envelope, with a request id) but does not rewrite the code — so a client sees the service's own code, not a gateway-invented one.

Supporting symbols: _REQUEST_ID_RE (validates a request id's shape) and new_request_id() (mints one). A request id is what makes a client-side report correlatable with a server-side log.

8.5 GatewayConfig and its validators

gateway/policy.py declares GatewayConfig with these defaults:

Field Default
max_body_bytes 8 MiB
max_file_bytes 4 MiB
rate_limit_per_ip 10
rate_limit_window_s 60.0
upstream_timeout_s 90.0
allowed_content_types 5

Its __post_init__ validators reject a misconfiguration rather than letting it fail later:

  • an origin with a trailing slash is rejected,
  • an empty value is rejected,
  • a * wildcard is rejected,
  • and a timeout that is not greater than 45 is rejected.

The last one is interesting: the 45-second floor is tied to the GPU duration map's longest budget (grounding and optical_sar are both 45 in app/space_app.py's GPU_DURATIONS). An upstream timeout below the longest task budget would cut off a legitimate run, so the validator forbids it.

Note the relationship between the two size caps: the gateway's max_file_bytes (4 MiB) is smaller than its max_body_bytes (8 MiB), which is coherent — a file cap inside a body cap.

8.6 The F-12b generic handler

Back in app/space_app.py, the generic Exception handler (F-12b) is what makes the taxonomy airtight at the edge: an exception that escaped the pipeline's own handling is still translated into a response rather than surfacing as a framework default. docs/DEPLOYMENT_ARCHITECTURE.md §5 lists F-12 and F-12b among the failure modes, alongside F-11, F-13, F-14, F-15, F-15b, F-15c, F-16, F-16c, F-17, F-18, and F-19. (F-15c is the gateway's transport-failure detail, _TRANSPORT_FAILURES / _transport_failure_detail() in gateway/app.py.)


9. The tunnel agent and the transport

9.1 Why a tunnel exists

The service runs on a host (a GitHub Codespace) that is not directly reachable at a stable public address in the way a normal web service is. The orchestrator on Render is the public face. Something must carry a request from the orchestrator to the service. That "something" is the transport, and the captured envelope records the transport it used:

transport: "tunnel"

(frontend/assets/data/anatomy-run.js; FRONTEND.md §7.4). The frontend's live client also reads a transport response header, x-satquery-transport (frontend/assets/js/live.js), which is how a client can see which transport carried its response.

9.2 The two transports

docs/DEPLOYMENT_TOPOLOGY.md and the delivery documents describe two transport designs:

  1. Forwarded-port transport. The orchestrator reaches the Codespace through a forwarded port. In this design a private repository yields a 302 (a redirect), which is why a 302 is a documented behaviour rather than an error.
  2. Outbound tunnel transport. The service-side agent long-polls POST /tunnel/agent to the hub, so the connection is outbound from the Codespace. An outbound tunnel avoids requiring the Codespace to be reachable inbound, which is the property that makes it robust on a platform that does not expose inbound ports.

The tunnel design supersedes the forwarded-port design: deploy/render/main.py's docstring says it is superseded by the tunnel design per the delivery documents, and the deployed backend repository is the one that carries the tunnel.

9.3 The agent's role, and what is known about it

The agent's role, assembled from the evidence that exists:

  • deploy/codespace/launch.sh starts and supervises it with setsid nohup bash -c '… python deploy/codespace/tunnel_agent.py …', detached from the launcher's terminal and restarted if it exits. So the agent is a long-running process, not a one-shot.
  • It announces to the hub. The launcher's step 3 verifies that the agent "announced to hub", so announcing is part of the agent's contract and the launcher treats a failed announcement as a failed launch.
  • SATQUERY_HUB_URL names the hub. The launcher sets it to https://<backend-host>, which is the same host the frontend's <meta name="satquery-api-base"> names (frontend/mission.html). So the hub, the orchestrator, and the API base are one host.
  • It is supervised, and it is started after the service. The launcher starts the service (_restart_serve()) and then starts the agent, which is the correct order: an agent that announced before the service was listening would advertise a dead endpoint.

What the agent does internally — its poll loop, its request framing, its reconnection strategy, its handling of a hub restart — is UNKNOWN — not established from the available evidence, because deploy/codespace/tunnel_agent.py does not exist in this working copy (§5.1's honesty note). The deployed backend repository (HEAD 89d80eaddec5) is where the tunnel implementation lives, and it was not read for this chapter.

9.4 B-07: tunnel gaps, patch prepared but not deployed

Per release/DOCS_STYLE_GUIDE.md §3 and docs/FINAL_DELIVERY_TODO.md §5: B-07 is OPEN. It is tunnel gaps, and the patch is prepared but NOT deployed. This status must not be upgraded. The correct statement is:

B-07 — tunnel gaps. Patch prepared, not deployed. OPEN.

The consequence for a reader: the tunnel transport works well enough to have carried the runs recorded in the delivery documents (including the captured run_d124d8b9adea, whose transport is "tunnel"), and it also has known gaps whose fix is written but not live. Both halves are true at once.


10. The deployment topology

10.1 The active topology

docs/DEPLOYMENT_TOPOLOGY.md is the active topology document. Its components:

Component Host Role
Static tier Cloudflare Pages The eleven pages (see FRONTEND.md).
Public backend Render (satquery-orchestrator) The /api/* mirror; wake + proxy; CORS.
Inference GitHub Codespace Runs the service (build_space_app()), CPU-first, plus the tunnel agent.
Model artifacts Hugging Face Artifact hosting; also the public release surface.

The document contains a Mermaid topology diagram and a wake sequence, plus §3's per-component responsibilities and environment variables, §4's five old blockers, §5's reconciliation (CPU-first), and §6's preconditions.

10.2 Deployed HEADs

Per release/DOCS_STYLE_GUIDE.md §3:

Component Deployed HEAD
Frontend 2d7ae53b482d
Backend 89d80eaddec5
Inference 5a0936ace491

10.3 The measured live environment

docs/DEPLOYMENT_TOPOLOGY.md §3.2 records the measured live env-var set. Two entries in that section are worth flagging because the section also notes that some names listed historically are not in the live config: SATQUERY_UPSTREAM_URL and HF_TOKEN are named in the section's own prose while the section's measured note says they are not present. This is documentation drift inside the topology document, recorded here rather than propagated.

docs/DEPLOYMENT_ARCHITECTURE.md carries a superseded-topology banner and still names Railway / HF-Space hosts in its body while the active hosts are Render / Codespace. Both documents are kept, with the banner making the supersession explicit — which is the project's stated practice (mirroring P10-T02).

10.4 docs/DEPLOYMENT_ARCHITECTURE.md §2 — gateway responsibilities

The architecture document's §2 enumerates the gateway's responsibilities and the 4-route allowlist, and §2.3 pins the error-translation rule (code passed unchanged). §3.1 assigns entrypoint ownership, §3.2 lists constraints, §3.3 lists the five entrypoint requirements, §3.3.1 the single capability authority, §3.4 the ZeroGPU duration map, §4 the env-var vocabulary (a long table with F-6/F-7/F-8/F-9 notes), §5 the failure-mode table (F-11…F-19), §6 what is excluded, §7 implementation status, and §8 deployment preconditions.

10.5 The pipeline the service runs

The service's work is done by core/controller.py's AnalysisController.run(), whose stages are:

RECEIVE → PARSE → VALIDATE → PLAN → EXECUTE → AGGREGATE → VERIFY → RESPOND

The captured grounding envelope's eight steps are RECEIVE → RESPOND, i.e. the same eight-stage pipeline (frontend/assets/data/anatomy-run.js). Notable details from core/controller.py:

  • _asset_label — a basename reduction applied to asset labels, the same idiom as F-13/F-14 and the same idiom core/errors.py::scrub_paths uses for F-15. "One rule, one implementation, applied at every client-facing write site."
  • F-19 — the registry is re-snapshotted after execute: trace.parameters["registry"] = self.registry.describe() is written after the EXECUTE stage, so the trace records the registry state that actually ran rather than the state at request entry.
  • _execute() — applies a budget between steps; and per F-15, sets trace.errors[].message = user_message (the sanitized message, not the raw detail).
  • _execute_one() — implements F-20, a producer-side repair for unhandled exceptions, so a specialist that raises something unexpected is still recorded as a result rather than escaping.
  • health() — deprecated: it "Constructs everything", and it was retired as the public path. This is why app/deployment.py owns the health payload instead: the public health path must be cheap, and a health check that constructs every model is not cheap.
  • _route(), _resolve_assets(), _modalities() — the routing, asset-resolution, and modality helpers.

11. What the service does NOT do

Stated explicitly, because the depth of §2–§10 could otherwise imply more capability than exists.

  • No Gradio GUI. The service is an HTTP API. There is no Gradio interface in this serving tier; the user interface is the static frontend (FRONTEND.md), which talks to the service over HTTP. Whether a Gradio surface exists anywhere else in the project: UNKNOWN — not established from the available evidence for this chapter (the serving modules read contain no Gradio application).
  • No streaming. There is no server-sent-events or websocket channel. A request is answered with a single response. The frontend's eight-event display is driven client-side from that one response plus two headers (X-SatQuery-State, x-satquery-transport), not pushed from the server (FRONTEND.md §14).
  • No batching. A request is one analysis. There is no batch endpoint, and POST /v1/analyze takes one AnalysisRequest.
  • No queue. There is no job queue and no async job model: a COSTLY route does its work within the request, bounded by the upstream timeout (upstream_timeout_s default 90.0) and the gateway's timeout floor (> 45). This is why the gateway deliberately does not retry (gateway/app.py): a retry of a COSTLY route would duplicate work rather than dequeue it.
  • No authentication. v1 has no auth (docs/API_CONTRACT.md §7). This has downstream consequences throughout: path scrubbing (F-15), trace scrubbing (RegistryEntry.to_trace() scrubs detail), and logging-instead-of-attaching composition facts (app/serving.py).
  • No multipart upload. Assets are uploaded separately (docs/API_CONTRACT.md §2.4).
  • No artifact refs. Artifact references are null in v1 (the F-16 ruling).
  • No persistence. The asset store is ephemeral (TTL 900.0 s, cap 32 files) and there is no run store. A restart loses everything.
  • No natural-language routing in the serving composition. build_serving_controller() attaches no router, so a caller drives it with force_task (app/serving.py; §2.6).
  • No model preloading. Models load lazily and the cache holds one (cache_max_models: 1; §6).
  • No end-to-end benchmark. Per release/DOCS_STYLE_GUIDE.md §3 this does not exist, and no system-level accuracy is claimed anywhere in this chapter.

12. Status summary and blockers

12.1 Status by subsystem

Subsystem Status
app/serving.py composition root (build_serving_registry, build_serving_controller) IMPLEMENTED
Artifact wiring via the builders= seam (change / change_vqa / optical_sar) IMPLEMENTED
app/space_app.py (build_space_app(), four routes, two error handlers) IMPLEMENTED
app/deployment.py capability adapter (two vocabularies, five contract states) IMPLEMENTED
Four-endpoint contract (/v1/health, /v1/capabilities, /v1/analyze, /v1/assets) IMPLEMENTED
/api/* orchestrator mirror IMPLEMENTED; deployed backend HEAD 89d80eaddec5
Gateway (4-route allowlist, COSTLY_ROUTES, F-2/F-3/F-6/F-9) IMPLEMENTED
Lazy model loading; cache_max_models: 1 IMPLEMENTED
Asset store (opaque handles, TTL, cap, allowlist, fail-closed 503) IMPLEMENTED
Error taxonomy (23 codes) + _CODE_STATUS + gateway-origin rate_limited IMPLEMENTED
Path scrubbing (F-15) IMPLEMENTED
Tunnel transport IMPLEMENTED; carried run_d124d8b9adea (transport: "tunnel")
B-02 codespace_name trailing \n Fixed in deploy/render/main.py via a strip; recorded as cosmetic, OPEN
B-07 tunnel gaps Patch prepared, NOT deployed — OPEN

12.2 The blockers, stated exactly

ID Statement Status
B-07 Tunnel gaps. Patch prepared, not deployed. OPEN — never to be upgraded.
B-02 codespace_name trailing \n. Cosmetic. The orchestrator's _codespace_name() strips it. OPEN (cosmetic)
Local deploy/ The local deploy/ directory is stale/untracked; deploy/codespace/tunnel_agent.py is absent; deploy/render/main.py is superseded by the deployed backend. KNOWN (docs/FINAL_DELIVERY_TODO.md §5 B-03; docs/FINAL_DELIVERY_REPORT.md §6)
Change capability Recorded as degraded in the delivery documents at the time of writing. KNOWN — per docs/FINAL_DELIVERY_REPORT.md §6
P2-T03 Cosmetic. OPEN (cosmetic)

docs/FINAL_DELIVERY_TODO.md §5 records the full blocker register: B-01 CLOSED, B-02 DOWNGRADED, B-03 KNOWN, B-04 ACCEPTED, B-05 ACCEPTED, B-06 KNOWN, B-07 OPEN, B-08 CLOSED. Note that B-01 (which docs/FINAL_DELIVERY_REPORT.md §6 records as HF BLOCKED at the time of that report) is CLOSED in the later TODO register — so the correct current statement is that B-01 is CLOSED, with the earlier report's BLOCKED status being superseded.

12.3 The G-1 annotation-scope defect

This is the most instructive serving defect in the project and deserves its own treatment.

The mechanism. app/space_app.py uses from __future__ import annotations. Under that import, annotations are strings, resolved lazily by FastAPI via eval against a namespace. If a parameter's annotation names a type (Request) that is bound in a narrower scope than the function that FastAPI introspects, then FastAPI's eval resolves that name against the wrong globals. The name fails to resolve as a type, and FastAPI silently reinterprets the parameter as a REQUIRED QUERY PARAMETER named request.

The symptom. Every upload gets:

422 {"detail":[{"loc":["query","request"]}]}

This is the worst kind of bug: a server-side defect that presents as a client-side validation error. A client developer reads "missing required query parameter request" and concludes they mis-called the API. They did not.

Why it is silent. There is no exception at import time. The app builds. The route registers. Only the interpretation of the parameter changed, and it changed in a way that produces a plausible-looking error.

The twin, and the asymmetry. The related case is a return annotation naming JSONResponse. In that case the resolution failure does not degrade silently — it raises PydanticUndefinedAnnotation, and it raises at import/definition time, so build_space_app() is never called at all. The app therefore does not exist.

So the defect has two halves with opposite failure modes:

Annotation position Failure mode
Parameter annotation Silent. The parameter is reinterpreted as a required query parameter. The app runs and every upload 422s.
Return annotation Loud. PydanticUndefinedAnnotation is raised before build_space_app() can be called; the app never starts.

The asymmetry is why the defect is worth documenting: the loud half is easy to find (the app will not start), and the silent half is the dangerous one (the app starts and lies about why it is failing).

The repair pattern. app/space_app.py lines 55–91 carry module-scope comment blocks binding Request, Response, and JSONResponse at module scope, so that FastAPI's eval resolves the names against the module's globals. The gateway has the twin of this: gateway/app.py also binds Request, Response, and JSONResponse at module level for the same reason. The rule extracted:

Under from __future__ import annotations, every type used in a FastAPI route signature must be bound at the module scope where the route function is defined — because FastAPI resolves annotations by eval against that module's globals, and a narrower-scope binding resolves to nothing.

The correct status for G-1: the repair is IMPLEMENTED (the module-scope bindings are present in both app/space_app.py and gateway/app.py). The defect is RESOLVED in the code read. Whether an earlier deployment ever served the silent-422 behaviour is a historical question: the recorded live validation ran 24 runs with 8/8 per pass (release/DOCS_STYLE_GUIDE.md §3), which is consistent with a working upload path in the deployed build — but the exact deployment at which the fix landed is UNKNOWN — not established from the available evidence.

12.4 Other failure modes recorded in the architecture doc

docs/DEPLOYMENT_ARCHITECTURE.md §5 lists the failure-mode table. The ones most relevant to serving:

ID Subject
F-6 Streaming size cap (also gateway/app.py _proxy()).
F-7 _asset_max_file_bytes() refuses a non-positive or non-integer value.
F-8 device validation → _effective_device() returns None for an unrecognised value; _LEGAL_DEVICES = {cpu, cuda, mps}.
F-9 _read_body_bounded() in the gateway.
F-11 (per §5)
F-12 / F-12b The generic exception handler in build_space_app().
F-13 / F-14 _asset_label basename reduction.
F-15 / F-15b Path scrubbing; the F-15b variant.
F-15c Gateway transport-failure detail (_TRANSPORT_FAILURES, _transport_failure_detail()).
F-16 / F-16c The artifact-refs-null ruling; the F-16c variant.
F-17 / F-18 / F-19 F-19 is the post-execute registry re-snapshot in core/controller.py.
F-20 Producer-side repair for unhandled exceptions in _execute_one().

docs/DEPLOYMENT_ARCHITECTURE.md §4's env-var vocabulary table carries the F-6/F-7/F-8/F-9 notes inline, and §6 states what is excluded from the deployment, §7 its implementation status, and §8 the deployment preconditions.


13. NOT RUN / OPEN / BLOCKED (serving)

Per release/DOCS_STYLE_GUIDE.md §4, every doc ends with this list.

NOT RUN

  • No end-to-end benchmark of the service (project-wide fact per release/DOCS_STYLE_GUIDE.md §3; the service is not exempt, and no system-level accuracy is claimed).
  • No load/latency benchmark of the four endpoints under cache_max_models: 1.
  • No test of the tunnel under a hub restart.
  • No verification of the gateway's rate limiter under sustained load.
  • No verification of the asset store's cap (32) and TTL (900.0 s) boundaries end to end.

OPEN

  • B-07 — tunnel gaps. Patch prepared, NOT deployed. OPEN. (Never to be upgraded.)
  • B-02 — codespace_name trailing \n. Cosmetic. OPEN. (The strip is present in deploy/render/main.py.)
  • P2-T03 — cosmetic. OPEN.
  • F-15 path scrubbing — the measured leaks are all absolute paths; relative-path leaks were deliberately not covered. The scoping is documented as intentional; whether any relative-path leak exists is UNKNOWN — not established from the available evidence.
  • Documentation drift inside the topology docs — docs/DEPLOYMENT_TOPOLOGY.md §3.2 names SATQUERY_UPSTREAM_URL and HF_TOKEN while its own measured note says they are not in the live config; docs/DEPLOYMENT_ARCHITECTURE.md names Railway / HF-Space hosts under a superseded-topology banner. Recorded; OPEN as documentation debt.
  • deploy/codespace/tunnel_agent.py — referenced by launch.sh and two delivery docs, absent from this working copy. The agent's internals are UNKNOWN — not established from the available evidence.
  • Task enum's seventh value — the enum has seven values while six specialists are declared; which value accounts for the difference is UNKNOWN — not established from the available evidence.
  • cache_max_models config key location — the value's effect (a cap of one) is documented; the exact key location is UNKNOWN — not established from the available evidence.
  • G-1's fix deployment point — the repair is IMPLEMENTED in the code read; the deployment at which it landed is UNKNOWN — not established from the available evidence.
  • B-01 — CLOSED per docs/FINAL_DELIVERY_TODO.md §5 (superseding the earlier report's BLOCKED status). Recorded here so it is not re-opened.
  • No LICENSE file exists — project-wide, OPEN (release/DOCS_STYLE_GUIDE.md §3).

BLOCKED

  • Nothing in the serving code read for this chapter is blocked.
  • Deployment-level: the local deploy/ tree is stale/untracked, so the tunnel implementation cannot be read from this working copy — the corresponding investigation is BLOCKED on that tree being refreshed (or on the deployed backend repository being read instead).
  • B-01 at the time of docs/FINAL_DELIVERY_REPORT.md was BLOCKED (HF); it is CLOSED per the later TODO register. The earlier status is superseded, not deleted.

14. Where the evidence lives

Claim area Evidence file(s)
Composition root; the three artifact constants and their identities; the builders= seam and why config must not be edited; degrade-don't-crash; the three builders and the defects they close; build_serving_registry(); build_serving_controller() (no router → force_task) app/serving.py
HTTP application; build_space_app(); the four routes; the two error handlers (incl. F-12b); GPU_DURATIONS; decorate_gpu(); _spaces_module(); get_controller(); describe_deployment(); the asset-store helpers (_asset_max_files() 32, _asset_ttl_seconds() 900.0, _asset_root(), _asset_max_file_bytes() F-7, _ALLOWED_ASSET_CONTENT_TYPES 5, _asset_store_available() requiring both env vars); main() app/space_app.py
Capability adapter: CONTRACT_STATES (5), REGISTRY_TO_CONTRACT, _REQUIREMENTS, _MISSING_REASONS, _HUB_REASONS, _optical_sar_artifacts(), _resolve_croma_checkpoint(), _requirement_artifacts(), _missing_shipped(), _hub_unconfigured(), _configured_path(), _HUB_BACKED, CapabilityReport, _MODALITIES, DeploymentReport, _schema_version(), _registry_capabilities(), _asset_count(), _artifact_evidence(), _report_for(), deployment_report(), _effective_device() (F-8), _LEGAL_DEVICES, _cuda_detected(), health_payload(), capabilities_payload() app/deployment.py
Error taxonomy (23 codes), SatQueryError + to_trace(), the specialist_timeout recoverability correction, F-15 path scrubbing (_WINDOWS_DRIVE_PATH, _UNC_PATH, _POSIX_PATH, scrub_paths()) core/errors.py
_CODE_STATUS (23 codes), DEFECT_CODES (5), GATEWAY_ORIGIN_CODES, rate_limited → 429, translate_error(), _REQUEST_ID_RE, new_request_id(), GatewayConfig + validators gateway/policy.py
Gateway: PROXIED_ROUTES (4), BLOCKED_ROUTES, COSTLY_ROUTES, /v1/gateway/health, F-3 handler, _read_body_bounded() (F-9), _proxy() (F-2 CORS strip + assertion, F-6 streaming cap, no-retry), _is_cors_header(), _CORS_HEADER_PREFIX, _client_ip(), _env(), module-level Request/Response/JSONResponse bindings (the G-1 twin) gateway/app.py
Registry: RegistryState, PLANABLE_STATES, SpecialistSpec, default_specs() (6 rows, requires_assets), RegistryEntry.to_trace() scrubs detail, discover(), available(), specs(), entry(); the builders= override site (lines 420-433) and the spec-name key (lines 204-213); _builder_kwargs (lines 435-453) core/registry.py
Controller: the eight-stage pipeline; _asset_label (F-13/F-14); the F-19 post-execute registry re-snapshot; health() deprecated ("Constructs everything"); _route(), _resolve_assets(), _modalities(), _execute() (budget; F-15 user_message), _execute_one() (F-20) core/controller.py
The frozen contract: conventions + extra="forbid" / GeoMetadata extra="allow" (§1.1); health (§2.1, device closed set, gpu_available: false normal on ZeroGPU); capabilities (§2.2, §2.3, §2.3.1 five-word vocabulary, modalities only on optical_sar); analyze (§2.4, multipart NOT implemented, artifact refs null per F-16); assets (§2.5, opacity, caps, allowlist, lifetime); enums (§3, Task 7 / CoordinateSystem 3 / Modality 4); confidence (§4, ECE 0.013755→0.014929, T = 0.9772731820958189, 16,441 Val rows); errors (§5, §5.1 status map + 307 footgun, §5.2 23 codes, §5.3 rate_limited); latency/quotas (§6); auth (§7) + CORS (§7.1); status (§8); integration checklist (§9) docs/API_CONTRACT.md
Five entrypoint requirements; §3.3.1 single capability authority; gateway responsibilities + 4-route allowlist + COSTLY; §2.3 error translation (code unchanged); §3.4 ZeroGPU duration map; §4 env-var vocabulary; §5 failure-mode table (F-11…F-19); §6 exclusions; §7 status; §8 preconditions; superseded-topology banner docs/DEPLOYMENT_ARCHITECTURE.md
Active topology; components; Mermaid topology + wake sequence; §3 per-component responsibilities/env vars; §4 five old blockers; §5 CPU-first reconciliation; §6 preconditions docs/DEPLOYMENT_TOPOLOGY.md
Orchestrator: _github_token(), _codespace_name() (strip = B-02), _codespace_port() 8000, _wake_timeout_s() 120, _upstream_timeout_s() 90, _DEV_ORIGINS, _PRODUCTION_ORIGINS, _allowed_origins(), error classes, _envelope(), ensure_codespace_up(), _proxy(), create_app() (four /api/* routes), _handle_orchestrator_error(); the superseded-by-tunnel docstring deploy/render/main.py
Orchestrator service declaration: start command, healthCheckPath: /api/health, env-var names, sync: false on secrets render.yaml
Service entrypoint: app = build_space_app(); uvicorn.run(host="0.0.0.0", port=...); PORT default 8000 deploy/codespace/serve.py
Launcher: preflight deps; port/stamp guards; _restart_serve(); the supervised tunnel-agent loop; the env vars (SATQUERY_DEVICE=cpu, SATQUERY_ASSET_ENABLED=1, SATQUERY_ASSET_DIR, SATQUERY_HUB_URL); the "announced to hub" verification deploy/codespace/launch.sh
Captured run: run_id, transport: "tunnel", config_hash, confidence + temperature + calibration_samples, warnings, steps frontend/assets/data/anatomy-run.js
Deployed HEADs (2d7ae53b482d, 89d80eaddec5, 5a0936ace491); B-07 OPEN patch prepared not deployed; B-02 cosmetic OPEN; no E2E benchmark; live validation 24 runs / 0 mock nodes / 94.4444 % release/DOCS_STYLE_GUIDE.md
Commits; live topology; E2E run-id table; metrics; blockers; test results (94 + 183 passed); truthfulness statement docs/FINAL_DELIVERY_REPORT.md
Status board; artifact inventory; real measured metrics; nine known blockers (incl. item 9 Cloudflare concatenation); blocker register B-01…B-08; evidence register E-01…E-14; final verification checklist docs/FINAL_DELIVERY_TODO.md