Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
08 — The API Contract
Parent: Architecture hub · Status tags: IMPLEMENTED · VERIFIED ·
MEASURED · NOT RUN · OPEN
Sources of truth for this chapter, all read in full before writing:
| Source | What it establishes |
|---|---|
core/schemas.py (462 lines) |
the binding typed contract: AnalysisRequest, ResultEnvelope, HealthStatus, SpecialistResult, ExecutionTrace, Task, … |
core/errors.py (315 lines) |
the 23-code error taxonomy, recoverable defaults, scrub_paths |
app/space_app.py (735 lines) |
the four Codespace endpoints, build_space_app(), the G-1 and return-annotation traps, the five entrypoint requirements |
gateway/policy.py (937 lines) |
_CODE_STATUS, DEFECT_CODES, GATEWAY_ORIGIN_CODES, GatewayConfig, admit(), CORS, body validation, translate_error |
gateway/app.py (691 lines) |
PROXIED_ROUTES, BLOCKED_ROUTES, COSTLY_ROUTES, _proxy(), the no-retry rule |
gateway/assets.py (516 lines) |
AssetStore, read_body_bounded, handle opacity, TTL, capacity |
deploy/render/main.py (532 lines) |
the /api/* gateway mirror (monorepo copy) |
deploy/render/codespaces.py (178 lines) |
the GitHub Codespaces control-plane client |
render.yaml (25 lines) |
the Render blueprint's env-var declarations |
docs/API_CONTRACT.md (916 lines) |
the client-facing specification, read fully |
docs/DEPLOYMENT_ARCHITECTURE.md §3.3, §2.3 |
the entrypoint requirements and the code-passthrough rule |
docs/DEPLOYMENT_TOPOLOGY.md (248 lines) |
the active topology and the measured live env-var set |
docs/FRONTEND_INTEGRATION.md (417 lines) |
the integration guide, and what it says is not guaranteed |
The one rule that governs this whole chapter. A claim here is only as good as the file it came from. Where the code and a document disagree, the code is authoritative and the disagreement is stated. Where neither answers, this chapter writes
UNKNOWN — not established from the available evidence.
1. Scope, and where this subsystem sits
The API contract is the outermost typed surface of SatQuery AI. Everything inside the system — router, planner, controller, specialists, evidence engine, confidence stage — is reachable only through four HTTP endpoints. There is no fifth door, no streaming channel, and no persistent session.
The chapter covers:
- the four Codespace endpoints and their
/api/*gateway mirror (§2–§4); - the cheap / COSTLY split and what each class is allowed to do (§3);
- the request and response envelopes, with real JSON (§5–§7);
- the error taxonomy, its machine codes and the
recoverableflag (§8); - the
x-satquery-transportheader and what it proves (§4.3); - the size and limit rules (§9);
- the CORS allowlist rule, which is never
*(§10); - the HARD RULE that the gateway must not retry
POST /api/infer(§11); - the verified examples (§12);
- the five entrypoint requirements (§13);
- the two annotation traps that make a FastAPI app silently wrong (§14);
- what is
NOT RUN/OPEN/BLOCKED(§15) and where the evidence lives (§16).
1.1 Two vocabularies, deliberately
The system publishes two path vocabularies and they are not interchangeable:
| Vocabulary | Owner | Path shape | Audience |
|---|---|---|---|
/v1/* |
the inference service (app/space_app.py) |
/v1/health, /v1/capabilities, /v1/analyze, /v1/assets |
the gateway, and any direct caller of the Codespace |
/api/* |
the orchestrator (deploy/render/main.py) |
/api/health, /api/infer, /api/capabilities, /api/assets |
the browser |
The browser talks only to /api/*. The /v1/* surface is the Codespace's own; the orchestrator
holds the security boundary and is "the only public door" (frontend/assets/js/live.js:54). The
naming is not cosmetic — live.js records why the frontend cannot simply use /v1/*:
"The orchestrator's proxied routes (
deploy/render/main.py). These are NOT the Space's own/v1/*routes -- the browser never talks to the Space directly; the orchestrator is the only public door." (frontend/assets/js/live.js:52-54)
1.2 The contract's own status
docs/API_CONTRACT.md §8 states the status of each element. Reproduced because it is the contract's
own honest self-assessment and it must not be softened:
| Element | Status (verbatim from docs/API_CONTRACT.md §8) |
|---|---|
Endpoint surface (/v1/health, /v1/capabilities, /v1/analyze, /v1/assets) |
Fixed — 3 by the plan, the 4th by the owner ruling of 2026-09-22 |
| Request/response shapes | Existing and tested — core/schemas.py |
Error taxonomy and code values |
Existing and tested — core/errors.py |
Error envelope ({"error": {...}}) |
Specified here. The gateway must produce it; the Space's own errors are translated by the gateway |
POST /v1/assets |
Implemented, Option A |
Multipart upload into /v1/analyze |
Not implemented, and not chosen — Option B was rejected |
| Authentication | Deliberately absent (plan §74) |
| Streaming / progress | Not in v1 |
| Rate-limit values | Not specified by the plan. The gateway must choose them; ask the maintainer |
| Asset TTL, size cap and content-type allowlist values | Deployment configuration, not contract constants |
"Nothing in the 'Status' column above may be treated as settled if it says 'Not in v1', 'Not implemented' or 'Not specified by the plan' — unless the row also names a decision that closed it. Those are gaps this document surfaces rather than fills." (
docs/API_CONTRACT.md§8)
2. The four endpoints
The surface is four endpoints. The original plan fixed three; the owner ruling of 2026-09-22
added the fourth by choosing Option A for upload (docs/API_CONTRACT.md §2).
| # | Method | Path (Codespace) | Purpose | Auth | Class |
|---|---|---|---|---|---|
| 1 | GET |
/v1/health |
Liveness + which models are loaded | none | cheap |
| 2 | GET |
/v1/capabilities |
What this deployment can actually do right now | none | cheap |
| 3 | POST |
/v1/analyze |
Run one analysis request | none | COSTLY |
| 4 | POST |
/v1/assets |
Upload one image out of band; returns an opaque handle | none | COSTLY |
The endpoint table is declared in code in app/space_app.py, where the routes are registered:
@api.get("/v1/health")
async def health() -> JSONResponse: ...
@api.get("/v1/capabilities")
async def capabilities() -> JSONResponse: ...
@api.post("/v1/assets")
async def assets(request: Request) -> JSONResponse: ...
@api.post("/v1/analyze")
async def analyze(payload: dict[str, Any]) -> JSONResponse: ...
(app/space_app.py:521, :549, :555, :661)
The four-endpoint statement is repeated in the entrypoint's own module docstring:
"The app itself is where the contract's four endpoints are served (
/v1/health,/v1/capabilities,/v1/analyze,/v1/assets-- the fourth per the owner ruling of 2026-09-22); the gateway sits in front of it and holds the security boundary." (app/space_app.py:412-416)
2.1 Why the fourth endpoint exists at all
AnalyzeRequest.assets is list[str] — asset handles, not bytes — and the original plan defines no
upload endpoint. docs/API_CONTRACT.md §2.5.1 records the gap as "NOT IN THE PLAN" and preserves
the phrase "because it is the finding, and a decision record that deletes the problem it solved is
not a record." The two options were:
| Option | Shape | Trade-off |
|---|---|---|
| A. Out-of-band upload — CHOSEN | POST /v1/assets → {"asset_id": "...", "expires_at": "..."}. Frontend uploads first, then calls /v1/analyze with the returned ids |
Keeps /v1/analyze JSON-only and lets the gateway enforce a size limit before the JSON body is parsed. Costs one extra round trip |
| B. Inline multipart — not chosen | /v1/analyze accepts multipart/form-data directly |
One round trip. Couples upload and analysis; a retry re-uploads |
The three things the frontend needs — a per-file size limit, a content-type allowlist, and an idempotency story for retries — were resolved by Option A. Note the third resolved to "there is none, a retry mints a new handle", and the contract explains why that is a decision rather than an omission:
"with no request key in the contract, a deduplicating server would have to hash payloads, and a content-hash handle is exactly the guessable identifier §2.5 forbids." (
docs/API_CONTRACT.md§2.5.1)
2.2 The multipart form on /v1/analyze is specified but not implemented
docs/API_CONTRACT.md §2.4 documents a multipart/form-data request shape for /v1/analyze:
| Part | Type | Notes |
|---|---|---|
assets |
file, repeatable | 1–2 image files. Field name repeats for the pair |
request |
text | A JSON string of the AnalysisRequest body with assets omitted |
and then states its own status plainly:
"Not yet implemented. The multipart entry point is part of the gateway's contract but the reference implementation serves the JSON form only. […] Build the frontend against the JSON form, which pairs with
POST /v1/assets." (docs/API_CONTRACT.md§2.4)
The shipped handler confirms this: analyze(payload: dict[str, Any]) reads a JSON body and validates
it with AnalysisRequest.model_validate(payload) (app/space_app.py:662-683). No multipart parsing
exists on that path.
3. Cheap versus COSTLY — the distinction that orders everything
The gateway splits its routes into two classes. This is not documentation prose; it is a tuple in the code:
#: Routes the gateway rate-limits, because they cost GPU quota or disk.
COSTLY_ROUTES: tuple[str, ...] = ("/v1/analyze", "/v1/assets")
(gateway/app.py:199)
and the rate limiter is applied only when the route is costly:
# 4. Rate limiting, only for routes that cost GPU quota. Rate-limiting
# health checks would make the frontend's load probe fail for no gain.
if cost:
allowed, remaining, retry_after = self.limiter.check(identity.key())
(gateway/policy.py:620-623)
| Route | Class | Rate-limited? | May load a model? | Cost |
|---|---|---|---|---|
GET /v1/health |
cheap | no | never | CPU only; no GPU, no weights |
GET /v1/capabilities |
cheap | no | never | filesystem + config reads |
POST /v1/analyze |
COSTLY | yes | yes, lazily | the only route that consumes GPU quota |
POST /v1/assets |
COSTLY | yes | no | writes disk; consumes one of a bounded number of handles |
docs/API_CONTRACT.md §2.4 states the analyze cost in one line: "Run one analysis. This is the only
endpoint that can consume GPU quota."
3.1 Why upload is COSTLY even though it touches no GPU
gateway/app.py records the reasoning, because grouping upload with analyze is the non-obvious call:
"
docs/API_CONTRACT.mdsection 6 tells the frontend to "serialize requests" and warns that every/v1/analyzecosts quota.POST /v1/assetsdoes not touch the GPU, but it does write to the Space's disk and consume one of a bounded number of handles (gateway/assets.py), so an unthrottled upload loop is a cheap denial of service against a 5-GPU-minute deployment. It is therefore rate-limited alongside analyze." (gateway/app.py:188-193)
and the two allowlists are kept separate on purpose:
"This is a separate allowlist from
PROXIED_ROUTESbecause the two answer different questions -- "may this reach the Space at all?" and "does it cost a metered resource?" -- and collapsing them would make the rate limiter's coverage depend on the proxy allowlist." (gateway/app.py:195-198)
3.2 The costly classification is passed explicitly, not derived from the path
_proxy() passes is_analyze=path in COSTLY_ROUTES rather than letting admit() infer it:
is_analyze=path in COSTLY_ROUTES,
(gateway/app.py:437)
with the reason recorded at the call site:
"Explicit rather than derived from the path.
policy.admit's own docstring says tests pass this so "a route rename cannot silently disable rate limiting"; passing it here means ADDING a costly route cannot silently miss the limiter either, which is exactly the mistake this would otherwise have made for/v1/assets." (gateway/app.py:432-436)
GatewayPolicy.admit() accepts the same override for the same reason:
"
is_analyze: override for the "this route costs GPU" test. Defaults to a path check. Tests pass it explicitly so a route rename cannot silently disable rate limiting." (gateway/policy.py:579-581)
3.3 What a cheap route is forbidden to do
Requirement 4 of the entrypoint requirements (§13) is the operative prohibition:
"Never load a model for a metadata request. Health and capabilities read artifact presence (filesystem) and configuration, not weights." (
docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
docs/API_CONTRACT.md §2.1 states the consequence for the caller:
"This endpoint is answered without loading any model and without importing torch — the device is resolved from configuration, not by probing the runtime. A liveness probe that built the world would consume GPU quota to say "I am alive"."
The corresponding implementation note in app/space_app.py records that this was not free:
"
docs/DEPLOYMENT_ARCHITECTURE.mdsection 3.3 […] two things there did import torch:build_serving_registry()(viaConfig.device_preference, a@propertythat calls_torch_cuda_available()) and any read of that property. Both were removed: the adapter enumerates capabilities fromdefault_specs()and resolves the device from environment and configuration only. The testtest_the_metadata_path_does_not_import_torchruns the import in a subprocess and assertstorch imported: False." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3, requirement 1)
4. The gateway mirror — /api/* ⇄ /v1/*
4.1 The mapping
deploy/render/main.py declares the mapping in its module docstring:
Proxied routes (never answered locally)
---------------------------------------
POST /api/infer -> POST {codespace}/v1/analyze
GET /api/capabilities-> GET {codespace}/v1/capabilities
POST /api/assets -> POST {codespace}/v1/assets
(deploy/render/main.py:22-26)
/api/* (browser) |
/v1/* (Codespace) |
Answered locally? |
|---|---|---|
GET /api/health |
— (none) | yes — the orchestrator's own liveness |
GET /api/capabilities |
GET /v1/capabilities |
no — proxied |
POST /api/infer |
POST /v1/analyze |
no — proxied |
POST /api/assets |
POST /v1/assets |
no — proxied |
The route registration in the monorepo copy:
@app.get("/api/health")
async def health() -> dict[str, Any]: ...
@app.post("/api/infer")
async def infer(request: Request) -> JSONResponse: ...
@app.get("/api/capabilities")
async def capabilities() -> JSONResponse: ...
@app.post("/api/assets")
async def assets(request: Request) -> JSONResponse: ...
(deploy/render/main.py:444, :468, :491, :497)
/api/health is deliberately not a proxy:
"Orchestrator liveness. Reports its own configuration; never answers for the Codespace (that is /api/capabilities)." (
deploy/render/main.py:446-447)
and the design reason for keeping the two healths apart is recorded in the other gateway
implementation, gateway/app.py:
"Separate from
/v1/healthon purpose: conflating them would make a gateway that is up but whose upstream is down indistinguishable from a gateway that is itself broken. This route never touches the Space, so it costs nothing." (gateway/app.py:243-247)
4.2 No second copy of the capability table
The orchestrator does not decide capabilities. deploy/render/main.py states this as a design
rule:
"There is deliberately no second copy of the capability table here; the gateway proxies
/v1/capabilitiesand nothing else decides that question." (deploy/render/main.py:28-29)
The handler confirms it:
@app.get("/api/capabilities")
async def capabilities() -> JSONResponse:
"""Proxy ``GET /v1/capabilities`` — no local capability table."""
base, _ = await ensure_codespace_up()
return await _proxy("GET", f"{base}/v1/capabilities")
(deploy/render/main.py:491-495)
This is the orchestrator-side expression of the "single capability authority" ruling that
docs/DEPLOYMENT_ARCHITECTURE.md §3.3.1 records for the inference side.
4.3 The x-satquery-transport header — the measured proof of the path taken
The measured transport value is tunnel. docs/FINAL_DELIVERY_TODO.md §6 E-03 records the
verification:
"
E-03| P3-T01 |curl …/api/capabilities,POST /api/infer {}| 200 (6× available:true); 422invalid_request, headerx-satquery-transport: tunnel| VERIFIED"
and §1.4:
"Tunnel up / warm | VERIFIED | health
tunnel.agent_connected:true;/api/inferreturns headerx-satquery-transport: tunnel" (docs/FINAL_DELIVERY_TODO.md§1.4)
The client reads it deliberately, before the response object is discarded:
/* Read the transport/state headers BEFORE parsing: they are evidence
about WHICH path served the request, and they are gone once the
response object is discarded. `x-satquery-transport: tunnel` is the
proof that Render forwarded to the Codespace rather than answering
locally. */
var state = resp.headers.get('X-SatQuery-State') || '';
var transport = resp.headers.get('x-satquery-transport') || '';
(frontend/assets/js/live.js:311-317)
| Header | Values | Meaning | Source |
|---|---|---|---|
x-satquery-transport |
tunnel (measured) |
the request was forwarded to the Codespace rather than answered locally | docs/FINAL_DELIVERY_TODO.md §6 E-03; frontend/assets/js/live.js:315 |
X-SatQuery-State |
waking | ready |
whether a cold start occurred | deploy/render/main.py:488 |
X-SatQuery-State is set on the /api/infer path in the monorepo copy:
out.headers["X-SatQuery-State"] = "waking" if woke else "ready"
(deploy/render/main.py:488)
The tunnel header is not set by the monorepo copy.
deploy/render/main.pyis 532 lines with no tunnel code at all; the deployedSatQuery-Backend/main.pyis 768–769 lines with it (docs/FINAL_DELIVERY_TODO.md§1.1;DELIVERY_REPORT_2026-09-25.md§4). The monorepo copy therefore documents the contract of the route, while thetunnelvalue is a measured live fact recorded in the delivery evidence. See §15 for the consequence.
docs/DEPLOYMENT_TOPOLOGY.md §2 states the measured transport shape in full:
"Measured 2026-09-25 (live). Transport is an outbound tunnel, not a polled forwarded port: the Codespace runs
deploy/codespace/tunnel_agent.py, which dials out toPOST /tunnel/agent(long-poll) and executes againsthttp://127.0.0.1:8000locally. When the Codespace is stopped the agent stops polling →GET /api/healthreportstunnel.agent_connected:falseandPOST /api/inferparks untilSATQUERY_TUNNEL_TIMEOUT_S(150 s), then returnstunnel_offline(503,recoverable:true)."
4.4 The gateway's own route allowlists
The deployed gateway's allowlists live in SatQuery-Backend/main.py, but the monorepo copy in
gateway/app.py declares the same three tuples and their reasoning:
PROXIED_ROUTES: tuple[str, ...] = (
"/v1/health",
"/v1/capabilities",
"/v1/analyze",
"/v1/assets",
)
BLOCKED_ROUTES: tuple[str, ...] = ()
COSTLY_ROUTES: tuple[str, ...] = ("/v1/analyze", "/v1/assets")
(gateway/app.py:170-199)
BLOCKED_ROUTES is empty, and should stay that way:
"EMPTY, and it should stay that way: this tuple exists so a route the contract discusses but the server does not implement answers 501 with a reason instead of a 404 that a frontend developer would debug as a typo. Nothing is in that state right now." (
gateway/app.py:177-184)
"An allowlist, not a passthrough: a gateway that forwarded arbitrary paths would expose every route the Space happens to serve, including ones the contract does not document." (
gateway/app.py:161-163)
5. GET /v1/health — the cheap liveness probe
5.1 Shape
The response shape is HealthStatus (core/schemas.py:430):
class HealthStatus(BaseModel):
model_config = ConfigDict(extra="forbid")
status: Literal["ok", "degraded", "error"] = "ok"
schema_version: str = SCHEMA_VERSION
models: dict[str, str] = Field(default_factory=dict)
device: str | None = None
gpu_available: bool = False
5.2 A real measured response
docs/API_CONTRACT.md §2.1 publishes the measured output of a deployment where the CROMA
checkpoint is not shipped:
{
"status": "degraded",
"schema_version": "1.0",
"models": {
"caption": "not_requested",
"change": "not_requested",
"change_vqa": "not_requested",
"grounding": "not_requested",
"optical_sar": "absent",
"vqa": "not_requested"
},
"device": "cpu",
"gpu_available": false
}
(docs/API_CONTRACT.md §2.1)
5.3 Field by field
| Field | Type | Notes (verbatim where quoted) |
|---|---|---|
status |
"ok" | "degraded" | "error" |
degraded = the service is up but at least one capability is not servable. Derived, not asserted: any absent capability makes the service degraded; any unavailable makes it error |
schema_version |
string |
Always present; "1.0" (core/schemas.py:21) |
models |
object<string,string> |
Per-capability state. Values are strings, not booleans, so a reason can be carried |
device |
string | null |
"cpu", "cuda", "mps", or null if unknown |
gpu_available |
boolean |
Whether a CUDA/MPS device was detected |
Every capability the registry resolves appears in models, and the set is identical to
capabilities[].task:
"the two endpoints are generated from one source, so they cannot enumerate different capabilities. A key is never absent; a capability that cannot be served is reported with a state, not by omission." (
docs/API_CONTRACT.md§2.1)
The implementation asserts this rather than trusting it:
payload = health_payload()
# Asserted rather than trusted: `HealthStatus` is `extra="forbid"`, so a
# key added to the payload without a key added to the model would make
# the response invalid against the project's own contract. Finding H-1
# was exactly this failure in the other direction.
from core.schemas import HealthStatus
HealthStatus.model_validate(payload)
return JSONResponse(payload)
(app/space_app.py:537-547)
5.4 device is a closed set, and null means the value was not understood
This is F-8, and it is worth restating because the failure mode is a false statement about the deployment:
"This row has always published four legal values, but the reader accepted any string and echoed it into the field, so
SATQUERY_DEVICE=garbageserved{"device": "garbage"}— a value the frontend has no rendering for. The reader now casefolds and validates against the set above; anything unrecognised is served asnull.nullis deliberately not a silent"cpu": reporting the CPU because the operator mistyped would be a false statement about the deployment, and it is the same mistake that F-7 fixed in a different variable." (docs/API_CONTRACT.md§2.1)
The invariant the frontend may rely on:
"if
device == "cuda"thengpu_availableistrue. The converse does not hold — a GPU may exist whiledeviceis"cpu"(the operator chose it, or the config did)." (docs/API_CONTRACT.md§2.1)
5.5 gpu_available: false is expected, not a fault
"Important for the frontend:
gpu_available: falseon a ZeroGPU Space is expected, not an error. ZeroGPU allocates the GPU only for the duration of a decorated call. Do not surface this as a fault." (docs/API_CONTRACT.md§2.1)
On the active topology the declared device is cpu outright (render.yaml:20-21,
docs/DEPLOYMENT_TOPOLOGY.md header), so device: "cpu" / gpu_available: false is the expected
warm state rather than a transient one.
6. GET /v1/capabilities — what this deployment can do right now
6.1 The rule
"What the deployment can do right now, derived from actual artifact presence — not from what the code could theoretically do." (
docs/API_CONTRACT.md§2.2)
"Every capability the registry resolves is listed, including ones this deployment cannot serve. A capability that cannot be served is reported
available: falsewith a reason, never omitted: omitting it would make it invisible to the frontend, which cannot disable an affordance it was never told about." (docs/API_CONTRACT.md§2.2)
6.2 A real measured response
docs/API_CONTRACT.md §2.2 publishes the measured output where four of the six capabilities lack
runtime dependencies and optical_sar lacks its CROMA checkpoint:
{
"schema_version": "1.0",
"capabilities": [
{
"task": "change",
"available": true,
"reason": null,
"requires_pair": true,
"max_assets": 2
},
{
"task": "change_vqa",
"available": true,
"reason": null,
"requires_pair": true,
"max_assets": 2
},
{
"task": "optical_sar",
"available": false,
"reason": "the CROMA backbone checkpoint (CROMA_base.pt) is not present in this deployment; without it optical/SAR fusion degrades to sensor-only; the trained optical/SAR fusion head is not present in this deployment; without it no fused prediction is produced",
"requires_pair": true,
"max_assets": 2,
"modalities": ["optical", "sar"]
},
{
"task": "caption",
"available": true,
"reason": "the SmolVLM weights are fetched from the Hugging Face Hub on first use and no local checkpoint_path is configured in this deployment",
"requires_pair": false,
"max_assets": 1
}
],
"deployment": {
"platform": "huggingface-spaces",
"zerogpu": true,
"lazy_load": true,
"cache_max_models": 1,
"torch_compile": false
}
}
(docs/API_CONTRACT.md §2.2 — abridged: the real response lists all six)
"note that
optical_sar's names two missing artifacts, because two are required and both are absent." (docs/API_CONTRACT.md§2.2)
6.3 Field by field
| Field | Type | Notes |
|---|---|---|
capabilities[].task |
string |
One of the Task enum values |
capabilities[].available |
boolean |
Whether the task can be served right now |
capabilities[].reason |
string | null |
Required when available is false. "A bare false with no reason is not compliant" |
capabilities[].modalities |
string[] |
Optional; present only for optical_sar |
capabilities[].requires_pair |
boolean |
Whether two assets are required |
capabilities[].max_assets |
integer |
Maximum assets accepted |
deployment.platform |
string |
Deployment target, e.g. "huggingface-spaces" |
deployment.zerogpu |
boolean |
Whether GPU work runs under ZeroGPU's per-call allocation |
deployment.lazy_load |
boolean |
true means models load on first use |
deployment.cache_max_models |
integer |
Resident-model cap. 1 means requests serialize |
deployment.torch_compile |
boolean |
Always false |
torch_compile carries a specific obligation:
"Always
false.torch.compileis unsupported on ZeroGPU and the config loader hard-fails ontrue(finding C-8). Echoed here so an operator can confirm the constraint from a single response." (docs/API_CONTRACT.md§2.2)
6.4 A reason on an available capability is not a defect
This is the single most likely misreading of the endpoint, and the contract calls it out twice:
"A
reasonon an available capability is not a defect. Three capabilities above areavailable: trueand still carry a reason — it reads "…fetched from the Hub on first use, no local checkpoint configured". That is not an error; it is a disclosure that the first request will be slow and will need egress. A frontend that treats a non-nullreasonas a failure will mislay every cold start." (docs/API_CONTRACT.md§2.2)
6.5 The capability state vocabulary — five words, closed
models in /v1/health uses a five-word closed vocabulary:
| Value | Meaning |
|---|---|
"loaded" |
Resident and ready |
"absent" |
The artifact is not present in this deployment. Permanent for this revision; not retryable |
"unavailable" |
Present but could not be loaded (corrupt, incompatible, dependency missing). This is a defect, distinct from absent |
"not_requested" |
Nothing has attempted to load it yet (normal with lazy_load: true) |
"evicted" |
Was loaded, was unloaded to make room (cache_max_models: 1) |
"
absentandunavailablemust not be conflated in the UI. Absent means "this build does not ship it"; unavailable means "this build ships it and it is broken"." (docs/API_CONTRACT.md§2.3)
"These five words are the complete permitted vocabulary. They are the contract's vocabulary and are not the registry's." (
docs/API_CONTRACT.md§2.3)
6.6 The translation layer, and what it must never leak
The adapter derives the contract state from the registry's declared spec table and the filesystem; it does not read a live registry state.
"
app/deployment.pyinspects the registry's declared spec table and the filesystem and derives the contract state from what it finds. It never callsbuild()/build_all(), because requirement 4 (DEPLOYMENT_ARCHITECTURE.md§3.3) forbids loading a model to answer a metadata request. A live registry state is therefore not observable on this path, and the registry's word is reconstructed from the contract state — not translated into it." (docs/API_CONTRACT.md§2.3.1)
The exhaustive table of what the adapter can produce:
| Contract state | When it is emitted | available |
Why |
|---|---|---|---|
not_requested |
All declared shipped artifacts are present, and nothing has attempted a load. The normal healthy state under lazy_load: true |
true |
Nothing is missing. Emitting loaded here would claim a model was resident, which cannot be known without loading one |
absent |
A required shipped artifact is not on disk in this deployment | false |
Nothing is broken; the deployment does not ship it |
unavailable |
Construction was attempted in this process and failed (defect path only) | false |
Present but broken — a genuine defect |
evicted |
(never emitted) | — | A runtime model-cache fact. No static inspection can observe it |
loaded |
(never emitted) | — | See above |
(docs/API_CONTRACT.md §2.3.1)
available: trueandmodels: "not_requested"coexist by design, and that is not a contradiction. "The two fields answer different questions:availableis "can this deployment serve this capability?" andnot_requestedis "has anything loaded it yet?". Underlazy_load: truethe healthy answer to the second is no, not yet — for every capability, including ones that will work perfectly on the first request." (docs/API_CONTRACT.md§2.3.1)
The registry's own vocabulary is internal and is never served on any endpoint.
6.7 The contract obligations that fall on the client
*"- The frontend MUST build its UI affordances from this response, not from a hardcoded list. A capability that is
available: falsemust be shown as disabled with itsreasondisplayed — never hidden, never silently downgraded to a different task.
- The deployment block echoes
configs/deploy.yaml. Notecache_max_models: 1: at most one model is resident. Concurrent requests for different specialists will evict each other, so the frontend must not assume parallel throughput."* (docs/API_CONTRACT.md§2.2)
6.8 The deployment block on the live deployment is stale metadata
docs/FINAL_DELIVERY_TODO.md §1.7 item 6 records this as a known defect:
"**
/api/capabilitiesdeploymentblock* claimshuggingface-spaces/zerogpu— stale metadata."*
The block is generated from configs/deploy.yaml, which is frozen paperwork describing an HF
Space + Gradio + ZeroGPU target that no longer matches the active Render + tunnel topology
(docs/DEPLOYMENT_TOPOLOGY.md §3.4; docs/DEPLOYMENT_DECISION.md §4). The capability list is
live and correct (six capabilities, available: true); the deployment block is not. This is
recorded as a defect rather than smoothed over, because a reader who trusts platform: "huggingface-spaces" would look for a Space that does not exist.
7. POST /v1/analyze — the one endpoint that runs inference
7.1 The request shape — AnalysisRequest
class AnalysisRequest(BaseModel):
model_config = ConfigDict(extra="forbid")
assets: list[str] = Field(min_length=1)
query: str
force_task: Task | None = None
run_id: str | None = None
(core/schemas.py:412-418)
Four fields, and extra="forbid" — an unknown field is a 422, not a silently ignored one.
A real request
{
"assets": ["asset_0", "asset_1"],
"query": "How has the built-up area changed between these two dates?",
"force_task": "change_vqa",
"run_id": "9f2c1c0e-4a5b-4f5e-9a2c-1b3d4e5f6a7b"
}
(docs/API_CONTRACT.md §2.4)
Field by field
| Field | Type | Required | Notes |
|---|---|---|---|
assets |
string[] |
yes | Minimum length 1. Values are asset handles returned by the upload step (§7.6), not base64 and not URLs |
query |
string |
yes | Natural language. Empty string is permitted by the schema but will route to an unsupported_query error in practice |
force_task |
string | null |
no | One of the Task values. Bypasses the intent router |
run_id |
string | null |
no | Client-supplied correlation id. If omitted the server generates one. The server always echoes a run_id in the response |
(docs/API_CONTRACT.md §2.4)
assets min_length 1 is enforced at two layers
The schema declares Field(min_length=1) (core/schemas.py:415), and the gateway refuses an empty
array before it can cost a round trip:
assets = parsed.get("assets")
if not isinstance(assets, list) or not assets:
return None, (
422,
translate_error(
"invalid_request",
"`assets` must be a non-empty array of asset handles.",
detail=f"assets={assets!r}",
)[1],
)
(gateway/policy.py:815-824)
The same function refuses a non-string entry, a missing/non-string query, an unknown force_task
value, and a non-string run_id — all as 422 invalid_request, all before the Space is called:
force_task = parsed.get("force_task")
if force_task is not None:
if not isinstance(force_task, str):
... # 422
allowed = _task_values()
if allowed is not None and force_task not in allowed:
... # 422 `force_task` is not a recognised task
(gateway/policy.py:857-880)
The force_task check was missing and its omission was not harmless:
"This check was MISSING and the omission was not harmless: a body with
force_task: "nonsense"was forwarded to the Space, whose schema rejected it -- so the client received a 502/upstream error for a defect entirely local to the request. That both mis-states the fault and spends a GPU-quota round trip on a body the Space cannot accept." (gateway/policy.py:846-851)
force_task is derived from the Task enum, never hard-coded
def _task_values() -> frozenset[str] | None:
try:
from core.schemas import Task
return frozenset(member.value for member in Task)
except Exception: # pragma: no cover - only in a broken install
return None
(gateway/policy.py:896-919)
and the asymmetry between the field check and the enum check is deliberate:
"An unknown-field check needs a field set, and the documented four field names are a short, stable list that has not changed since the contract was written. The
Taskvalues are a different case: they are the router's vocabulary and the schema is explicit that it may grow (core/schemas.py). A stale copy here would silently reject a newly-added task at the gateway, before the Space could accept it -- a gate that fails closed on valid input, which is worse than no gate." (gateway/policy.py:900-909)
Unknown fields are rejected, and the correction matters
docs/API_CONTRACT.md §1.1 records a corrected claim — the earlier text asserted forward
compatibility on read, and the code says otherwise:
*"
ResultEnvelope,HealthStatus,AnalysisRequestand every other contract-facing model incore/schemas.pysetsextra="forbid"— with exactly one exception,GeoMetadata, which setsextra="allow"."* (docs/API_CONTRACT.md§1.1)
GeoMetadata is the one open surface:
class GeoMetadata(BaseModel):
model_config = ConfigDict(extra="allow")
(core/schemas.py:120-121)
"It is the geospatial descriptor attached to
AssetMetadata.geoandSpecialistResult.geospatial, so it is reachable in every/v1/analyzeresponse. The reason is that a raster reader supplies whatever tags the source file carries, and forbidding unknown keys there would discard provenance a caller may need." (docs/API_CONTRACT.md§1.1)
The corrected consequences, stated plainly:
- Additive changes are not free. Adding a field to a response breaks any client that validates strictly.
- A version bump is required when a field is added, not only when one is removed.
- Clients should be written permissively even though the server is strict — a client-side robustness measure, not a server guarantee.
7.2 The response shape — ResultEnvelope
class ResultEnvelope(BaseModel):
model_config = ConfigDict(extra="forbid")
run_id: str
result: SpecialistResult
trace: ExecutionTrace
schema_version: str = SCHEMA_VERSION
(core/schemas.py:421-427)
7.3 A real response
docs/API_CONTRACT.md §2.4 publishes the measured shape:
{
"run_id": "9f2c1c0e-4a5b-4f5e-9a2c-1b3d4e5f6a7b",
"schema_version": "1.0",
"result": {
"task": "change_vqa",
"answer": "The built-up area increased...",
"labels": [],
"regions": [],
"boxes": [
{
"x1": 0.12, "y1": 0.34, "x2": 0.56, "y2": 0.78,
"label": "expanded built-up area",
"score": 0.81,
"coordinate_system": "normalized_0_1"
}
],
"masks": [],
"change_map": null,
"evidence": [
{
"evidence_id": "ev_001",
"type": "change_map",
"score": 0.72,
"source_specialist": "change_vqa",
"coordinate_system": "normalized_0_1",
"coordinates": [0.12, 0.34, 0.56, 0.78],
"artifact_ref": null,
"payload": {}
}
],
"confidence": {
"raw": 0.991,
"calibrated": 0.987,
"method": "temperature_scaling",
"components": {},
"degraded": false,
"degradation_reason": null
},
"geospatial": {},
"execution_trace": null,
"schema_version": "1.0",
"warnings": [],
"degraded": false
},
"trace": {
"run_id": "9f2c1c0e-4a5b-4f5e-9a2c-1b3d4e5f6a7b",
"task": "change_vqa",
"intent": null,
"query": "How has the built-up area changed?",
"modalities": ["optical"],
"workflow": [],
"steps": [],
"timings": {},
"selected_models": [],
"parameters": {},
"config_hash": "78f1e3700da15aa1",
"inputs": [],
"outputs": [],
"errors": [],
"fallbacks": [],
"contradiction": false,
"validation": {},
"confidence": null,
"started_at": "2026-09-22T04:12:21.000Z",
"finished_at": "2026-09-22T04:12:29.400Z",
"schema_version": "1.0"
}
}
(docs/API_CONTRACT.md §2.4)
"
traceabove lists every fieldExecutionTracedefines. Fields leftnullor empty here are genuinely optional, not omitted from the contract — the model uses defaults, so they will normally be present in a real response." (docs/API_CONTRACT.md§2.4)
7.4 The fields a client must read correctly
| Field | Why it matters |
|---|---|
result.confidence.value |
NOT a JSON field. It is a Python @property on ConfidenceBreakdown and is not serialised. Read calibrated if it is non-null, otherwise raw |
result.confidence.method |
"uncalibrated" or "temperature_scaling" |
result.confidence.degraded / degradation_reason |
Whether the confidence is trustworthy. Display the reason verbatim when set |
result.degraded + result.warnings |
The result is served but something was degraded |
result.answer |
"" for non-VQA tasks. Empty is valid |
result.boxes[].coordinate_system |
Read this per box |
result.boxes[] flat geometry |
x1, y1, x2, y2 are flat fields on the box, not a nested box object. Region is the one with a nested box |
result.evidence[] shape |
Every item carries evidence_id, type, score, source_specialist, coordinate_system, coordinates, artifact_ref, payload. Note score — not value — and source_specialist — not source. artifact_ref is always null in v1 |
result.evidence[].type |
One of 11 EvidenceType values |
trace.steps[].state |
ControllerState — the pipeline stage. detail and duration_ms accompany it |
trace.steps |
Observable facts only. Never chain-of-thought (plan §26). Safe to display |
trace.config_hash |
The frozen config identity. 78f1e3700da15aa1 for this revision |
(docs/API_CONTRACT.md §2.4)
The value property is real and private to Python:
@property
def value(self) -> float:
return self.calibrated if self.calibrated is not None else self.raw
(core/schemas.py:271-273)
"NOT a JSON field. […] (verified:
model_dump()yields onlycalibrated, components, degradation_reason, degraded, method, raw)." (docs/API_CONTRACT.md§2.4)
The client implements the rule with ??:
const shown = c.calibrated ?? c.raw; // NOT c.value — it is not serialised
(docs/FRONTEND_INTEGRATION.md §4.2)
7.5 The confidence contract
ConfidenceBreakdown (core/schemas.py:259-273):
| Field | Meaning |
|---|---|
raw |
The uncalibrated score |
calibrated |
The post-calibration score, or null |
method |
"uncalibrated" or "temperature_scaling" |
components |
A string -> float map. May be empty. Diagnostic only — do not compute a confidence from it |
degraded |
Whether this confidence should be trusted |
degradation_reason |
Why, when degraded is true |
The four rules the frontend MUST follow:
"1. Display
calibratedwhen it is notnull; otherwise displayraw. 2. Displaymethodnext to the value.temperature_scalingmeans a fitted correction was applied;uncalibratedmeans it was not. 3. Never present a confidence as a percentage without its method. A raw 0.99 and a calibrated 0.99 do not mean the same thing. 4. Whendegradedistrue, showdegradation_reason. Confidence that is degraded is not a quality signal." (docs/API_CONTRACT.md§4)
The measured caveat, recorded honestly:
"The R-02 calibration fit (
artifacts/calibration_v001.json,T = 0.9772731820958189, 16,441 Val rows) found that the raw softmax was already near-calibrated (ECE 0.013755) and that temperature scaling made ECE very slightly worse (0.014929) while improving NLL marginally (0.689741 → 0.689631). The frontend must not imply thattemperature_scalingis inherently "more accurate" thanuncalibrated." (docs/API_CONTRACT.md§4)
7.6 Artifact refs are null in v1, and why
"Every
artifact_refandchange_mapin a v1 response isnull. This is a deliberate contract, not a missing value." (docs/API_CONTRACT.md§2.4)
F-16 (owner ruling 2026-09-23): never expose filesystem paths.
"The specialists do render their artifacts — the change map and the optical/SAR views are written server-side — but their location is an operator fact, not a client-facing one. A response that carried the server's path would disclose the deployment's directory layout to an unauthenticated caller, and nothing the frontend can do requires it." (
docs/API_CONTRACT.md§2.4)
"No
artifact://URI is fabricated in its place. v1 has no artifact-serving endpoint, so a URI would be a promise the service cannot keep — strictly worse thannull, because the frontend would build a link that 404s." (docs/API_CONTRACT.md§2.4)
What replaces the ref:
| Removed | Replaced by |
|---|---|
change_map path |
null, plus the change statistics in the CHANGE_MAP evidence's payload (total_change_pixels, n_components_kept, threshold) |
view artifact_ref path |
null, plus payload.rendered / payload.retrievable / payload.retrieval |
| — | an explicit warnings[] entry saying the artifact is NOT retrievable |
The schema carries the ruling in the field description:
artifact_ref: str | None = Field(
default=None,
description=(
"Reference to an externally retrievable artifact. NEVER a "
"filesystem path (F-16, owner ruling 2026-09-23): v1 exposes no "
"artifact-serving endpoint, so this is null unless a deployment "
"supplies a client-fetchable reference. An artifact may still be "
"written server-side where configured; being written is not the "
"same as being retrievable."
),
)
(core/schemas.py:225-235)
7.7 The handle → path translation happens in exactly one place
The controller needs a path to inspect; the contract carries handles. The handler is the one place that translates:
handle -> AssetStore.get() -> path -> AnalysisRequest.assets
(app/space_app.py:674)
store = get_asset_store()
try:
handles = store.resolve_many(list(request.assets))
except UnknownAssetError as exc:
... # input_error: "One or more asset handles are unknown or have expired."
# Rebuild the request with resolved PATHS. `model_copy` rather than
# mutating, because `AnalysisRequest` is the contract's model and a
# handler must not rewrite a validated request in place.
request = request.model_copy(
update={"assets": [str(handle.path) for handle in handles]}
)
(app/space_app.py:709-724)
"A handle that is unknown or expired is refused HERE, with a named error, rather than being passed to the controller as a path that does not exist -- which would surface as a raster read failure and name the wrong cause." (
app/space_app.py:676-678)
resolve_many is all-or-nothing:
"All-or-nothing: a partial resolution would let an analysis start with one of a required pair missing, which the specialists would then reject with a pairing error that names the wrong cause. Failing here names the real one." (
gateway/assets.py:400-406)
7.8 The success status
/v1/analyze returns 200 with the envelope. deploy/render/main.py's _proxy() passes the upstream
status through unchanged:
"Connection/transport errors and non-JSON upstream bodies are translated into the v1 envelope (
502,recoverable: true); the upstream status is otherwise passed through unchanged." (deploy/render/main.py:372-376)
8. POST /v1/assets — the upload endpoint (Option A)
8.1 The request
multipart/form-data with exactly one part, the file. The Content-Type of the part is the declared
type (docs/API_CONTRACT.md §2.5).
The shipped client sends raw bytes, not multipart, and says why:
/**
* Upload ONE File and return its asset handle.
*
* The body is the raw bytes with the derived Content-Type -- not multipart.
* `/v1/assets` reads the raw body (gateway/assets.py `read_body_bounded`), so
* wrapping the file in a form would store the multipart wrapper as the image.
*/
(frontend/assets/js/live.js:190-196)
return fetch(SQ.live.url('assets'), {
method: 'POST',
headers: { 'Content-Type': contentType },
body: file,
signal: opts.signal
})
(frontend/assets/js/live.js:215-220)
Note the divergence and do not smooth it over.
docs/API_CONTRACT.md§2.5 anddocs/FRONTEND_INTEGRATION.md§3.2 both describe the upload asmultipart/form-datawith a part namedfile, and the integration guide even warns "Do not setContent-Typemanually." The shipped client (live.js) sends raw bytes with an explicitContent-Typederived from the file extension. The Space's handler reads the raw body —raw, too_large = await read_body_bounded(request, _asset_max_file_bytes())(app/space_app.py:610) — and takes the type from the header —store.put(raw, content_type=request.headers.get("content-type"))(app/space_app.py:625-628). Raw-body upload is what the deployed path exercises; the multipart description in the two documents is not what the shipped client does. This is recorded rather than resolved, because only the raw path has been run.
The content type is derived from the extension on purpose:
/*: Sent explicitly rather than relying on `File.type`. Deliberate: a GeoTIFF
arrives as `""` in Chrome and Firefox, and an empty Content-Type is rejected
by the store. Deriving it from the extension means the client and the server
agree on the same fact. */
(frontend/assets/js/live.js:75-78)
8.2 The response 201
{
"asset_id": "asset_7c6f64a4a4c821e25d518467a1cc5d47",
"content_type": "image/png",
"bytes": 20481,
"expires_at": "2026-09-22T04:42:21.000Z"
}
(docs/API_CONTRACT.md §2.5)
The handler returns exactly this shape:
return JSONResponse(status_code=201, content=handle.to_response())
(app/space_app.py:659)
and to_response() deliberately omits the path:
def to_response(self) -> dict[str, Any]:
"""The `POST /v1/assets` response body.
`path` is deliberately absent. Returning it would hand a client a
server-side filesystem location -- an information disclosure, and an
invitation to construct a path directly instead of via a handle.
"""
return {
"asset_id": self.asset_id,
"content_type": self.content_type,
"bytes": self.bytes,
"expires_at": self.expires_at_iso,
}
(gateway/assets.py:219-231)
8.3 Opacity — the handle is the access control
"
asset_idisasset_+ 32 hex characters, fromsecrets.token_hex(16). It is 128 bits of entropy and carries no information about the upload — no filename, no type, no index, no position. There is no auth in v1 (§7), so this handle is the access control for the uploaded bytes" (docs/API_CONTRACT.md§2.5)
def _new_handle() -> str:
"""An opaque, unguessable handle.
`secrets`, not `random`: the handle is this endpoint's only access control
(`API_CONTRACT.md` section 7 -- there is no auth in v1). 128 bits from
`token_hex(16)` is not brute-forceable, and the prefix keeps a handle
recognisable in a log without making it derivable.
"""
return f"asset_{secrets.token_hex(16)}"
(gateway/assets.py:466-474)
The stored filename is derived from the content type, never the client's filename:
#: Content type -> stored suffix. Derived from the TYPE, never from the
#: client-supplied filename, so a hostile name cannot influence a path.
_SUFFIXES: Mapping[str, str] = {
"image/tiff": ".tif",
"image/geotiff": ".tif",
"image/png": ".png",
"image/jpeg": ".jpg",
"application/octet-stream": ".bin",
}
(gateway/assets.py:477-485)
8.4 The three guarantees the frontend depends on
| Concern | Guarantee |
|---|---|
| Opacity | 128 bits of entropy from secrets.token_hex(16); no filename, type, index or position |
| Size limit | A per-file byte cap, configurable per deployment. It is enforced at two layers and a client should rely on both. The gateway refuses an over-limit body from a declared Content-Length and, since the F-6 fix, while reading the bytes — so omitting the header does not evade it. The Space does the same since the F-9 fix, via the shared gateway/assets.py::read_body_bounded. An over-limit upload gets 413 and writes nothing |
| Content-type allowlist | A closed list of exactly five types. A request with no declared type is refused rather than defaulted. A disallowed type gets 415. Media-type parameters are ignored |
| Retries | There is no idempotency key. A retry is a new upload that mints a new handle |
(docs/API_CONTRACT.md §2.5)
The five allowed types, declared identically on both layers:
#: The content types this endpoint accepts. Mirrors
#: `GatewayConfig.allowed_content_types`; the Space's copy exists because the
#: Space validates independently (defence in depth) rather than trusting that
#: the gateway is the only caller.
_ALLOWED_ASSET_CONTENT_TYPES: tuple[str, ...] = (
"image/tiff",
"image/geotiff",
"image/png",
"image/jpeg",
"application/octet-stream",
)
(app/space_app.py:381-391)
and on the gateway side:
allowed_content_types: tuple[str, ...] = (
"image/tiff",
"image/geotiff",
"image/png",
"image/jpeg",
"application/octet-stream",
)
(gateway/policy.py:229-236)
"**
image/tiffis the type the geospatial specialists need* — a client that uploads only PNG/JPEG can serve the VQA, caption and grounding tasks but not the change or optical/SAR ones."* (docs/API_CONTRACT.md§2.5)
An absent type normalises to the empty string so the allowlist refuses it:
def _normalise_content_type(content_type: str | None) -> str:
"""Lower-case the type and strip parameters, or '' when absent.
`image/tiff; charset=binary` is a legitimate header and the parameter is not
part of the type. `None` becomes `''` so the allowlist check refuses it --
defaulting an absent type to `application/octet-stream` would make the
allowlist unenforceable for exactly the clients that omit the header.
"""
if not content_type:
return ""
return content_type.split(";", 1)[0].strip().lower()
(gateway/assets.py:488-498)
A zero-byte upload is refused before the raster reader ever sees it:
if not data:
# A zero-byte upload cannot be a raster. Refused here rather than
# downstream so the failure names the upload, not the reader.
raise AssetTooLargeError("the uploaded file is empty")
(gateway/assets.py:313-316)
8.5 Errors on this endpoint
413 over the size limit · 415 unsupported or absent content type · 503 the asset store is not
configured on this deployment · 400 for a malformed body. All use the §9 envelope
(docs/API_CONTRACT.md §2.5).
The handler maps each store error onto an existing taxonomy code:
except AssetTooLargeError as exc:
status, body = translate_error("oversized_image", "The uploaded file is too large.", detail=exc.detail)
except UnsupportedContentTypeError as exc:
status, body = translate_error("raster_read_error", "This file type is not accepted.", detail=exc.detail)
except AssetStoreFullError as exc:
status, body = translate_error("model_unavailable", "The upload buffer is full. Please wait and retry.", detail=exc.detail, recoverable=True)
except AssetStoreError as exc: # pragma: no cover - defensive
status, body = translate_error("input_error", "The upload could not be stored.", detail=exc.detail)
(app/space_app.py:629-657)
gateway/assets.py explains why the store's own error family is not part of core/errors.py:
"Deliberately NOT a
core.errors.SatQueryError:core/errors.pyis the taxonomy the contract publishes (API_CONTRACT.mdsection 5.2, 23 codes), and none of those codes means "this handle is unknown". Adding one would move the taxonomy, which the contract forbids the gateway from doing (docs/DEPLOYMENT_ARCHITECTURE.mdsection 2.3). The route therefore maps these onto existing codes, and says which." (gateway/assets.py:156-163)
The two layers do not map identically. The Space maps
UnsupportedContentTypeErrortoraster_read_error, whiledocs/FRONTEND_INTEGRATION.md§3.3's table namesunsupported_bandsfor the415case. The code is authoritative: the shipped handler emitsraster_read_error(app/space_app.py:638). The integration guide's table is a client-facing approximation and disagrees with the code on this one row.
8.6 When upload is disabled
The endpoint fails closed. _asset_store_available() requires both variables:
def _asset_store_available() -> bool:
"""Whether the upload endpoint is enabled.
Off by default in a deployment that has not set `SATQUERY_ASSET_DIR`, and
ON when it has -- so turning on the fourth endpoint is an explicit operator
action rather than something that starts writing to a temp directory
unbidden. `/v1/capabilities` is where a client learns which it is.
"""
import os
return bool(os.environ.get("SATQUERY_ASSET_ENABLED", "")) and bool(
os.environ.get("SATQUERY_ASSET_DIR")
)
(app/space_app.py:394-406)
if not _asset_store_available():
# A deployment that has not enabled uploads says so, with the reason
# and the switch, rather than accepting bytes it cannot keep.
status, body = translate_error(
"model_unavailable",
"Asset upload is not enabled on this deployment.",
detail=(
"Set SATQUERY_ASSET_ENABLED=1 and SATQUERY_ASSET_DIR to a "
"writable path to enable POST /v1/assets. See "
"docs/DEPLOYMENT_ARCHITECTURE.md."
),
recoverable=False,
)
(app/space_app.py:578-591)
docs/DEPLOYMENT_TOPOLOGY.md §3.3 states the same requirement in the env-var table: "Both required
for /v1/assets; fails closed (503) otherwise".
8.7 Lifetime, capacity, and the refusal-not-eviction rule
"Handles expire on a TTL and are refused on read once lapsed — a lapsed handle is rejected even if nothing has swept it, so a client never succeeds by racing a cleanup job. Capacity is bounded, and a live handle is never evicted to make room: when the store is full it refuses (
503) rather than invalidating a handle a client is about to use. A handle is single-use in practice — consuming it in/v1/analyzedoes not consume it, so the same handle may be analysed repeatedly until it expires." (docs/API_CONTRACT.md§2.5)
Expiry is checked on read:
def get(self, asset_id: str, *, now: float | None = None) -> AssetHandle:
"""Resolve a handle, or raise `UnknownAssetError`.
Expiry is evaluated here rather than trusted to a sweeper: a handle past
its deadline is unknown even if nothing has run `sweep()`. That is what
makes the TTL a guarantee instead of a housekeeping hope.
"""
(gateway/assets.py:370-376)
UnknownAssetError deliberately merges expired and never issued:
"The two are one error on purpose: a client cannot act on the difference (both mean "upload again"), and distinguishing them would report whether a handle had ever existed -- a small information leak about other clients' uploads, which matters precisely because there is no auth." (
gateway/assets.py:183-189)
Sizing defaults, and the env-var overrides:
| Knob | Default | Env var | Source |
|---|---|---|---|
| Per-file cap | 4 * 1024 * 1024 (4 MiB) |
SATQUERY_MAX_FILE_BYTES |
app/space_app.py:359; gateway/policy.py:324 |
| Capacity (handles) | 32 |
SATQUERY_ASSET_MAX_FILES |
app/space_app.py:278 |
| TTL | 900.0 s (15 min) |
SATQUERY_ASSET_TTL_S |
app/space_app.py:298 |
| Root | tempfile.gettempdir()/satquery-assets |
SATQUERY_ASSET_DIR |
app/space_app.py:310 |
| Enable switch | off | SATQUERY_ASSET_ENABLED |
app/space_app.py:404 |
The measured per-file limit is 4,194,304 bytes.
HANDOFF_NEXT_AGENT.md§4 (session workspace) records it as a hard constraint: "Per-file upload limit 4,194,304 bytes (HTTP 413 above it)." The default in code is4 * 1024 * 1024= 4,194,304, so the default and the measured value agree.
The capacity default was hardcoded until the STEP 8 audit:
"These two were hardcoded until the STEP 8 audit recorded the resulting scaling limit: a deployment with ample
SATQUERY_ASSET_DIRand a burst of concurrent users could hit the ceiling before the TTL reaped anything and answer503with disk free. That is a refusal, not corruption -- the store never evicts a live handle -- but it is a limit an operator should be able to raise without editing code." (app/space_app.py:254-261)
stats() was removed rather than left wired:
"F-11 (owner ruling 2026-09-23):
stats()lived here. It reportedlive/capacity/ttl_seconds/max_file_bytes/sweeps/capacity_refusalsfor an operator, but no production module ever called it --docs/DEPLOYMENT_ARCHITECTURE.mdsection 5 pointed an operator at an instrument the deployment did not expose." (gateway/assets.py:449-457)
9. Size and limit rules
9.1 Two caps, both enforced twice
| Cap | Default | Where declared | Enforced at |
|---|---|---|---|
| Whole-request body | 8 * 1024 * 1024 (8 MiB) |
GatewayConfig.max_body_bytes (gateway/policy.py:219) |
gateway: Content-Length check (admit step 3) and while reading (_read_body_bounded) |
| Per-file upload | 4 * 1024 * 1024 (4 MiB) |
GatewayConfig.max_file_bytes (gateway/policy.py:221) and _asset_max_file_bytes() (app/space_app.py:359) |
gateway and Space, both while reading |
The invariant that ties them together:
if self.max_file_bytes > self.max_body_bytes:
raise ValueError(
"max_file_bytes exceeds max_body_bytes; the per-file cap would "
"be unreachable and the body check would fire first"
)
(gateway/policy.py:273-277)
9.2 The Content-Length check is declarative; the read-time check is not
The admit ladder's step 3:
# 3. Body size, BEFORE the body is read. This is the check that saves
# quota and memory; everything downstream has already buffered it.
if content_length is not None and content_length > self.config.max_body_bytes:
status, body = translate_error(
"oversized_image",
"The request body is too large.",
detail=(
f"Content-Length {content_length} exceeds the "
f"{self.config.max_body_bytes} byte limit"
),
request_id=request_id,
)
return PolicyDecision.refused(request_id, status, body, headers)
(gateway/policy.py:606-618)
F-6's measurement is the reason a second, read-time check exists:
*"Measured through this stack on 2026-09-22 with
max_body_bytesat 8 MiB:Content-Length declared, 12 MiB -> 413 oversized_image, peak 0.2 MiB, 0 bytes read (the cap worked) Content-Length omitted, 12 MiB -> 502 model_unavailable, peak 13.9 MiB, 12 MiB read (the cap was SKIPPED)and allocation tracked the body size exactly with no ceiling -- 1/8/16/32/64 MiB in produced 3.0/8.1/16.0/32.0/64.0 MiB allocated. So the header check protects the common case and bounds nothing in the hostile one."* (
gateway/app.py:460-478)
The read-time cap is unconditional and shared by both layers:
async def read_body_bounded(request: Any, limit: int) -> tuple[bytes, bool]:
"""Read a request body, refusing it the moment it exceeds `limit`.
...
The `Content-Length` header is deliberately NOT consulted here, not even as a
fast path. The F-6 measurement is the reason: a declared length drew `413`
with 0.2 MiB peak, but an **omitted** one drew `502` with a 13.9 MiB peak, so
anything keyed on that header holds only for clients that tell the truth.
"""
chunks: list[bytes] = []
total = 0
async for chunk in request.stream():
total += len(chunk)
if total > limit:
return b"", True
chunks.append(chunk)
return b"".join(chunks), False
(gateway/assets.py:100-148)
The F-9 measurement on the Space side:
"Measured 2026-09-22 with the cap at 1 MiB: a 64 MiB body produced a peak allocation of 128 MiB and a 16 MiB body 32 MiB, tracking body size linearly with no ceiling, and the
413came only after everything had been held." (app/space_app.py:595-598)
9.3 The F-7 correction — one variable, and it must mean one value
SATQUERY_MAX_FILE_BYTES is read by both layers, and until F-7 each layer parsed it separately:
"Measured on four inputs (probe
probe_f7_cap_parsers.py):'abc'and'4e6'made the gateway raise at startup while the Space silently returned the 4 MiB default;'0'and'-1'were accepted by the gateway while the Space rejected them only when the first upload arrived. Neither layer was right in both directions. Both now refuse an unparsable or non-positive value, naming the variable, and a cross-layer agreement test drives the whole matrix through both real parsers." (docs/API_CONTRACT.md§2.5)
The Space's reader now refuses rather than defaulting:
try:
value = int(raw)
except ValueError:
raise ValueError(
f"SATQUERY_MAX_FILE_BYTES={raw!r} is not an integer. It is NOT "
f"defaulted, because the gateway refuses this same value at startup "
f"and silently substituting a different cap here would leave the two "
f"layers disagreeing about what 'too large' means -- the exact "
f"failure this shared variable exists to prevent."
) from None
if value <= 0:
raise ValueError(...)
(app/space_app.py:360-378)
The gateway refuses it at startup:
max_file_bytes = _int("SATQUERY_MAX_FILE_BYTES", 4 * 1024 * 1024)
if max_file_bytes <= 0:
raise ValueError(
f"SATQUERY_MAX_FILE_BYTES={max_file_bytes} is not positive. A "
f"non-positive per-file cap would make every upload fail on the "
f"Space while the gateway kept admitting it; the Space's AssetStore "
f"rejects the same value, so it is refused here to fail at startup "
f"with the variable named rather than on the first upload"
)
(gateway/policy.py:324-332)
"A deployment whose cap is malformed no longer starts at all, at either layer, instead of quietly running on a limit nobody chose." (
docs/API_CONTRACT.md§2.5)
9.4 The pixel budget is a different limit, owned by the planner
The per-file byte cap is not the image-size limit. core/planner.py owns a pixel budget of
25,000,000, and exceeding it is oversized_image (recoverable: True — retry at reduced resolution):
class OversizedImageError(InputError):
code = "oversized_image"
user_message = "The image exceeds the configured pixel budget."
# Recoverable via downscale.
(core/errors.py:147-154)
HANDOFF_NEXT_AGENT.md §8 item 6 and the release chapter 03 §25 record the value as 25,000,000; this
chapter does not restate the planner's rules (see
03 — Request Lifecycle §25).
9.5 Rate limiting — fairness, not protection
#: Per-IP request budget for `POST /v1/analyze`.
rate_limit_per_ip: int = 10
#: Window for the per-IP budget, in seconds.
rate_limit_window_s: float = 60.0
(gateway/policy.py:222-225)
The limiter is in-memory and fixed-window:
"A distributed limiter needs shared state, and the plan forbids Redis-cluster infrastructure (
docs/DEPLOYMENT_ARCHITECTURE.mdsection 6). An in-memory limiter in a single-instance Railway service loses its counters on restart -- which is correct behaviour for protecting a daily GPU budget, because the thing being protected (the Space's quota) is unaffected by a gateway restart. The limiter therefore protects the budget, not a billing invariant." (gateway/policy.py:362-373)
A denied request does not increment the counter:
if state.count >= self.limit:
retry_after = self.window_s - (now - state.window_start)
return False, 0, max(0.0, retry_after)
(gateway/policy.py:396-398)
"A denied request does not increment the counter -- otherwise a client hammering the endpoint would push its own reset further away on every rejected attempt." (
gateway/policy.py:386-389)
The client identity is best-effort, and the contract says so:
"
X-Forwarded-Foris used because the gateway sits behind Railway's proxy, but it is attacker-controlled absent a trusted proxy, so this is a best-effort guard, not a security boundary. The runbook says so." (gateway/policy.py:408-413)
and the limiter bounds request count, not request size:
"It bounds request COUNT (10/60s, measured engaging at exactly 10), not request SIZE -- ten admitted 16 MiB requests are 160 MiB of unaccounted memory. That is the F-5 distinction again: a fairness control is not a protection control." (
gateway/app.py:475-478)
Rate-limit values are not specified by the plan (docs/API_CONTRACT.md §8), so the numbers above
are this implementation's defaults, not contract constants.
10. CORS — an explicit allowlist, never *
10.1 The rule
"The gateway sets CORS explicitly to the deployed frontend origin. It does not use a wildcard. A preflight
OPTIONSis answered by the gateway, not by the HF Space." (docs/API_CONTRACT.md§7.1)
The configuration refuses to start rather than guess:
if not self.allowed_origins:
raise ValueError(
"allowed_origins must be non-empty. An empty allowlist would "
"either block every browser or, if the app fell back to '*', "
"expose the Space. Refuse to start rather than guess."
)
if "*" in self.allowed_origins:
raise ValueError(
"allowed_origins must not contain '*'. A wildcard exposes the "
"Space to any origin (docs/DEPLOYMENT_ARCHITECTURE.md 2.1)"
)
(gateway/policy.py:248-258)
10.2 A disallowed origin receives no CORS headers at all
def build_cors_headers(
origin: str | None, allowed: Sequence[str], *, request_headers: str = ""
) -> dict[str, str]:
"""Explicit CORS headers, or none.
...
A disallowed origin receives **no CORS headers at all**, which is what makes
the browser block the response. Echoing the origin back with
`Access-Control-Allow-Origin: <origin>` regardless would defeat the
allowlist entirely.
"""
if not origin or origin not in allowed:
return {}
return {
"Access-Control-Allow-Origin": origin,
"Access-Control-Allow-Methods": "GET, POST, OPTIONS",
"Access-Control-Allow-Headers": request_headers or "Content-Type, X-Request-Id",
"Access-Control-Max-Age": "600",
"Vary": "Origin",
}
(gateway/policy.py:429-454)
docs/DEPLOYMENT_TOPOLOGY.md §3.2 lists the rule among the gateway's responsibilities: "CORS
allowlist (never *)".
10.3 Preflight is answered by the gateway, never forwarded
# 1. Preflight is answered here, never forwarded. The Space has no CORS
# configuration and forwarding OPTIONS would waste a round trip.
if method == "OPTIONS":
return PolicyDecision.allowed(request_id, headers)
(gateway/policy.py:589-592)
10.4 F-2 — the response leg strips every upstream CORS header
This is a security rule, not hygiene, and the bypass was measured:
*"The bypass was measured, not theorised. With the allowlist set to
["https://app.example.com"]and the Space answering withAccess-Control-Allow-Origin: *:
GET /v1/healthfromhttps://evil.example.netreturned two values foraccess-control-allow-origin,*and the allowlisted origin;POST /v1/analyzefrom the same disallowed origin returned200withAccess-Control-Allow-Origin: *andAccess-Control-Allow-Credentials: true.Both are the failure
docs/DEPLOYMENT_ARCHITECTURE.md§2.1 exists to prevent: the allowlist is bypassed by a header the gateway never inspected."* (gateway/policy.py:705-723)
The fix is a prefix filter on the response leg:
return {
key: value
for key, value in upstream.items()
if key.lower() not in DEFAULT_HOP_BY_HOP
and key.lower() not in blocked
# F-2. A prefix rule, not an enum: RFC 6648 discourages new `Access-`
# headers, but the CORS family has grown (`-Allow-Credentials`,
# `-Expose-Headers`, `-Max-Age`, `-Allow-Methods`, `-Allow-Headers`)
# and an enum would silently miss whichever is added next.
and not key.lower().startswith("access-control-")
}
(gateway/policy.py:741-753)
and _proxy() pins the filter with an assertion rather than trusting it:
assert not any(_is_cors_header(k) for k in out_headers) or decision.headers, (
"a CORS header reached the response without a policy decision; the "
"upstream's headers are no longer filtered (see F-2)"
)
out_headers.update(decision.headers)
(gateway/app.py:592-596)
10.5 The orchestrator's own allowlist assembly
deploy/render/main.py assembles the list from three sources and re-checks for a wildcard, because
CORSMiddleware does not run GatewayConfig.__post_init__:
raw = os.environ.get("SATQUERY_ALLOWED_ORIGINS", "")
origins: list[str] = [o.strip() for o in raw.split(",") if o.strip()]
...
origins.extend(_PRODUCTION_ORIGINS)
if include_dev:
origins.extend(_DEV_ORIGINS)
if "*" in origins:
raise ValueError(
"SATQUERY_ALLOWED_ORIGINS must not contain '*'. A wildcard exposes "
"the deployment to any origin (docs/DEPLOYMENT_ARCHITECTURE.md 2.1)."
)
(deploy/render/main.py:194-208)
The production origin is hard-coded so an env-var typo cannot take the site down:
#: The production frontend origin. Listed here rather than only in the
#: environment so that a deployment which forgets `SATQUERY_ALLOWED_ORIGINS`
#: still serves the real frontend -- an empty allowlist would otherwise take the
#: live site down, which is a worse failure than the one this guards.
_PRODUCTION_ORIGINS: tuple[str, ...] = ("https://satquery.pages.dev",)
(deploy/render/main.py:145-149)
The dev origins are enumerated host:port pairs, never a regex or a suffix match:
_DEV_ORIGINS: tuple[str, ...] = tuple(
f"http://{host}:{port}"
for host in ("localhost", "127.0.0.1")
for port in ("3000", "5500", "5173", "8000", "8080")
)
(deploy/render/main.py:139-143)
"This list is deliberately EXPLICIT, never a wildcard or a suffix match. It cannot be used to reach the deployment from an arbitrary host: only a browser running on the developer's own machine can send
Origin: http://localhost:*." (deploy/render/main.py:135-138)
The measured live allowlist is the single production origin:
SATQUERY_ALLOWED_ORIGINS=https://satquery.pages.dev (docs/DEPLOYMENT_TOPOLOGY.md header note).
11. The HARD RULE — the gateway must NOT retry POST /api/infer
11.1 The rule, stated three times in the sources
"Render must not retry
POST /api/inferon its own — a retry would consume inference a second time. The client decides on retry. (Matches the gateway contract inDEPLOYMENT_ARCHITECTURE.md§2.2.)" (docs/DEPLOYMENT_TOPOLOGY.md§2)
The code says it at the one place it could be violated — the transport-failure branch:
except Exception as exc: # network-level failure
# NO RETRY. A retry on /v1/analyze would spend GPU quota twice
# (docs/DEPLOYMENT_ARCHITECTURE.md section 2.2).
(gateway/app.py:550-552)
and the frontend says it to the user:
"Never automatically retry
POST /v1/analyze. Each attempt consumes GPU quota, and on a 5-minute daily budget an auto-retry loop can exhaust the day. Retries must be an explicit user action." (docs/FRONTEND_INTEGRATION.md§6.1)
11.2 Why: the retry is not free, and it is not the gateway's call
Three independent reasons, each recorded:
- It costs inference twice. A retry on
/v1/analyzespends GPU quota a second time (gateway/app.py:551). - The gateway cannot know whether the first attempt succeeded. A transport failure is ambiguous: the upstream may have completed the work and failed to answer. Only the client holds the intent.
- The plan's boundary excludes it.
docs/DEPLOYMENT_ARCHITECTURE.md§2.2 is the "what the gateway must NOT do" list, anddocs/API_CONTRACT.md§8 records that the rate-limit values are the gateway's to choose while the retry policy is not.
11.3 What the gateway does instead
It classifies the failure and returns the envelope, with a stable request_id the client can quote:
_log.error(
"upstream transport failure for request_id=%s: %s: %s",
decision.request_id,
type(exc).__name__,
exc,
exc_info=exc,
)
status, err_body = translate_error(
"model_unavailable",
"The analysis service is not reachable.",
detail=_transport_failure_detail(exc),
recoverable=True,
request_id=decision.request_id,
)
return JSONResponse(status_code=502, content=err_body, headers=decision.headers)
(gateway/app.py:565-579)
recoverable: True is the honest signal: a retry may help, but the client decides.
11.4 The client's side of the rule
The shipped client offers no auto-retry. It surfaces which step failed and what the server said:
"HONEST FAILURE. When the backend is unreachable, the caller is told which step failed and what the server said. Nothing is synthesised to fill the gap -- no invented answer, no placeholder confidence." (
frontend/assets/js/live.js:24-26)
"
GET /v1/healthand/capabilitiesare cheap and may be polled." (docs/FRONTEND_INTEGRATION.md§6.1)
So the asymmetry is: polling the cheap endpoints is fine; retrying the costly one is a user decision.
12. The error contract
12.1 The envelope
"Every non-2xx response body has this shape:" (
docs/API_CONTRACT.md§5)
{
"error": {
"code": "pair_misaligned",
"message": "The images are not sufficiently co-registered for spatial analysis.",
"detail": "RMSE 4.21 px exceeds the 2.0 px budget",
"recoverable": false,
"request_id": "req_01H...",
"run_id": "9f2c1c0e-..."
}
}
(docs/API_CONTRACT.md §5)
"
codeis stable and comes fromcore/errors.py.messageis operator-safe (SatQueryError.user_message).detailis technical and may be absent." (docs/API_CONTRACT.md§5)
translate_error builds it:
body: dict[str, Any] = {
"error": {
"code": code if known else "satquery_error",
"message": message or "An internal error occurred.",
"detail": detail or None,
"recoverable": bool(recoverable),
"request_id": request_id,
"run_id": run_id,
}
}
return status, body
(gateway/policy.py:161-171)
Note the two or None / or "An internal error occurred." defaults: detail is emitted as null
rather than "" when absent, and message can never be empty.
12.2 The code is passed through unchanged
"The
codeis passed through unchanged.docs/DEPLOYMENT_ARCHITECTURE.mdsection 2.3 forbids the gateway from inventing or remapping codes: the taxonomy incore/errors.pyis the single source of truth, and a gateway that renamed anything would make the frontend's error handling unpredictable." (gateway/policy.py:136-139)
An unrecognised code is treated as satquery_error and mapped to 500 — never to a success
status — and the original code is preserved in detail for diagnosis:
known = code in _CODE_STATUS
status = _CODE_STATUS.get(code, 500)
if not known:
# Keep the original code visible in `detail` for diagnosis, but present
# a code the contract defines. Swallowing it entirely would hide a real
# defect behind a generic one.
detail = f"unmapped error code {code!r}" + (f"; {detail}" if detail else "")
(gateway/policy.py:153-159)
12.3 The HTTP status mapping
| Status | When | recoverable |
|---|---|---|
400 |
Malformed JSON, missing required field, or a malformed upload body | false |
404 |
The path is not an endpoint at all (routing_error) |
false |
405 |
The path exists but not for this method (routing_error). GET /v1/assets is the common case |
false |
413 |
Upload exceeds the per-file size limit | false |
415 |
Upload's content type is absent or not on the allowlist | false |
422 |
Schema violation (unknown field with extra="forbid", wrong enum value, assets empty) |
false |
429 |
Rate limited. Honours Retry-After |
true |
500 |
Unexpected internal failure | false |
503 |
A required model is absent or unavailable; or GPU quota exhausted; or the asset store is unconfigured or full |
depends |
504 |
The specialist exceeded its budget (specialist_timeout) |
true |
(docs/API_CONTRACT.md §5.1)
413 and 415 are upload-only. A 415 is the expected answer to an upload with no declared
Content-Type — the server refuses rather than guessing (docs/API_CONTRACT.md §5.1).
12.4 404 and 405 carry the same envelope — through both layers
"**
404and405carry this same envelope*, which is worth stating because they are the two statuses a proxy framework raises before any handler runs. A client should therefore not special-case them: parseerror.codeas usual."* (docs/API_CONTRACT.md§5.1)
This is F-3 on the gateway and F-12 / F-12b on the Space. Measured before the fix, direct to the Space:
GET /v1/whocares -> 404 {"detail":"Not Found"}
GET /v1/assets -> 405 {"detail":"Method Not Allowed"}
an unwrapped failure -> 500 text/plain, no envelope at all
(app/space_app.py:448-451)
Both layers register a handler for Starlette's HTTPException:
@api.exception_handler(StarletteHTTPException)
async def _contract_envelope_for_transport_errors(
request: Request, exc: StarletteHTTPException
) -> JSONResponse:
code = "routing_error" if exc.status_code < 500 else "satquery_error"
status, body = translate_error(
code,
"This endpoint does not exist."
if exc.status_code == 404
else str(exc.detail),
detail=(
f"{request.method} {request.url.path} -> {exc.status_code}; "
f"see docs/API_CONTRACT.md sections 2 and 5"
),
recoverable=False,
)
return JSONResponse(
status_code=exc.status_code,
content=body,
headers=getattr(exc, "headers", None),
)
(app/space_app.py:465-491; the gateway's copy is at gateway/app.py:302-332)
"the status is taken from the EXCEPTION, not from the code -- exactly as the gateway handler does it, so the two layers cannot drift." (
app/space_app.py:454-457)
An unhandled failure also answers with the envelope, and the client gets a fixed, operator-safe message:
@api.exception_handler(Exception)
async def _contract_envelope_for_unhandled_failures(
request: Request, exc: Exception
) -> JSONResponse:
"""F-12b: an unwrapped failure answers with the envelope, not plain text.
...
**F-15 is enforced here as well as in the handlers.** The client is told
the code and a fixed, operator-safe message; the exception's own text is
recorded SERVER-SIDE only. [...] A traceback in `detail` would disclose
internal paths and library versions to an unauthenticated caller.
"""
_log.exception(
"unhandled failure on %s %s", request.method, request.url.path
)
status, body = translate_error(
"satquery_error",
"An internal error occurred.",
detail="",
recoverable=False,
)
return JSONResponse(status_code=status, content=body)
(app/space_app.py:493-519)
12.5 A trailing slash is a 307, and it is a real footgun
*"Starlette's default
redirect_slashesbehaviour applies:GET /v1/analyze/answers307withLocation: http://<gateway-host>/v1/analyze. Two consequences a client must handle, both verified against the running app on 2026-09-22:
- The
Locationis built from the gateway's own host, not from the client's request URL, so a redirect followed naively after aPOSTmay not land where the caller expects. Do not rely on it.- A
307preserves the method and body, so aPOST /v1/analyze/will re-send the body to/v1/analyze— which is fine, but it is a second request against the rate limiter, and/v1/analyzeis aCOSTLY_ROUTE."* (docs/API_CONTRACT.md§5.1)
"Use exact paths with no trailing slash. This is also why the runbook's
SATQUERY_SPACE_URLis normalised with a trailing-slash strip." (docs/API_CONTRACT.md§5.1)
The same normalisation exists for the orchestrator's upstream URL, and it refuses rather than trims:
if self.upstream_url.endswith("/"):
raise ValueError(
"upstream_url must not have a trailing slash; a doubled slash "
"produces a 404 from the Space that looks like an outage"
)
(gateway/policy.py:239-243)
and the environment reader strips it so the operator cannot cause the failure:
upstream_url=env.get("SATQUERY_SPACE_URL", "").rstrip("/"),
(gateway/policy.py:335)
12.6 The complete code taxonomy — 23 codes
From core/errors.py. The frontend should map these to user-facing copy; the user_message field is
a safe default (docs/API_CONTRACT.md §5.2).
code |
Meaning | Suggested UX |
|---|---|---|
satquery_error |
Base class — the fallback when a more specific code does not apply | Generic failure. Treat an unexpected occurrence as a defect |
input_error |
The uploaded input could not be read | Ask the user to re-upload |
raster_read_error |
Not a readable TIFF/GeoTIFF | "This file is not a readable GeoTIFF" |
missing_crs |
No coordinate reference system | "This image has no georeferencing" |
unsupported_bands |
Band layout unsupported | Explain expected bands |
oversized_image |
Exceeds the pixel budget | Offer downsampling |
pair_incompatible |
The two images do not match | Prompt for a better pair |
pair_misaligned |
Not co-registered | Explain alignment requirement |
temporal_pair_invalid |
Two distinct acquisitions required | Ask for a second date |
routing_error |
Request could not be interpreted | Offer force_task |
unsupported_query |
No specialist supports this | Show the capability list |
invalid_request |
Inputs do not support the task | Suggest a valid task |
workflow_plan_error |
Workflow could not be planned | Retry; report if persistent |
specialist_error |
A specialist failed | Generic failure |
model_load_error |
A model could not be loaded | Defect — surface it |
model_unavailable |
Model not available in this environment | Disable the capability |
out_of_memory |
OOM; retry at reduced resolution | Suggest a smaller image |
specialist_timeout |
Processing timed out | Offer retry |
schema_validation_error |
The system produced a malformed result | Defect — always report |
coordinate_error |
Invalid spatial coordinates | Defect |
confidence_range_error |
Confidence out of range | Defect |
leakage_violation |
A data isolation rule was violated | Defect — never user-facing |
benchmark_freeze_error |
The benchmark is not frozen | Evaluation-only |
(docs/API_CONTRACT.md §5.2)
"Render
user_messageas the default and override specific codes with better copy. Do not invent a mapping fromdetail— it is not stable." (docs/API_CONTRACT.md§5.2)
The class hierarchy that produces these codes (core/errors.py):
classDiagram
class SatQueryError { code = "satquery_error" }
class InputError { code = "input_error" }
class RasterReadError { code = "raster_read_error" }
class MissingCRSError { code = "missing_crs"; recoverable default True }
class UnsupportedBandsError { code = "unsupported_bands" }
class OversizedImageError { code = "oversized_image"; recoverable default True }
class PairCompatibilityError { code = "pair_incompatible" }
class PairMisalignmentError { code = "pair_misaligned" }
class TemporalPairError { code = "temporal_pair_invalid" }
class RoutingError { code = "routing_error" }
class UnsupportedQueryError { code = "unsupported_query" }
class InvalidRequestError { code = "invalid_request" }
class WorkflowPlanError { code = "workflow_plan_error" }
class SpecialistError { code = "specialist_error" }
class ModelLoadError { code = "model_load_error" }
class ModelUnavailableError { code = "model_unavailable"; recoverable default True }
class OutOfMemoryError { code = "out_of_memory"; recoverable default True }
class SpecialistTimeoutError { code = "specialist_timeout"; recoverable default True }
class SchemaValidationError { code = "schema_validation_error" }
class CoordinateError { code = "coordinate_error" }
class ConfidenceRangeError { code = "confidence_range_error" }
class LeakageError { code = "leakage_violation" }
class BenchmarkFreezeError { code = "benchmark_freeze_error" }
SatQueryError <|-- InputError
InputError <|-- RasterReadError
InputError <|-- MissingCRSError
InputError <|-- UnsupportedBandsError
InputError <|-- OversizedImageError
SatQueryError <|-- PairCompatibilityError
PairCompatibilityError <|-- PairMisalignmentError
PairCompatibilityError <|-- TemporalPairError
SatQueryError <|-- RoutingError
RoutingError <|-- UnsupportedQueryError
SatQueryError <|-- InvalidRequestError
SatQueryError <|-- WorkflowPlanError
SatQueryError <|-- SpecialistError
SpecialistError <|-- ModelLoadError
SpecialistError <|-- ModelUnavailableError
SpecialistError <|-- OutOfMemoryError
SpecialistError <|-- SpecialistTimeoutError
SatQueryError <|-- SchemaValidationError
SchemaValidationError <|-- CoordinateError
SchemaValidationError <|-- ConfidenceRangeError
SatQueryError <|-- LeakageError
SatQueryError <|-- BenchmarkFreezeError
12.7 The recoverable flag — what it means and where it comes from
recoverable is a field on the base exception, defaulting to False:
class SatQueryError(Exception):
"""Base class for every SatQuery failure.
Attributes:
code: stable machine-readable identifier, used in traces.
user_message: text safe to show the operator.
detail: technical detail for the execution trace (never chain-of-thought).
recoverable: whether the controller may continue with a fallback.
"""
code: str = "satquery_error"
user_message: str = "An internal error occurred."
def __init__(
self,
detail: str = "",
*,
user_message: str | None = None,
recoverable: bool = False,
context: dict[str, Any] | None = None,
) -> None:
(core/errors.py:82-106)
Four subclasses override the default to True, each with a stated reason:
| Class | recoverable |
Reason |
|---|---|---|
MissingCRSError |
True |
"Degraded, not fatal: non-geospatial analysis may still be possible." (core/errors.py:135) |
OversizedImageError |
True |
"Recoverable via downscale." (core/errors.py:150) |
ModelUnavailableError |
True |
"Recoverable: the controller degrades the workflow." (core/errors.py:220) |
OutOfMemoryError |
True |
"Recoverable: retry at lower resolution." (core/errors.py:230) |
SpecialistTimeoutError |
True |
see below |
SpecialistTimeoutError carries the longest justification in the file, and it documents a defect
that was fixed:
*"
: Recoverable, perdocs/API_CONTRACT.mdsection 5.1, which maps 504 withrecoverable: true`. Two independent reasons:
docs/API_CONTRACT.mdis the frozen frontend-facing contract. A frontend that readsrecoverable: falsewill not offer a retry for the one failure the contract explicitly tells it to retry.- The plan's Failure Matrix (§57) lists Timeout with the recovery "abort specialist" and the fallback "partial result" -- i.e. the controller continues rather than failing the request. A terminal
recoverable=Falsecontradicts that.Inheriting
FalsefromSatQueryErrorwas the defect this default corrects. Note the controller currently only reuses.codefor its budget-skip trace entry (core/controller.py:464), so nothing in the pipeline constructed this class and the wrong default was never observable from the inside -- only from a client."* (core/errors.py:240-255)
The client's use of the flag is one branch:
if (!res.ok) {
const { error } = await res.json();
if (error.recoverable) { /* offer a retry affordance */ }
else { /* terminal: explain, do not offer retry */ }
}
(docs/FRONTEND_INTEGRATION.md §5)
The shipped client keeps the flag on its error object so a caller can branch on it:
err.recoverable = !!opts.recoverable;
(frontend/assets/js/live.js:160)
12.8 DEFECT_CODES — the five that mean the system is broken
#: Codes that indicate the *system* is broken, not the request. The frontend is
#: instructed to surface these rather than swallow them (`API_CONTRACT.md` 5.2).
DEFECT_CODES: frozenset[str] = frozenset(
{
"model_load_error",
"schema_validation_error",
"coordinate_error",
"confidence_range_error",
"leakage_violation",
}
)
(gateway/policy.py:95-105)
"For these, show a generic failure and capture the
request_idso it can be reported. Do not attempt to explain them to the user." (docs/FRONTEND_INTEGRATION.md§5)
12.9 Gateway-origin codes — a separate set, and why they must stay separate
There is exactly one code a client can receive that is not in the taxonomy:
code |
Meaning | Suggested UX |
|---|---|---|
rate_limited |
Gateway-origin. The proxy's per-IP rate limit refused the request; it never reached the Space. 429, and Retry-After is set |
Wait Retry-After seconds, then retry. Not a bug |
(docs/API_CONTRACT.md §5.3)
GATEWAY_ORIGIN_CODES: frozenset[str] = frozenset({"rate_limited"})
_CODE_STATUS["rate_limited"] = 429
(gateway/policy.py:120-122)
"
tests/unit/test_gateway_responsibilities.pyasserts that §5.2 andcore/errors.pyare in exact one-to-one correspondence (23 codes), andtests/unit/test_gateway_policy.pyasserts that a gateway-origin code may never shadow a taxonomy code. Both hold only if the two sets stay disjoint — sorate_limitedis documented here, beside the taxonomy rather than inside it, and the correspondence test keeps its meaning." (docs/API_CONTRACT.md§5.3)
The rule that still holds:
"the gateway never invents a code for an error that ORIGINATED in the Space. Those pass through unchanged." (
gateway/policy.py:118-119)
12.10 The status map is total, and an unknown code cannot become a 200
#: HTTP status per SatQuery error code. From `docs/API_CONTRACT.md` section 5.1.
#:
#: The mapping is TOTAL over the taxonomy in `core/errors.py`: every code either
#: appears here or falls through to 500, and `tests/unit/test_gateway_policy.py`
#: asserts that no code silently maps to the wrong class. A gateway that let an
#: unknown code produce a 200 would turn a defect into a success.
_CODE_STATUS: dict[str, int] = {
"input_error": 400,
"raster_read_error": 400,
"missing_crs": 400,
"unsupported_bands": 400,
"oversized_image": 413,
"pair_incompatible": 422,
"pair_misaligned": 422,
"temporal_pair_invalid": 422,
"routing_error": 422,
"unsupported_query": 422,
"invalid_request": 422,
"workflow_plan_error": 500,
"specialist_error": 500,
"model_load_error": 503,
"model_unavailable": 503,
"out_of_memory": 503,
"specialist_timeout": 504,
"schema_validation_error": 500,
"coordinate_error": 500,
"confidence_range_error": 500,
"leakage_violation": 500,
"benchmark_freeze_error": 500,
# The base class. A bare `satquery_error` means no specific code applied,
# which is an internal failure, not a client error.
"satquery_error": 500,
}
(gateway/policy.py:61-93)
Note a real divergence between the map and the contract table.
_CODE_STATUSmapsrouting_errorto 422, whiledocs/API_CONTRACT.md§5.1 maps404/405torouting_error. The Space's and gateway's transport handlers take the status from the exception (status_code=exc.status_code) and usetranslate_erroronly for the body, which is exactly why the divergence is invisible on that path — and it is stated here rather than left for a reader to trip over.
12.11 detail never carries a path
F-15 (owner ruling 2026-09-23): sanitize all client-facing exception messages; retain full exception details only in server-side diagnostics. The scrubber reduces absolute paths to their basename:
def scrub_paths(text: str | None) -> str | None:
"""Reduce every absolute filesystem path in `text` to its final component.
...
Exception messages in this repo routinely embed an absolute path:
`specialists/optical_sar/croma.py` raises *"CROMA requires the vendored
'use_croma.py', which is not in 'C:\\\\...\\\\empty_vendor_dir'"* and
`specialists/change/stanet.py` raises *"could not read encoder weights from
C:\\\\..."*. Those strings reach client-visible fields, and
`API_CONTRACT.md` section 7 records that v1 has **no auth**.
"""
(core/errors.py:39-68)
"URLs are left intact on purpose:
https://github.com/antofuller/CROMAappears inside one of the very messages this scrubs, and mangling it would be a worse defect than the one being repaired." (core/errors.py:65-67)
Relative paths are deliberately not matched:
"A rule broad enough to catch
artifacts/change/head.ptalso catchesand/orand the path segments of a URL, and a scrubber that mangles ordinary prose is a worse defect than the disclosure it fixes." (core/errors.py:24-27)
Two further disclosures were closed on the same ruling and are worth naming because they are the same class of bug:
- F-13.
trace.inputsechoedrequest.assetsafter handles had become paths. Measured end-to-end: "the client sent the handleasset_d243f7f85d8c2f3c02981f0af9737f01and received backC:\Users\anish\sq_scratch\...\assets\asset_d243...7f01.tif." The fix is_asset_label(a)— the basename (core/controller.py:263-277). - F-14. The same disclosure again, one record later, in the
PARSEstep'sdetail(core/controller.py:288-300).
Both now use the one _asset_label rule, because "a second copy of an existing rule is a second
thing that can drift from it."
12.12 The transport-failure classifier — a classification, not an exception dump
F-15c. The gateway's transport branch used to publish the exception's own class name and message:
#: Transport failures, mapped to a CONTRACT-level classification. Ordered, most
#: specific first. Matched on the exception's MRO class names rather than with
#: `isinstance`, so the classifier keeps working when `httpx` is absent (the
#: caller passes `httpx=None` in that case) and does not couple the client-facing
#: vocabulary to a third-party type hierarchy that can be renamed.
_TRANSPORT_FAILURES: tuple[tuple[str, str], ...] = (
("TimeoutException", "the upstream did not respond within the gateway timeout"),
("ConnectError", "the upstream could not be reached"),
("ProxyError", "the gateway's egress proxy refused the connection"),
)
_TRANSPORT_FAILURE_FALLBACK = "the upstream request failed at the transport layer"
(gateway/app.py:125-136)
"what the client needs is which kind of transport failure this was -- unreachable, timed out, or refused by the proxy -- because those imply different operator actions. What it must not receive is the exception's own text, which names the gateway's HTTP client, its internals, and potentially a proxy URL or a path. The full exception still reaches the operator through
_log.error(..., exc_info=exc)in the caller, so nothing is lost -- it is moved, not deleted." (gateway/app.py:141-149)
12.13 A non-JSON upstream body is a defect, on every status
content_type = upstream.headers.get("content-type", "")
if "application/json" not in content_type and path != "/v1/health":
status, err_body = translate_error(
"schema_validation_error",
"The analysis service returned a malformed response.",
detail=(
f"content-type={content_type!r} for {path} "
f"(upstream status {upstream.status_code})"
),
request_id=decision.request_id,
)
# 502 regardless of the upstream's own status: the fault the CLIENT can
# act on is "the gateway's upstream misbehaved", and echoing e.g. a 404
# from a reverse proxy would suggest the API path itself was wrong.
return JSONResponse(status_code=502, content=err_body, headers=decision.headers)
(gateway/app.py:616-630)
"The guard covers EVERY status, not only 2xx. It originally read
upstream.status_code < 400, which meant a non-JSON 4xx/5xx from the upstream -- a proxy error page, an HTML 502 from a load balancer, a plain-text stack trace from a misconfigured Space -- was forwarded verbatim. […] the upstream's own wording was passed through unfiltered, which is how this was found: a sandbox egress proxy returned a 502 whose body disclosedos error 10061." (gateway/app.py:600-612)
/v1/health is exempt because "a liveness probe may legitimately answer non-JSON, and it consumes no
GPU quota." (gateway/app.py:614-615)
The orchestrator applies the same rule in its own _proxy():
except Exception:
_log.error(
"non-JSON upstream response from %s (status %s)", url, resp.status_code
)
return _envelope(
"schema_validation_error",
"The inference engine returned a non-JSON response.",
f"upstream status {resp.status_code}",
status=502,
recoverable=True,
)
(deploy/render/main.py:404-414)
12.14 The orchestrator's own error codes
deploy/render/main.py defines three orchestrator-local codes, which are not in the 23-code
taxonomy because they describe the proxy, not an analysis:
| Class | code |
status | recoverable |
|---|---|---|---|
WakeTimeout |
wake_timeout |
504 |
True |
OrchestratorConfigError |
orchestrator_config_error |
500 |
False |
OrchestratorUpstreamError |
upstream_unreachable |
502 |
True |
(deploy/render/main.py:247-272)
and two more codes minted at the call sites:
code |
Where | status |
|---|---|---|
upstream_error |
deploy/render/main.py:395 (generic httpx.HTTPError) |
502 |
invalid_request |
deploy/render/main.py:480 (body was not valid JSON) |
400 |
wake_timeout's message is deliberately actionable:
class WakeTimeout(OrchestratorError):
"""The Codespace did not reach `available` within the wake timeout."""
code = "wake_timeout"
message = "The inference engine did not start in time. Please retry shortly."
status = 504
recoverable = True
(deploy/render/main.py:247-253)
and the orchestrator's envelope carries four fields, not six — it has no request_id or run_id
to offer at that layer:
body = {
"error": {
"code": code,
"message": message,
"detail": detail or None,
"recoverable": bool(recoverable),
}
}
(deploy/render/main.py:283-290)
This is a real shape difference between the two layers' envelopes. The gateway's envelope carries
request_idandrun_id(gateway/policy.py:161-170); the orchestrator's carries neither (deploy/render/main.py:283-290). A client that readserror.request_idmust tolerate its absence when the error originated at the orchestrator.
12.15 tunnel_offline — a code that exists only on the deployed backend
docs/DEPLOYMENT_TOPOLOGY.md §2 records that when the tunnel agent is absent, POST /api/infer
parks until SATQUERY_TUNNEL_TIMEOUT_S (150 s) and then returns tunnel_offline (503,
recoverable:true). That code is minted by the deployed SatQuery-Backend/main.py, which is not
in this repository (B-03). It is therefore:
IMPLEMENTED (not in this repository) — the code and its semantics are recorded in
docs/DEPLOYMENT_TOPOLOGY.md §2 and docs/FINAL_DELIVERY_TODO.md §5 (B-07), but the source line that
emits it was not read for this chapter.
13. The five entrypoint requirements
docs/DEPLOYMENT_ARCHITECTURE.md §3.3 fixes five requirements for the inference entrypoint. They are
reproduced in app/space_app.py's module docstring under the heading "THE FIVE ENTRYPOINT
REQUIREMENTS", and each is implemented in a way a reader can check.
13.1 Requirement 1 — import cheaply and without torch
"Import cheaply and without torch.
GET /v1/healthand/v1/capabilitiesmust answer on CPU with no GPU and no model load. This follows the existing project convention:app/serving.pyimportsbuild_*functions lazily inside the builders precisely so that importing the module does not pull the model stack." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
The implementation defers both the web framework and the model stack:
def get_controller() -> Any:
"""Build (once) and return the serving controller.
Requirement 2: this delegates to `app.serving.build_serving_controller()`, […]
The call is deferred to first use, not performed at import: requirement 1
says importing this module must not pull the model stack.
"""
global _CONTROLLER
if _CONTROLLER is None:
from app.serving import build_serving_controller
_CONTROLLER = build_serving_controller()
return _CONTROLLER
(app/space_app.py:168-185)
def build_space_app() -> Any:
"""...
`FastAPI` is imported HERE, inside the function, which is correct and
deliberate: requirement 1 says importing this module must not pull the web
framework, and a name used only at build time does not need to be in the
module namespace. The annotation subjects (`Request`, `Response`,
`JSONResponse`) are the opposite case and ARE module-level -- see the import
comment for why the distinction is load-bearing.
"""
from fastapi import FastAPI
...
(app/space_app.py:409-432)
The ZeroGPU decorator is applied conditionally so the module imports everywhere:
def _spaces_module() -> Any | None:
"""Import `spaces` if present.
ZeroGPU Spaces ship it; a CPU-only machine does not. Returning None rather
than raising keeps the module importable everywhere, which requirement 1 of
section 3.3 demands.
"""
try:
import spaces # type: ignore[import-not-found]
return spaces
except Exception:
return None
(app/space_app.py:129-141)
The decoration has never executed.
app/space_app.py's own implementation note is explicit: "The consequence is recorded indocs/PHASE19_FINAL_HARDENING.md: the ZeroGPU decoration has never executed here. It is specified from finding C-8 and the frozengpu_duration_*values, and that is all it is." (app/space_app.py:40-42)
13.2 Requirement 2 — reuse app.serving's composition root
"Reuse
app.serving.build_serving_controller()is the existing, tested composition root. It already wires the change head and the change-VQA head through the registry'sbuilders=override — the mechanism that avoids editingconfigs/base.yamland therefore avoids movingConfig.hash. The entrypoint must not reimplement this wiring; doing so would duplicate the F2 train/serve-skew fix." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
The single call site, with the reasoning at the call:
"""Requirement 2: this delegates to `app.serving.build_serving_controller()`,
the existing composition root that wires the change head and the change-VQA
head through the registry's `builders=` override. That override is a call-site
argument rather than config, which is what keeps `Config.hash` unchanged
(requirement 5)."""
(app/space_app.py:171-175)
13.3 Requirement 3 — degrade, do not crash, but corrupt artifacts raise ModelLoadError
"Degrade, do not crash.
app/serving.pydocuments the contract: absent artifacts degrade; corrupt artifacts raiseModelLoadError. A capability whose artifacts are absent reportsavailable: falsewith a reason, and/v1/healthreturnsstatus: "degraded". A Space that refuses to boot because an optional artifact is absent is a worse failure than one that serves a reduced capability set." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
The corrected statement of what a bare host reports:
"Corrected 2026-09-22. This requirement previously added "On a Space with no artifacts, every capability reports
available: false". That is no longer the behaviour and was never the right target:changeandchange_vqaare shipped in-repo, so they report available even on a bare host. The accurate statement is that an unservable capability is reported unavailable and still listed — never omitted from the enumeration, because a frontend cannot disable an affordance it was never told about." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
The two error classes that carry the distinction are separate types, not one type with a flag:
class ModelLoadError(SpecialistError):
code = "model_load_error"
user_message = "A required model could not be loaded."
class ModelUnavailableError(SpecialistError):
code = "model_unavailable"
user_message = "A required model is not available in this environment."
# Recoverable: the controller degrades the workflow.
(core/errors.py:212-225)
model_load_error is a DEFECT code (§12.8) and maps to 503; model_unavailable is recoverable
and also maps to 503. The distinction is recoverable, and it is carried into the UI.
13.4 Requirement 4 — never load a model for a metadata request
"Never load a model for a metadata request. Health and capabilities read artifact presence (filesystem) and configuration, not weights." (
docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
The adapter is the implementation, and its docstring names the method it must not reuse:
def describe_deployment() -> dict[str, Any]:
"""Report what this deployment can do, without loading any model.
...
It does **not** reuse `AnalystController.health()`, which its own docstring
documents as "Constructs everything" -- the opposite of what a metadata
request may do. Per the ruling, that method is retired as a public API path.
"""
from app.deployment import deployment_report
return deployment_report().as_internal()
(app/space_app.py:188-215)
"The adapter reads the spec table plus the filesystem and never calls
build()orbuild_all(). It is asserted bytest_it_does_not_use_the_constructing_controller_health, which fails if the metadata path reachesAnalysisController— the class whosehealth()constructs everything." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3, requirement 4)
13.5 Requirement 5 — honour the config hash; never merge deploy.yaml
"Honour the config hash. The entrypoint must not mutate the config, merge
deploy.yamlinto the registry, or otherwise moveConfig.hashoff78f1e3700da15aa1." (docs/DEPLOYMENT_ARCHITECTURE.md§3.3)
Three places in the code that exist because of this requirement:
The GPU duration table. change_vqa has no key of its own and reuses change:
#: ZeroGPU duration per task, from `configs/deploy.yaml` -- the FROZEN values,
#: not new guesses. `change_vqa` has no key of its own and reuses `change`,
#: because adding a key would move `Config.hash` off `78f1e3700da15aa1`
#: (`docs/DEPLOYMENT_ARCHITECTURE.md` section 3.4).
GPU_DURATIONS: dict[str, int] = {
"vqa": 20,
"caption": 20,
"grounding": 45,
"change": 30,
"optical_sar": 45,
"change_vqa": 30,
}
(app/space_app.py:105-116)
A missing duration is a programming error, not a default:
if task not in GPU_DURATIONS:
raise KeyError(
f"no gpu_duration_* is declared for {task!r}; add it to "
f"configs/deploy.yaml (which moves Config.hash) or map it to an "
f"existing task. Do not guess a duration."
)
(app/space_app.py:152-157)
The asset sizing knobs are read from the environment, not from configs/base.yaml:
"""Read from the environment rather than from `configs/base.yaml` on purpose:
adding a key there moves `Config.hash` off `78f1e3700da15aa1` and invalidates
the frozen Phase-9 benchmark. Sizing is deployment state, and the same
reasoning already governs `SATQUERY_MAX_FILE_BYTES`."""
(app/space_app.py:263-266)
The builders= override is a call-site argument, not config — which is what requirement 2's
docstring calls out as "what keeps Config.hash unchanged" (app/space_app.py:174-175).
The frozen hash itself is 78f1e3700da15aa1, and it appears in the response:
"config_hash": "78f1e3700da15aa1"
(docs/API_CONTRACT.md §2.4; docs/DEPLOYMENT_TOPOLOGY.md §5)
14. The two annotation traps — how a FastAPI app becomes silently wrong
Both traps have the same root cause and opposite diagnosability. They are documented at length in the source because both fired in this codebase, and because a future handler can re-introduce either by moving one import.
14.1 G-1 — an unresolvable parameter annotation is silently reinterpreted
The module uses from __future__ import annotations, so request: Request is a string at
runtime. FastAPI resolves it with eval(annotation, func.__globals__), so the name must be in the
module's globals:
# `Request` is bound at MODULE scope, and this is load-bearing rather than
# stylistic -- it is the G-1 defect, and it would be re-introduced here by a
# local import.
#
# This module uses `from __future__ import annotations`, so `request: Request`
# in the `/v1/assets` handler below is a *string* at runtime. FastAPI resolves
# it with `eval(annotation, func.__globals__)`, so the name must be in this
# module's globals. If `Request` were imported inside `build_space_app`, it would
# be a local of that function, `eval` would fail, and FastAPI would **not raise**
# -- it would silently reinterpret the parameter as a required *query* parameter
# named `request`, answering every upload with
# `422 {"detail":[{"loc":["query","request"]}]}` and never entering the handler.
(app/space_app.py:55-66)
The same trap is documented in gateway/app.py with the measured symptom:
*"FastAPI then does not raise -- it silently falls back to treating the parameter as a query parameter named
request. The observable consequence was that every POST to/v1/analyzeand/v1/assetsreturned422 {"detail":[{"loc":["query","request"],"msg":"Field required"}]}without ever entering the handler: the request body was never read, the gateway's own validation never ran, and the error envelope was FastAPI's
{"detail": ...}rather than the contract's{"error": {...}}. No test caught this because no test could import FastAPI when this file was written."* (gateway/app.py:63-73)
14.2 The return annotation — an unresolvable one raises
# `Response` is bound for the same reason, and its necessity was discovered the
# hard way. Every route in this module is annotated `-> JSONResponse` while
# `JSONResponse` was imported INSIDE `build_space_app`. With
# `from __future__ import annotations`, FastAPI evaluates that return annotation
# against `space_app.__globals__`, where `JSONResponse` did not exist, so
# `add_api_route` raised:
#
# pydantic.errors.PydanticUndefinedAnnotation: name 'JSONResponse' is not defined
#
# and `build_space_app()` could not be called at all. This is the SAME class of
# defect as G-1 -- a name needed by `eval` at route-registration time bound in a
# narrower scope than the annotation evaluator can see -- and it is why every
# route annotation subject in this file is now module-level.
#
# Note this one FAILS LOUDLY, where G-1 failed silently. The difference is
# whether the unresolved name is a parameter annotation (FastAPI falls back to a
# query parameter) or a return annotation (FastAPI has no fallback and raises).
(app/space_app.py:73-89)
14.3 The asymmetry, stated as the lesson
| Unresolved name is… | FastAPI behaviour | Symptom |
|---|---|---|
| a parameter annotation | falls back to a required query parameter | 422 {"detail":[{"loc":["query","request"]}]}, handler never runs, silent |
| a return annotation | no fallback; raises at add_api_route |
PydanticUndefinedAnnotation: name 'JSONResponse' is not defined; the app cannot be built, loud |
"Note the asymmetry with G-1, which is worth internalising: an unresolvable PARAMETER annotation is silently reinterpreted (FastAPI treats it as a query parameter and the handler never runs), while an unresolvable RETURN annotation raises. Same root cause, opposite diagnosability. Binding the name here fixes both and makes the difference moot." (
gateway/app.py:101-105)
14.4 What this means for a maintainer
Three rules, all derived from the above:
- Every annotation subject must be module-level.
Request,Response,JSONResponseare imported at module scope in both files with# noqa: E402and a comment (app/space_app.py:71,:90,:91;gateway/app.py:85,:86,:106). - A name used only at build time may stay inside the function.
FastAPIis imported insidebuild_space_appdeliberately, and the docstring says why the distinction is load-bearing (app/space_app.py:418-423). - A route rename is not the only way to disable rate limiting. The
is_analyzeoverride exists because a path-derived cost check can be silently wrong (§3.2) — the same class of "make the implicit explicit" repair.
15. What is NOT RUN, OPEN or BLOCKED for this topic
The style guide requires this list explicitly, and it must not be softened.
| Item | Status | Detail |
|---|---|---|
Multipart upload into /v1/analyze |
NOT IMPLEMENTED | Option B was rejected; the JSON form plus /v1/assets is the surface (docs/API_CONTRACT.md §2.4, §2.5.1) |
/v1/assets exercised against a live deployment end to end |
NOT RUN at the contract level | docs/FRONTEND_INTEGRATION.md §9: "Upload is defined but not yet exercised against a live deployment. […] no request has traversed the real gateway-to-Space path, because no egress to it exists in the build environment." The live path that has been exercised is the deployed /api/assets → tunnel → /v1/assets chain (E-03, E-05), which is a different statement from the integration suite's |
| Multipart upload through the real gateway | NOT RUN | The shipped client sends raw bytes (§8.1); the multipart description has no live evidence |
GET /v1/analyze/ trailing-slash behaviour |
VERIFIED (2026-09-22) | "verified against the running app on 2026-09-22" (docs/API_CONTRACT.md §5.1) |
| End-to-end benchmark of the API surface (latency, throughput) | NOT RUN — none exists | "Latency is not characterized. No cold-start or throughput measurement has been taken against a live Space." (docs/FRONTEND_INTEGRATION.md §9) |
| Rate-limit values as contract constants | NOT SPECIFIED BY THE PLAN | "The gateway must choose them; ask the maintainer" (docs/API_CONTRACT.md §8) |
| Streaming / progress API | NOT IN v1 | docs/API_CONTRACT.md §8 |
| Authentication | DELIBERATELY ABSENT | plan §74; docs/API_CONTRACT.md §7 |
tunnel_offline code |
IMPLEMENTED (not in this repository) | recorded in docs/DEPLOYMENT_TOPOLOGY.md §2; the emitting source is in the private SatQuery-Backend repo |
The deployment block of /api/capabilities |
STALE — known defect | claims huggingface-spaces/zerogpu on a Render+tunnel deployment (docs/FINAL_DELIVERY_TODO.md §1.7 item 6) |
gateway/app.py route wiring actually executing |
NOT RUN | "the route wiring below has never executed. Its logic is not untested -- every decision it makes lives in gateway.policy, which has 51 passing tests -- but the FastAPI plumbing […] is specified and reviewed, not run." (gateway/app.py:22-27) |
The x-satquery-transport: tunnel header as a contract field |
MEASURED, not contracted | It is a live fact (E-03) and a client read (live.js:315), but it is not in docs/API_CONTRACT.md |
model_load_error observed from a live client |
NOT RUN | The class and its 503 mapping exist; no measurement of it reaching a client was found |
| UNKNOWN — not established from the available evidence | — | the exact HTTP status the deployed tunnel returns for each failure mode; the rate-limit values in force on the live Render service; whether /v1/assets is enabled on the live Codespace (SATQUERY_ASSET_ENABLED/SATQUERY_ASSET_DIR are not in the measured live env-var list in docs/DEPLOYMENT_TOPOLOGY.md) |
15.1 A note on what "VERIFIED" means for this contract
docs/API_CONTRACT.md §8 is explicit that the contract's own verification has a boundary:
"
docs/ITEM5_INTEGRATION_SUITE_SCOPE.mdrecords what the 26-test integration suite proves (the app's boundary, in-process) and what only a live deployment can prove (reachability, cold start, memory ceilings). Read it before treating a greentests/integrationrun as evidence about a deployment — no test in this repository dials a network address, including this contract's own/v1/*examples." (docs/API_CONTRACT.md§8)
So this chapter distinguishes three grades of claim, and never blends them:
- shape — verified against the Pydantic models and the handlers (in-process);
- behaviour — verified against a running app on a named date (the
307, the404/405envelopes, the two body-cap measurements); - deployment — verified against the live Render + tunnel stack (E-03, E-05, E-11, E-14), and only for the paths those runs exercised.
16. Where the evidence lives
| Claim class | File | What it establishes |
|---|---|---|
| Request/response shapes | core/schemas.py |
AnalysisRequest (:412), ResultEnvelope (:421), HealthStatus (:430), SpecialistResult (:325), ExecutionTrace (:296) |
| Error taxonomy | core/errors.py |
23 codes, recoverable defaults, scrub_paths |
| The four endpoints | app/space_app.py |
:521 health, :549 capabilities, :555 assets, :661 analyze |
| Entrypoint requirements | app/space_app.py:17-29; docs/DEPLOYMENT_ARCHITECTURE.md §3.3 |
the five requirements |
| G-1 and the return-annotation trap | app/space_app.py:55-91; gateway/app.py:53-106 |
both traps, with the measured symptom |
Status map, DEFECT_CODES, gateway-origin codes |
gateway/policy.py |
:61-93, :95-105, :120-122 |
| CORS | gateway/policy.py:429-454, :698-753; deploy/render/main.py:139-216 |
allowlist, F-2, dev origins |
| Body caps, F-6, F-7, F-9 | gateway/policy.py:606-618; gateway/assets.py:100-148; app/space_app.py:313-378, :593-621 |
both caps and both measurements |
| Asset store | gateway/assets.py |
handle opacity (:466), TTL (:370), capacity (:328), allowlist (:318) |
/api/* mirror |
deploy/render/main.py |
route table (:22-26), handlers (:444-508) |
| Gateway allowlists | gateway/app.py:170-199 |
PROXIED_ROUTES, BLOCKED_ROUTES, COSTLY_ROUTES |
| The no-retry rule | gateway/app.py:550-552; docs/DEPLOYMENT_TOPOLOGY.md §2; docs/FRONTEND_INTEGRATION.md §6.1 |
three independent statements |
| Verified live example | docs/FINAL_DELIVERY_TODO.md §6 E-03 |
POST /api/infer {} → 422 invalid_request, x-satquery-transport: tunnel |
| Live validation | docs/FINAL_DELIVERY_TODO.md §6 E-11, E-14 |
8/8 × 3 passes, mock_nodes=0, all /api/* → onrender.com |
| Client-facing specification | docs/API_CONTRACT.md (916 lines) |
read fully; the authority for everything a client author reads |
| Integration guidance | docs/FRONTEND_INTEGRATION.md (417 lines) |
§9 "what is NOT guaranteed" |
| Active topology | docs/DEPLOYMENT_TOPOLOGY.md (248 lines) |
the /api/* route list (§3.2), the measured transport (§2) |
| Blueprint | render.yaml (25 lines) |
the declared Render env vars |
16.1 Cross-references
| For… | Read |
|---|---|
| the topology, the wake flow, the tunnel, the health payload | 02 — Deployment Topology |
| the controller's nine states, the planner, the raster contract, the events | 03 — Request Lifecycle |
| the evidence record, the confidence rules, the trace's fields | 06 — Evidence and Confidence |
| the static tier, the Analyze console, the client, the harness | 09 — Frontend |
| the health payload's every field, the trace as an observability object, the runbook | 10 — Observability and Operations |