Instructions to use thundercode/SatQuery with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use thundercode/SatQuery with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Download docs/architecture/06-evidence-and-confidence.md from thundercode/SatQuery: direct link, hf CLI and curl.
- Browser
- Download file 108 kB
-
https://huggingface.co/thundercode/SatQuery/resolve/main/docs/architecture/06-evidence-and-confidence.md
- Command line
-
hf download hf://thundercode/SatQuery/docs/architecture/06-evidence-and-confidence.md
-
curl -L -o 06-evidence-and-confidence.md https://huggingface.co/thundercode/SatQuery/resolve/main/docs/architecture/06-evidence-and-confidence.md
06 β Evidence and Confidence
Parent: Architecture hub Β· Sibling chapters: 01 System overview Β· 05 Specialists Β· 07 Configuration freeze
Primary sources read for this chapter (all under C:/Users/anish/satquery-ai/):
| Source | What it establishes here |
|---|---|
evidence/engine.py (714 lines) |
the aggregation pipeline, purity contract, identity/sort/claim keys, _deduplicate, _record_agreement, _renumber, evidence_type_for, evidence_from_box/region/geospatial, confidence_for, evidence_digest |
evidence/confidence.py (434 lines) |
the honesty rule, TemperatureCalibration, _is_effective, _logit/_sigmoid, _EPS, calibrate, calibrate_result, load_calibration and its ordered candidate search |
core/schemas.py (462 lines) |
Evidence, EvidenceType, ConfidenceBreakdown, ExecutionTrace, TraceStep, ModelRef, SpecialistResult, CoordinateSystem, and every validator |
configs/base.yaml (Β§evidence, Β§confidence) |
evidence.max_items: 32, confidence.temperature_scaling: true, confidence.calibration_file: calibration_v001.json |
artifacts/calibration_v001.json |
the measured fitted temperature and the reliability diagram |
frontend/assets/js/core.js |
SQ.EVENT_NAMES β the eight execution events |
frontend/assets/js/mission.js |
markState() and the .trace__fill width formula |
docs/ARCHITECTURE_FREEZE.md Β§1, Β§3, Β§5 |
the layer verbs and the frozen non-negotiables |
docs/PHASE13_EVIDENCE_ENGINE.md |
the phase record for this package (77 tests) |
docs/API_CONTRACT.md Β§2.4, Β§3.3, Β§4 |
the client-facing shape of evidence and confidence |
docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5 |
the F-16 owner ruling on artifact refs |
docs/STEP7_BACKEND_CHAIN_REPORT.md Β§7, Β§8 |
the measured calibration and evidence integration results |
1. Where this subsystem sits
docs/ARCHITECTURE_FREEZE.md Β§5 gives every layer exactly one verb:
Router understands; policy engine decides; specialists compute; VLM explains; evidence engine proves.
Two consecutive stages implement that last verb:
flowchart LR
SPEC["Specialists<br/>compute"] -->|"SpecialistResult.evidence"| EV["Evidence engine<br/>aggregate()<br/>collect β dedup β sort β<br/>annotate β renumber β cap"]
EV -->|"EvidenceCollection"| CONF["Confidence<br/>confidence_for() / calibrate()"]
CONF -->|"ConfidenceBreakdown"| RES["ResultEnvelope<br/>+ ExecutionTrace"]
style EV fill:#1f6feb22,stroke:#1f6feb
style CONF fill:#1f6feb22,stroke:#1f6feb
The freeze Β§1 pipeline diagram places them adjacently:
+---------+---------+---------+
| | | |
VQA/CAP GROUNDING CHANGE OPTICAL-SAR
| | | |
+---------+---------+---------+
|
v
Evidence engine <- this chapter, part A
|
v
Confidence calibration <- this chapter, part B
|
v
Result normaliser -> JSON + trace + PDF
Division of labour, stated precisely. evidence/engine.py does not re-derive any
specialist's claim. The module docstring is explicit (evidence/engine.py:10-14):
This module is the "proves" step. Specialists each emit
Evidencefor what they computed (seeGroundingSpecialist._build_evidence). This engine does not re-derive any of that. Its job is aggregation:collect across specialists -> order -> deduplicate -> renumber -> bound
Status of this subsystem.
| Component | Status | Basis |
|---|---|---|
| Evidence aggregation | IMPLEMENTED + VERIFIED |
evidence/engine.py; 77 tests in tests/unit/test_evidence_engine.py (docs/PHASE13_EVIDENCE_ENGINE.md Β§7) |
| Dedup / ordering / renumber / cap | IMPLEMENTED + VERIFIED |
pinned by named tests (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.2βΒ§2.5) |
| Confidence honesty rule (uncalibrated pass-through) | IMPLEMENTED + VERIFIED |
evidence/confidence.py:296-311 |
| Temperature-scaling transform | IMPLEMENTED + VERIFIED |
monotonicity + endpoint-safety tests (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3) |
| Calibration artifact fitted | MEASURED |
artifacts/calibration_v001.json |
| Calibration improves ECE | REJECTED β it does not |
ECE 0.013755 β 0.014929 (worse); see Β§7 |
| Evidence engine wired into the controller | IMPLEMENTED (Phase 14) |
docs/PHASE13_EVIDENCE_ENGINE.md Β§8 recorded it as "not yet wired"; the live run path emits the eight events (Β§9) |
| Artifact rendering / retrieval | OPEN β deliberately null in v1 |
F-16 ruling, Β§3.2 |
Part A β The evidence schema
2. Evidence β the canonical record
Evidence is defined once, in core/schemas.py:216-253. Every specialist emits it; nothing
invents a second shape (docs/ARCHITECTURE_FREEZE.md Β§3: "Single schema for every
specialist. No specialist invents its own shape.").
class Evidence(BaseModel):
model_config = ConfigDict(extra="forbid")
evidence_id: str = Field(default_factory=lambda: _new_id("ev"))
type: EvidenceType
source_specialist: str
coordinate_system: CoordinateSystem | None = None
coordinates: list[float] | None = None
score: float | None = Field(default=None, ge=0.0, le=1.0)
artifact_ref: str | None = Field(
default=None,
description=(
"Reference to an externally retrievable artifact. NEVER a "
"filesystem path (F-16, owner ruling 2026-09-23): v1 exposes no "
"artifact-serving endpoint, so this is null unless a deployment "
"supplies a client-fetchable reference. An artifact may still be "
"written server-side where configured; being written is not the "
"same as being retrievable."
),
)
payload: dict[str, Any] = Field(default_factory=dict)
2.1 Every field, exhaustively
| Field | Type | Required | Default | Meaning | Note |
|---|---|---|---|---|---|
evidence_id |
str |
no | _new_id("ev") β ev_<12 hex> |
the item's identity | run-local, engine-assigned after aggregation (Β§5.8); a uuid before that |
type |
EvidenceType |
yes | β | the kind of proof | closed 11-member enum, Β§3 |
source_specialist |
str |
yes | β | which specialist made this claim | a single string, not a set β this is load-bearing, Β§5.7 |
coordinate_system |
CoordinateSystem | None |
no | None |
the frame the coordinates live in | mandatory in practice for spatial types β see the validator, Β§4 |
coordinates |
list[float] | None |
no | None |
the geometry | a flat list; a box is [x1, y1, x2, y2] |
score |
float | None |
no | None |
reliability, bounded | ge=0.0, le=1.0 β Pydantic rejects out-of-range |
artifact_ref |
str | None |
no | None |
a retrievable artifact reference | always null in v1 β Β§2.3 |
payload |
dict[str, Any] |
no | {} |
observations that support the claim without being it | where corroborated_by is written, Β§5.7 |
2.2 What Evidence deliberately does not have
These absences are load-bearing; each one is why a downstream design decision exists.
| Absent field | Why it is absent | Consequence |
|---|---|---|
any ordering field (index, rank, order) |
the engine owns order; a specialist must not pre-empt it | order is defined by _sort_key, Β§5.6 |
value |
the field is called score |
docs/PHASE19_FINAL_HARDENING.md Β§3.6 records that the docs once said value and that was one of nine validated defects |
source |
the field is called source_specialist |
same defect class; a frontend using source renders nothing |
summary |
not in the schema | same defect class |
contributing_specialists: list[str] |
proposed and NOT adopted | see Β§5.7 and docs/PHASE13_EVIDENCE_ENGINE.md Β§5 |
label at top level |
labels live in payload |
evidence_from_box writes payload["label"] |
a retrievable boolean |
retrieval is not a v1 capability | artifact_ref being null is the signal, Β§2.3 |
model_config = ConfigDict(extra="forbid") is what makes all of the above enforceable: a
response carrying value instead of score is a 422, not a silently-ignored key. The
conformance test asserts this directly β docs/STEP7_BACKEND_CHAIN_REPORT.md Β§D records
"extra="forbid" on all eight client-facing models; GeoMetadata is the one documented
extra="allow" exception."
2.3 artifact_ref is permanently null in v1 β owner ruling F-16
This is the single most mis-documented field in the system, and the ruling is recorded in
three places: the field's own description (core/schemas.py:225-235), the client contract
(docs/API_CONTRACT.md Β§2.4 "Artifact refs β null in v1, and why"), and the audit
(docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5).
The ruling (F-16, owner, 2026-09-23): never expose filesystem paths.
docs/API_CONTRACT.md:586-602 states the contract in full:
Every
artifact_refandchange_mapin a v1 response isnull. This is a deliberate contract, not a missing value.F-16 (owner ruling 2026-09-23): never expose filesystem paths. The specialists do render their artifacts β the change map and the optical/SAR views are written server-side β but their location is an operator fact, not a client-facing one. A response that carried the server's path would disclose the deployment's directory layout to an unauthenticated caller, and nothing the frontend can do requires it.
No
artifact://URI is fabricated in its place. v1 has no artifact-serving endpoint, so a URI would be a promise the service cannot keep β strictly worse thannull, because the frontend would build a link that 404s.
The docstring's own phrasing is the crispest statement of the principle:
being written is not the same as being retrievable.
flowchart TB
subgraph Server["Server side (operator facts)"]
R["specialist renders change map / views"]
W["file written where configured<br/>change.artifact_dir"]
R --> W
end
subgraph Client["Client-facing contract (v1)"]
N["artifact_ref = null"]
P["payload statistics<br/>total_change_pixels<br/>n_components_kept<br/>threshold"]
WARN["warnings[] entry:<br/>artifact NOT retrievable"]
end
W -.->|"deliberately NOT exposed"| X["filesystem path<br/>(never sent)"]
R --> N
R --> P
R --> WARN
style X fill:#f8514922,stroke:#f85149
style N fill:#3fb95022,stroke:#3fb950
What replaces the ref (docs/API_CONTRACT.md:604-610):
| Removed | Replaced by |
|---|---|
change_map path |
null, plus the change statistics in the CHANGE_MAP evidence's payload (total_change_pixels, n_components_kept, threshold) |
view artifact_ref path |
null, plus payload.rendered / payload.retrievable / payload.retrieval |
| β | an explicit warnings[] entry saying the artifact is NOT retrievable |
Two live carriers, and the half-implementation that followed. The audit records that the
ruling was initially applied to only one carrier
(docs/DEPLOYMENT_ARCHITECTURE.md:361):
Two live carriers β
changeandcromaβ and both had to be brought into line: fixing only thecromacarrier left the ruling half implemented.
docs/STATUS.md:21 names the lesson: "The finding worth carrying forward: the F-16 ruling
was HALF IMPLEMENTED, and the suite was green."
F-16c β the consequence the fix introduced, measured and left open. The ruling removed a
proxy and thereby changed a degraded signal. The SpecialistResult validator used to read
(core/schemas.py:362-372, retained as a comment):
if self.task is Task.CHANGE and self.change_map is None and not self.regions:
... "change analysis produced no spatial output" ... degraded = True
change_map was a proxy for "a map was produced". With the ref permanently null, the
clause collapsed to not regions, and a successful no-change analysis began reporting
degraded: true. The clause was therefore removed under ruling F-16c
(core/schemas.py:362-395). The comment left in its place is worth quoting because it states
the reasoning better than a summary could:
That conflated two different things, and the conflation WAS the defect.
degradedmeans "the analysis could not be fully performed": no trained detector, or co-registration too poor to support a spatial claim. Both are set by the specialist itself (specialists/change/specialist.py:345and:365) and neither is this validator's to invent. "No change was detected" is a NORMAL, successful outcome.
And, on why no replacement clause was added:
No replacement clause is added, deliberately. A narrower "empty answer => degraded" rule was tried and backed out: it is not what the ruling asked for, it invented a semantic the specialist already owns, and it made a pre-existing, unrelated fixture (
test_change_with_regions_is_not_degraded, a result with regions and no answer) fail. A fix that forces edits to tests it has nothing to do with is signalling over-reach, not diligence. CHANGE is the one task whosedegradedflag is now set entirely by its specialist.
Status: F-16 CLOSED (both carriers conform). F-16c RESOLVED by clause removal β the
narrower replacement rule is REJECTED. mask_ref is not an exemption: the audit
corrected an earlier working note that had listed it alongside artifact_ref
(docs/DEPLOYMENT_ARCHITECTURE.md:868-871).
3. EvidenceType β the closed 11-member vocabulary
class EvidenceType(str, Enum):
IMAGE_CROP = "image_crop"
TILE = "tile"
BOUNDING_BOX = "bounding_box"
MASK = "mask"
CHANGE_MAP = "change_map"
OPTICAL_VIEW = "optical_view"
SAR_VIEW = "sar_view"
JOINT_FEATURE_REGION = "joint_feature_region"
STATISTIC = "statistic"
GEOLOCATION = "geolocation"
AVAILABILITY_MASK = "availability_mask" # C-1: modality trust evidence
core/schemas.py:65-76. All eleven members, with what each one asserts:
| # | Member | Wire value | Asserts | Typically emitted by |
|---|---|---|---|---|
| 1 | IMAGE_CROP |
image_crop |
a rectangular raster crop exists | preprocessing / tiling |
| 2 | TILE |
tile |
one tile of the tiling policy was examined | preprocessing/tiling.py |
| 3 | BOUNDING_BOX |
bounding_box |
a rectangle locates a referent | grounding, change |
| 4 | MASK |
mask |
a per-pixel region exists | grounding, change |
| 5 | CHANGE_MAP |
change_map |
a bitemporal difference map exists | change specialist |
| 6 | OPTICAL_VIEW |
optical_view |
the optical rendering of an optical/SAR pair | optical-SAR specialist |
| 7 | SAR_VIEW |
sar_view |
the SAR rendering of an optical/SAR pair | optical-SAR specialist |
| 8 | JOINT_FEATURE_REGION |
joint_feature_region |
a region defined in CROMA's joint embedding space | optical-SAR specialist |
| 9 | STATISTIC |
statistic |
a measured scalar or count, with no geometry | VQA, caption, unsupported, and any geometry-less region |
| 10 | GEOLOCATION |
geolocation |
the raster is georeferenced and where it sits | geospatial layer |
| 11 | AVAILABILITY_MASK |
availability_mask |
which modality channels were actually available | sensor adapter (C-1) |
3.1 availability_mask β the member the freeze prose omits
docs/ARCHITECTURE_FREEZE.md Β§3's prose list of evidence types stops at ten members and does
not name availability_mask. The enum member is nevertheless real, and the engine says
so explicitly (evidence/engine.py:441-447):
The members are exactly those of
core.schemas.EvidenceType. The freeze's section 3 prose list omitsavailability_mask; that member is real (C-1: modality trust evidence) and is included here because a missing member would otherwise force a wrong fallback.
Its existence follows from finding C-1: the channel-availability mask is a first-class fusion
input and is never a CROMA input (core/schemas.py:7; docs/PHASE0_CONTRACT_VALIDATION.md
Β§1.2). Because the mask is a first-class input, it is also a first-class observation β
"these twelve optical channels were present and these two SAR channels were not" is a fact a
consumer can audit. docs/API_CONTRACT.md:579 lists all eleven values for the frontend.
3.2 Which types are spatial
The engine's own validator names the spatial set (core/schemas.py:240-247) β this is the
authoritative list, not a prose paraphrase:
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
| Type | Spatial? | Why |
|---|---|---|
BOUNDING_BOX |
yes | a box is meaningless without a frame |
MASK |
yes | pixels need a frame |
CHANGE_MAP |
yes | a map is a raster |
TILE |
yes | a tile is a rectangle of a raster |
IMAGE_CROP |
yes | a crop is a rectangle of a raster |
JOINT_FEATURE_REGION |
yes | a region in a feature grid still needs a frame |
STATISTIC |
no | a scalar has no geometry |
GEOLOCATION |
no (by this validator) | geo bounds are already self-describing via payload["crs"]; evidence_from_geospatial sets CoordinateSystem.GEO when bounds are given |
OPTICAL_VIEW |
no (by this validator) | a whole-frame view is not a sub-region claim |
SAR_VIEW |
no (by this validator) | as above |
AVAILABILITY_MASK |
no (by this validator) | a channel-presence vector is not spatial |
4. The validator that refuses coordinates without a frame
@model_validator(mode="after")
def _spatial_needs_crs(self) -> "Evidence":
spatial = {
EvidenceType.BOUNDING_BOX,
EvidenceType.MASK,
EvidenceType.CHANGE_MAP,
EvidenceType.TILE,
EvidenceType.IMAGE_CROP,
EvidenceType.JOINT_FEATURE_REGION,
}
if self.type in spatial and self.coordinates and self.coordinate_system is None:
raise ValueError(
f"evidence type '{self.type.value}' carries coordinates "
"but no coordinate_system"
)
return self
core/schemas.py:238-253. Three things about this are deliberate:
It is a
model_validator(mode="after"), not afield_validator. The rule is cross-field β it depends ontype,coordinatesandcoordinate_systemtogether, so it cannot be expressed on any single field.The trigger is
self.coordinates, notis not None. An empty list is falsy and does not trip the rule; a non-empty list does. A spatial item with no coordinates at all is legal (aMASKwhose geometry lives only in amask_ref).It raises rather than defaulting. Defaulting to
normalized_0_1would silently mislabel pixel or geo coordinates as normalized, which is the failure mode the finding exists to prevent.docs/ARCHITECTURE_FREEZE.mdΒ§3 puts the requirement bluntly:Every spatial object carries
coordinate_systemβ{normalized_0_1, pixel, geo}.
docs/ARCHITECTURE_FREEZE.md Β§2.3 states the origin: internal box coordinates are
normalised 0β1, always with an explicit coordinate_system field (finding C-5).
4.1 CoordinateSystem β three values, exact spellings
class CoordinateSystem(str, Enum):
"""Never omit this. A bare box is meaningless without it. (C-5)"""
NORMALIZED_0_1 = "normalized_0_1"
PIXEL = "pixel"
GEO = "geo"
core/schemas.py:57-62.
| Member | Wire value | Frame |
|---|---|---|
NORMALIZED_0_1 |
normalized_0_1 |
fractions of the frame, [0, 1] |
PIXEL |
pixel |
integer-ish pixel indices |
GEO |
geo |
a projected/geographic CRS named in payload["crs"] |
Documentation hazard, measured.
docs/PHASE19_FINAL_HARDENING.mdΒ§3.6 records that the docs once spelled thesenormalized/geographic. The real values arenormalized_0_1/geo. "A frontend using the documented names would send values the server rejects with 422."docs/STEP7_BACKEND_CHAIN_REPORT.mdΒ§D confirms the live enum is exactly{normalized_0_1, pixel, geo}.
Evidence's coordinate_system is CoordinateSystem | None β the field is optional in
type but effectively mandatory in practice, because the validator above enforces it for
every spatial type that carries coordinates.
Part B β The aggregation pipeline
5. EvidenceEngine.aggregate() β collect β dedup β sort β annotate β renumber β cap
5.1 Why the module exists: three problems a specialist cannot solve
The module docstring enumerates them (evidence/engine.py:18-36); each is real and none is
fixable inside a single specialist:
| # | Problem | Why a specialist cannot fix it |
|---|---|---|
| 1 | Stable identity | Evidence.evidence_id defaults to a random uuid β fine within one result, useless once four specialists' evidence is merged into one trace. "Nothing can be cited." |
| 2 | Order | Evidence has no ordering field, so the collection's order follows specialist completion order β the same inputs produce different JSON on different runs. "Pure functions must not do that, and this whole system's reproducibility argument rests on it." |
| 3 | Duplication | the VQA specialist emits a STATISTIC carrying its answer, and that answer travels in SpecialistResult.answer; two specialists that georeference the same asset emit the same GEOLOCATION. "Neither is wrong, and neither knows about the other." |
5.2 The pipeline, in one place
flowchart TB
A["results: SpecialistResult<br/>or Sequence[SpecialistResult]<br/>or evidence: Iterable[Evidence]"] --> B["collect<br/>extend(result.evidence)"]
B --> C["sources = sorted({source_specialist})<br/><i>recorded BEFORE the cap</i>"]
C --> D["_deduplicate<br/>identity = _identity_key<br/>payloads MERGED on collision"]
D --> E["sorted(key=_sort_key)<br/>type β Β· specialist β Β·<br/>score β Β· coordinates β"]
E --> F["_record_agreement<br/>_claim_key groups β<br/>payload['corroborated_by']"]
F --> G["annotated[:max_items]<br/>dropped_over_limit = total β len(capped)"]
G --> H["_renumber β evidence_001β¦"]
H --> I["EvidenceCollection"]
style D fill:#1f6feb22,stroke:#1f6feb
style F fill:#1f6feb22,stroke:#1f6feb
The implementation, verbatim (evidence/engine.py:324-350):
def _canonicalise(self, raw: Iterable[Evidence]) -> EvidenceCollection:
"""Dedup -> sort -> renumber -> cap. The whole pipeline, in one place."""
items = list(raw)
sources = sorted({item.source_specialist for item in items})
deduped, dropped_duplicates = self._deduplicate(items)
ordered = sorted(deduped, key=_sort_key)
# Two specialists can make the same claim, and `deduplicate` keeps both
# (different `source_specialist` -> different key). Record the agreement
# now that ordering is fixed, so the annotation is part of the pure
# pipeline rather than a post-hoc edit a caller might forget.
annotated = self._record_agreement(ordered)
total = len(annotated)
capped = annotated[: self.max_items]
dropped_over_limit = total - len(capped)
return EvidenceCollection(
items=[
self._renumber(item, index) for index, item in enumerate(capped, start=1)
],
sources=sources,
dropped_duplicates=dropped_duplicates,
dropped_over_limit=dropped_over_limit,
total_before_limit=total,
truncated=dropped_over_limit > 0,
)
Note the ordering choice: annotate comes after sort but before cap. Annotating before
sorting would be equivalent (the annotation is a payload key, which is not in _sort_key),
but the docstring's stated reason is that the annotation must be inside the pure pipeline
rather than "a post-hoc edit a caller might forget".
5.3 The entry point and its argument rules
def aggregate(
self,
results: SpecialistResult | Sequence[SpecialistResult] | None = None,
*,
evidence: Iterable[Evidence] | None = None,
) -> EvidenceCollection:
evidence/engine.py:280-320. Two argument invariants, both raising:
| Condition | Behaviour | Reason given in the docstring |
|---|---|---|
both results and evidence are None |
ValueError("aggregate() requires results= or evidence=") |
there is nothing to aggregate |
| both are supplied | ValueError("aggregate() accepts results= or evidence=, not both") |
ambiguity about precedence |
| empty list | returns an empty EvidenceCollection |
"Never raises on empty input." |
Only result.evidence is read. The docstring explains why reading both forms would
double-count (evidence/engine.py:292-297):
Specialists state their evidence twice β in
SpecialistResult.evidenceand viaSpecialist.produce_evidence, which currently delegates to the same list β so onlyresult.evidenceis read. Reading both would double-count every item and inflate the duplicate count.
5.4 The purity contract
"""THE PURITY CONTRACT
-------------------
`aggregate` is pure and deterministic:
* source results are never mutated -- `model_copy` is used to rebuild the
collection rather than editing `Evidence.evidence_id` in place;
* ids are assigned from the sorted position, not from input order;
* dedup is order-insensitive by construction (the key is built from the
sorted specialist list);
* no clock, no RNG, no I/O.
Same inputs -> byte-identical output. `tests/unit/test_evidence_engine.py`
pins this.
"""
evidence/engine.py:38-50.
| Property | Mechanism | Pinned by |
|---|---|---|
| source non-mutation | model_copy(update={...}) at every write β _renumber (:430), payload merge (:385), agreement (:424) |
"the original keeps its uuid and the aggregated copy is a different object" (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.1) |
| ids from sorted position | enumerate(capped, start=1) after sorted(...) |
Β§5.8 |
| order-insensitive dedup | key content-only; _record_agreement sorts names |
Β§5.5, Β§5.7 |
| no clock / RNG / I/O | the module imports only hashlib, collections.abc, dataclasses, typing |
evidence/engine.py:80-102 |
The dedup merge rebuilds rather than mutates, and the comment says why
(evidence/engine.py:379-382):
Survivor's own keys win; the loser only fills gaps. Mutating the survivor in place would be fine here because
mergedholds the same object, but rebuilding keeps the "never mutate a source result" contract true even if the caller kept a reference.
5.5 Deduplication β identity is defined, not assumed
The module docstring (evidence/engine.py:52-68) states the principle:
Evidence.evidence_idis a uuid default, so it cannot be part of an identity key β two structurally identical items from two runs would never dedup. The key is therefore the content of the observation.
def _identity_key(item: Evidence) -> tuple[Any, ...]:
coords = (
tuple(_round(c) for c in item.coordinates)
if item.coordinates is not None
else None
)
return (
item.type.value,
item.source_specialist,
item.coordinate_system.value if item.coordinate_system else None,
coords,
_round(item.score),
)
evidence/engine.py:125-142.
| Key component | Included? | Rationale |
|---|---|---|
type |
yes | a box and a mask are different claims even at the same coordinates |
source_specialist |
yes | keeps two specialists' identical claims as two items (Β§5.7) |
coordinate_system |
yes | the same four numbers in pixel and in geo are different claims |
rounded coordinates |
yes | disagreeing geometry β two claims; suppressing it would be a silent contradiction |
rounded score |
yes | different reliability is a different observation |
evidence_id |
NO | "including it would mean dedup never fires in production" |
payload |
NO | "a detail of the claim, not a different claim" |
artifact_ref |
NO | "two renderings of one claim are still one claim" |
The docstring on the exclusions is the clearest statement of intent
(evidence/engine.py:60-68):
payloadandartifact_refare deliberately EXCLUDED. Two items that agree on the same geolocation, one carrying acrsin its payload and one not, have made the same claim about the world; deduplicating them is correct, and the surviving item's payload is merged with the discarded one's so thecrsis not lost. Conversely two items that share a type and score but disagree on coordinates are two different claims and are both kept β which matters, because suppressing a spatial disagreement would be exactly the silent contradiction the freeze forbids.
The rounding constant.
#: Rounding applied before an evidence item is turned into a dedup key. Six
#: decimals on normalized coordinates is ~1e-6 of the frame, far below any
#: meaningful spatial difference, but enough to absorb float noise from two
#: specialists rounding the same value two different ways.
_KEY_PRECISION = 6
def _round(value: float | None) -> float | None:
return None if value is None else round(float(value), _KEY_PRECISION)
evidence/engine.py:114-122. _round is None-safe: a None coordinate stays None and
does not become 0.0 in the identity key (unlike the sort key, Β§5.6, where a missing
coordinate becomes 0.0 because a total order needs a value).
Payloads are MERGED on collision.
dropped += 1
# Survivor's own keys win; the loser only fills gaps.
payload = {**item.payload, **existing.payload}
if payload != existing.payload:
updated = existing.model_copy(update={"payload": payload})
seen[key] = updated
merged[merged.index(existing)] = updated
evidence/engine.py:378-387. The spread order is {**loser, **survivor} β because Python
dict unpacking is last-wins, existing.payload (the survivor's) overrides the loser's on
key collision. This is the "survivor's own keys win; the loser only fills gaps" rule
expressed directly in the merge order.
| Behaviour | Result |
|---|---|
survivor had crs, loser did not |
survivor's crs survives |
loser had crs, survivor did not |
crs is filled in from the loser β the information is not lost |
both had crs with different values |
survivor's wins; the loser's value is dropped |
| no payload difference | no model_copy is performed (the if payload != existing.payload guard) |
docs/PHASE13_EVIDENCE_ENGINE.md Β§2.3 names the test that pins this:
test_deduplication_ignores_payload_differences_and_merges_them.
What dedup does not merge. Because the key includes source_specialist, "only
same-specialist repeats are merged here; two specialists agreeing is handled by
_record_agreement, which preserves the second specialist's identity rather than discarding
it" (evidence/engine.py:356-361).
5.6 The total sort order β and why each level exists
def _sort_key(item: Evidence) -> tuple[Any, ...]:
if item.score is None:
score_key: tuple[int, float] = (1, 0.0)
else:
score_key = (0, -float(item.score))
coords = (
tuple(_round(c) or 0.0 for c in item.coordinates)
if item.coordinates is not None
else ()
)
return (item.type.value, item.source_specialist, score_key, coords)
evidence/engine.py:165-190.
| # | Key | Direction | Why it exists (verbatim from the docstring) |
|---|---|---|---|
| 1 | type |
ascending | "groups like with like, so a reader sees all the boxes together" |
| 2 | source_specialist |
ascending | "stable and meaningful, unlike the uuid" |
| 3 | score |
descending | "within a type, the strongest claim leads" β implemented as -score so the tuple stays uniformly ascending |
| 4 | coordinates |
ascending | "the tie-breaker that makes the order total. Without it two items of the same type, specialist and score would sort by Python's stable-sort insertion order, which reintroduces input-order dependence." |
docs/PHASE13_EVIDENCE_ENGINE.md Β§2.2 restates key 4 and names the pinning test:
Key 4 is not decoration. Without it, two items of the same type, specialist and score would sort by Python's stable-sort insertion order β which reintroduces exactly the input-order dependence the sort exists to remove. Pinned by
test_equal_scores_order_deterministically_by_coordinates.
Unscored items sort last. "Items with no score sort AFTER items with a score: a missing
score is not evidence of strength." (evidence/engine.py:178-179). The mechanism is a
(0, β¦) / (1, β¦) discriminator in score_key: a scored item leads with 0, an unscored
item leads with 1, so every scored item precedes every unscored item within the same
(type, specialist) group.
flowchart LR
subgraph G1["type=bounding_box, specialist=grounding"]
direction TB
A1["score 0.91"] --> A2["score 0.62"] --> A3["score None<br/><i>unscored last</i>"]
end
subgraph G2["type=statistic, specialist=vqa"]
direction TB
B1["score 0.74"]
end
G1 --> G2
style A3 fill:#d2992222,stroke:#d29922
5.7 _record_agreement β corroboration, and why nothing is ever removed
The problem. source_specialist is a single str, not a set (the schema is frozen), so
two specialists making the same claim are two different items under _identity_key and
both survive. The docstring calls this "the correct conservative behaviour"
(evidence/engine.py:70-77):
That is the correct conservative behaviour: collapsing them would require the engine to pick a winner, and the engine has no basis for that. What it does instead is keep both and record the disagreement-free agreement in the collection's
sourcesand in each item'spayloadvia_record_agreement.
The claim key β deliberately distinct from the identity key.
def _claim_key(item: Evidence) -> tuple[Any, ...]:
"""Identity of the *claim*, ignoring which specialist made it.
Used only to detect corroboration. Deliberately distinct from
`_identity_key`: that one answers "is this the same item", this one answers
"are these two specialists saying the same thing about the world".
"""
coords = (
tuple(_round(c) for c in item.coordinates)
if item.coordinates is not None
else None
)
return (
item.type.value,
item.coordinate_system.value if item.coordinate_system else None,
coords,
_round(item.score),
)
evidence/engine.py:145-162. The only difference from _identity_key is the absent
source_specialist component.
flowchart TB
I["_identity_key =<br/>(type, specialist, crs, coords, score)"] --> IQ{"same item?"}
C["_claim_key =<br/>(type, crs, coords, score)"] --> CQ{"same claim about the world?"}
IQ -->|"yes β merge payloads"| DEDUP["_deduplicate"]
CQ -->|"yes, β₯2 distinct specialists β annotate"| AGR["_record_agreement"]
style I fill:#1f6feb22,stroke:#1f6feb
style C fill:#3fb95022,stroke:#3fb950
The annotation.
AGREEMENT_KEY = "corroborated_by"
...
out = list(items)
for indexes in groups.values():
if len(indexes) < 2:
continue
names = sorted({items[i].source_specialist for i in indexes})
if len(names) < 2:
# Same specialist restating itself across the collection. The
# identity key would normally have merged these, so reaching
# here means the payloads differed; it is not corroboration.
continue
for i in indexes:
item = out[i]
others = [n for n in names if n != item.source_specialist]
payload = {**item.payload, EvidenceCollection.AGREEMENT_KEY: others}
out[i] = item.model_copy(update={"payload": payload})
return out
evidence/engine.py:391-425.
Why the payload, and not a schema field. The constant's own comment states it
(evidence/engine.py:234-238):
Payload key under which
_record_agreementstashes the specialists that made the identical claim. A single string field cannot hold a set, so the agreement is recorded in the payload β which is exactly what the payload is for: observations that support the evidence but are not the claim.
The contributing_specialists: list[str] field that would have made this first-class was
proposed and rejected β docs/PHASE13_EVIDENCE_ENGINE.md Β§5 records the full
ARCHITECTURE CHANGE entry. The three reasons:
| # | Reason |
|---|---|
| 1 | Adding a field is a frozen-contract change, and the freeze rule requires evidence that the current design is insufficient. No such evidence exists: identity comes from evidence_id; provenance is already served by EvidenceCollection.sources + Evidence.source_specialist; corroboration is recorded in payload["corroborated_by"]. |
| 2 | Because source_specialist is a single string, two specialists making the identical claim are two items and both survive. Collapsing them would require the engine to pick a winner, and the engine has no basis for that. |
| 3 | Reversing this later is cheap: the engine already computes _claim_key(), which groups items by claim ignoring the specialist. Adopting the field would mean emitting one item per group instead of N. |
Decision: proposal REJECTED; core/schemas.py unchanged. Adopted interim =
source_specialist (single) + EvidenceCollection.sources (run-level provenance) +
payload["corroborated_by"] (per-claim corroboration). "Revisit only if a downstream
consumer needs corroboration as a first-class, queryable field."
Nothing is ever removed. The docstring is unambiguous (evidence/engine.py:401-403):
Nothing is removed: dropping a member would throw away a specialist's attribution, and the freeze does not permit silently discarding a specialist's evidence.
The two guards inside the loop are worth naming because they are the two ways a non-corroboration could masquerade as one:
| Guard | Condition | Why |
|---|---|---|
| group size | len(indexes) < 2 β skip |
a lone item is not agreement |
| distinct specialists | len(names) < 2 β skip |
the same specialist restating itself is not corroboration; if this is reached, the payloads differed (otherwise _identity_key would have merged them) |
corroborated_by is always the other specialists (others = [n for n in names if n != item.source_specialist]), so an item never lists itself.
Excluded from the digest. "The corroboration annotation in the payload IS excluded,
because it is a derived observation about the collection, not part of the claim."
(evidence/engine.py:692-693) β consistent with payload being excluded from
_identity_key.
5.8 Renumbering to evidence_001β¦
ID_PREFIX = "evidence"
@staticmethod
def _renumber(item: Evidence, index: int) -> Evidence:
"""Give one item its canonical id, leaving the source untouched."""
return item.model_copy(update={"evidence_id": f"{ID_PREFIX}_{index:03d}"})
evidence/engine.py:104-107, :427-430.
| Property | Value | Reason |
|---|---|---|
| format | evidence_001 β¦ |
zero-padded to three digits |
| why three digits | "so lexical sort matches numeric sort up to 999 items -- well past the evidence.max_items bound of 32" (:104-107) |
|
| assignment basis | sorted position, not input order | enumerate(capped, start=1) after sorted |
| scope | run-local | "Ids restart at 001 for each aggregate call β they are run-local citations, not global identities." (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.4) |
| source untouched | model_copy |
the source keeps its ev_<uuid> |
Why this matters at all (evidence/engine.py:21-25):
Evidence.evidence_iddefaults to a random uuid, which is fine within one result but useless when the controller merges four specialists' evidence into the trace of a single run. Nothing can be cited. The engine renumbers toevidence_001,evidence_002, β¦ so a downstream artefact (report, UI, audit) can reference one item deterministically.
5.9 The cap, and loss accounting that is never silent
DEFAULT_MAX_ITEMS = 32
evidence/engine.py:109-112, mirroring configs/base.yaml:
evidence:
max_items: 32
coordinate_system_default: normalized_0_1
EvidenceEngine.__init__ refuses a non-positive cap (evidence/engine.py:267-276):
if max_items < 1:
raise ValueError(f"max_items must be >= 1, got {max_items}")
a cap of zero would silently discard every piece of evidence, which is a policy decision this engine has no business making on its own. (
evidence/engine.py:258-261)
The accounting fields, all five:
| Field | Type | Meaning | Non-zero means |
|---|---|---|---|
sources |
list[str] |
sorted specialist names that contributed anything, BEFORE the cap | provenance β "Provenance is a property of the run, not of the surviving items." |
dropped_duplicates |
int |
items merged into an existing claim | normal and healthy β two specialists agreed |
dropped_over_limit |
int |
items the cap discarded | a warning β a specialist's evidence did not survive |
total_before_limit |
int |
deduplicated count prior to capping | the denominator for the two drop counts |
truncated |
bool |
dropped_over_limit > 0 |
a boolean summary so a trace can branch without recomputing |
evidence/engine.py:196-217; restated in docs/PHASE13_EVIDENCE_ENGINE.md Β§2.5.
Losses are counted, never silent.
And on why sources is recorded pre-cap:
sourcesis recorded pre-cap deliberately: provenance is a property of the run, not of the surviving items. A specialist whose evidence did not fit under the cap still ran, and the trace must still say so. (docs/PHASE13_EVIDENCE_ENGINE.mdΒ§2.5)
This is the same principle as core/config.py's refusal to silently truncate and
evidence/confidence.py's refusal to fabricate a number: a bounded output must declare
what the bound cost.
5.10 EvidenceCollection β the returned container
@dataclass(frozen=True)
class EvidenceCollection:
items: list[Evidence] = field(default_factory=list)
sources: list[str] = field(default_factory=list)
dropped_duplicates: int = 0
dropped_over_limit: int = 0
total_before_limit: int = 0
truncated: bool = False
evidence/engine.py:193-250. It is frozen=True β a dataclass decorator that blocks
attribute rebinding on the collection, complementing the model_copy discipline that keeps
the individual Evidence items unmutated.
| Method | Returns | Purpose |
|---|---|---|
__len__ |
int |
len(items) |
__iter__ |
iterator | iterate the items directly |
ids() |
list[str] |
[item.evidence_id for item in self.items] |
by_type(t) |
list[Evidence] |
filter by EvidenceType (identity comparison: item.type is evidence_type) |
by_specialist(name) |
list[Evidence] |
filter by source_specialist |
summary() |
dict[str, Any] |
the observable trace facts |
AGREEMENT_KEY |
"corroborated_by" |
the payload key for corroboration |
summary() β the trace-facing view.
def summary(self) -> dict[str, Any]:
"""Observable trace facts. No chain-of-thought, no interpretation."""
return {
"returned": len(self.items),
"total_before_limit": self.total_before_limit,
"dropped_duplicates": self.dropped_duplicates,
"dropped_over_limit": self.dropped_over_limit,
"truncated": self.truncated,
"sources": list(self.sources),
"types": sorted({item.type.value for item in self.items}),
}
evidence/engine.py:240-250. Note types is a sorted set of the type values actually
present in the surviving items β so a consumer can see which evidence kinds survived the cap
without walking the list. The docstring's first line is a constraint, not a description: "No
chain-of-thought, no interpretation." This mirrors docs/ARCHITECTURE_FREEZE.md Β§5's "Every
result carries an observable execution trace. No chain-of-thought."
5.11 evidence_digest() β reproducibility as an assertion
def evidence_digest(collection: EvidenceCollection | Sequence[Evidence]) -> str:
items = (
collection.items
if isinstance(collection, EvidenceCollection)
else list(collection)
)
hasher = hashlib.sha256()
for item in items:
hasher.update(repr(_identity_key(item)).encode("utf-8"))
hasher.update(b"\n")
return hasher.hexdigest()
evidence/engine.py:680-704.
| Design choice | Consequence |
|---|---|
built from _identity_key, in order |
changes when the claims change |
not from evidence_id |
does not change when a specialist restates the same claim under a new uuid |
not from payload |
payload is a detail of the claim, not the claim |
repr(...) + b"\n" separator |
unambiguous framing; no two distinct sequences can hash the same |
| corroboration annotation excluded | "it is a derived observation about the collection, not part of the claim" |
The docstring states the contract it enables (evidence/engine.py:687-690):
That is what makes it useful as a reproducibility assertion:
aggregate(inputs_a) is reproducible iff digest(a) == digest(b)against the same input, even across processes where the uuid defaults differ.
5.12 The canonical vocabulary: evidence_type_for
@staticmethod
def evidence_type_for(result: SpecialistResult) -> EvidenceType:
from core.schemas import Task
mapping = {
Task.GROUNDING: EvidenceType.BOUNDING_BOX,
Task.CHANGE: EvidenceType.CHANGE_MAP,
Task.OPTICAL_SAR: EvidenceType.JOINT_FEATURE_REGION,
Task.VQA: EvidenceType.STATISTIC,
Task.CAPTION: EvidenceType.STATISTIC,
Task.UNSUPPORTED: EvidenceType.STATISTIC,
}
return mapping.get(result.task, EvidenceType.STATISTIC)
evidence/engine.py:434-458.
This is the ONE place that knows the mapping from specialist task to evidence vocabulary, so no specialist needs to branch on it. Note it returns the primary type; a specialist's own
produce_evidenceemits its real per-artefact types, and those are preserved byaggregate.
| Task | Primary EvidenceType |
Why |
|---|---|---|
grounding |
BOUNDING_BOX |
the task is localisation |
change |
CHANGE_MAP |
the task is a change map |
optical_sar |
JOINT_FEATURE_REGION |
the claim lives in CROMA's joint embedding space |
vqa |
STATISTIC |
an answer is a measured scalar, not geometry |
caption |
STATISTIC |
as above |
unsupported |
STATISTIC |
there is no claim to localise |
| (fallback) | STATISTIC |
.get(..., STATISTIC) β an unknown task degrades to a non-spatial claim rather than a wrong spatial one |
change_vqa is absent from the mapping. Task.CHANGE_VQA exists in the enum
(core/schemas.py:46) but has no row, so it falls through to STATISTIC β which is correct:
change-VQA returns a short answer, not a map. The docstring's note that this returns the
primary type matters here: a change-VQA result may still carry CHANGE_MAP evidence
emitted by the detector it shares, and aggregate preserves that.
5.13 The three evidence_from_* constructors
These exist so the controller does not hand-build Evidence objects β the module docstring
calls them "Convenience for the controller when it needs evidence for a box that a specialist
emitted but did not itself describe."
evidence_from_box
def evidence_from_box(self, box, *, source_specialist, payload=None, asset_ref=None) -> Evidence:
return Evidence(
type=EvidenceType.BOUNDING_BOX,
source_specialist=source_specialist,
coordinate_system=box.coordinate_system,
coordinates=[box.x1, box.y1, box.x2, box.y2],
score=box.score,
artifact_ref=asset_ref,
payload={"label": box.label, **(payload or {})},
)
evidence/engine.py:462-485.
The
coordinate_systemis copied from the box, never assumed βcore.schemasrequires it andEvidencerejects spatial coordinates without one.
The coordinates are flattened to [x1, y1, x2, y2] β matching the flat-Box shape that
docs/PHASE19_FINAL_HARDENING.md Β§3.6 records as a corrected documentation defect.
evidence_from_region β a three-way branch
evidence/engine.py:487-540. A region may carry a mask, a box, both, or neither, and each
case maps to a different evidence type:
flowchart TB
R["Region"] --> M{"region.mask_ref?"}
M -->|yes| MASK["EvidenceType.MASK<br/>artifact_ref = asset_ref or mask_ref<br/>coords = box if box else None"]
M -->|no| B{"region.box?"}
B -->|yes| BOX["EvidenceType.BOUNDING_BOX<br/>coords from box<br/>score = region.score or box.score"]
B -->|no| STAT["EvidenceType.STATISTIC<br/>no coordinates<br/>score = region.score"]
style MASK fill:#1f6feb22,stroke:#1f6feb
style BOX fill:#1f6feb22,stroke:#1f6feb
style STAT fill:#d2992222,stroke:#d29922
A region carrying a mask reference becomes
MASKevidence; one with only a box becomesBOUNDING_BOXevidence. A region with neither is reported as aSTATISTICrather than a spatial claim, because there is no geometry to prove. (evidence/engine.py:496-500)
Two details worth naming:
- The
MASKbranch'scoordinatesareNonewhenregion.box is Noneβ a mask without a bounding box is legal, and the validator (Β§4) only fires when coordinates are present. - The
BOUNDING_BOXbranch readscoordinate_systemfromregion.box.coordinate_system(the box's own frame), not fromregion.coordinate_system. The frame of the geometry is the frame of the geometry.
evidence_from_geospatial β may return None
def evidence_from_geospatial(self, geo, *, source_specialist, bounds=None, score=None) -> Evidence | None:
if not geo.has_crs and not geo.crs:
return None
return Evidence(
type=EvidenceType.GEOLOCATION,
source_specialist=source_specialist,
coordinate_system=CoordinateSystem.GEO if bounds else None,
coordinates=list(bounds) if bounds else None,
score=score,
payload={
"crs": geo.crs,
"bounds": list(geo.bounds) if geo.bounds else None,
"width": geo.width,
"height": geo.height,
"is_georeferenced": geo.is_georeferenced,
},
)
evidence/engine.py:542-572.
Returns
Nonewhen there is no CRS to report. Emitting aGEOLOCATIONitem without a CRS would assert a placement the data does not support, which is the same failure mode the grounding degenerate-box guard exists to prevent.
This is the honesty rule applied to geometry: a claim that cannot be supported is not
emitted, rather than emitted with a placeholder. Note the guard is not geo.has_crs and not geo.crs β either signal is sufficient, so a raster that reports a CRS string without the
has_crs flag still produces evidence.
6. confidence_for() β calibrating through the engine's artifact
def confidence_for(
self,
result: SpecialistResult | ConfidenceBreakdown,
*,
extra_components: dict[str, float] | None = None,
degraded: bool | None = None,
degradation_reason: str | None = None,
) -> ConfidenceBreakdown:
evidence/engine.py:576-637. It accepts either form so "a caller that already has the pieces
does not rebuild a SpecialistResult to use it."
The degradation rule β the specialist's verdict WINS.
The specialist's own degradation verdict WINS unless explicitly overridden: it knows things the engine does not (a zero-shot fallback, a failed input-quality gate), and overwriting
degraded=Falsehere would launder a degraded result into a confident-looking one. (evidence/engine.py:589-592)
| Situation | Result |
|---|---|
degraded=None, specialist said degraded=True |
stays True |
degraded=None, specialist said degraded=False |
stays False |
degraded=True explicitly |
True, and if the specialist had not flagged it, the reason is sourced from result.warnings[0] β "so the reason is a fact rather than a restatement of the boolean" (:618-622) |
degraded=True and no reason and no warnings |
falls back to the literal string "aggregate degraded" (:628-629) |
degradation_reason explicitly passed |
that string is used verbatim |
calibrate_result() (evidence/engine.py:639-643) is the narrower entry point:
def calibrate_result(self, result: SpecialistResult) -> ConfidenceBreakdown:
"""Calibrate a result's existing breakdown, preserving its provenance."""
return calibrate_result(result.confidence, self.calibration)
from_config() (evidence/engine.py:647-665):
@classmethod
def from_config(cls, config, *, calibration=None, base_dir=None) -> "EvidenceEngine":
return cls(
max_items=int(config.get("evidence.max_items", DEFAULT_MAX_ITEMS)),
calibration=calibration,
)
calibrationis explicit rather than auto-loaded: whether a fitted artifact exists is an operational fact the caller may know better than the config file does, and silently loading one would make the engine's behaviour depend on filesystem state.
The base_dir parameter is accepted for signature symmetry with
evidence.confidence.load_calibration but is not used by this classmethod β the engine
does not resolve the artifact itself.
Module-level shorthand (evidence/engine.py:668-677):
def aggregate_evidence(results, *, max_items=DEFAULT_MAX_ITEMS, calibration=None) -> EvidenceCollection:
return EvidenceEngine(max_items=max_items, calibration=calibration).aggregate(results)
Part C β The confidence system
7. The honesty rule
evidence/confidence.py:12-30 is the reason the module exists at all. It is worth quoting in
full because it is the single most important design statement in this subsystem:
THE HONESTY RULE (the reason this module exists at all)
A calibration that claims to be fitted when it is not is a FALSE CLAIM OF RELIABILITY β strictly worse than reporting nothing, because a downstream consumer will trust it. So:
no fitted artifact -> pass the raw score through unchanged, set `method="uncalibrated"`, leave `calibrated=None`
calibrated=Noneis what makes it honest:ConfidenceBreakdown.valuethen returnsraw, and any consumer that wants to distinguish "we calibrated this" from "we did not" can readmethodorcalibrated is None. We never fill in a plausible-looking number to make a schema field look complete.
7.1 The three states
docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1 tabulates the decision:
| State | calibrated |
method |
value |
|---|---|---|---|
| no fitted artifact | None |
"uncalibrated" |
falls back to raw |
fitted artifact, T β 1 |
mapped value | "temperature_scaling" |
calibrated |
fitted artifact, T = 1 |
None |
"uncalibrated" |
falls back to raw |
The method strings are module constants (evidence/confidence.py:73-78):
METHOD_TEMPERATURE = "temperature_scaling"
METHOD_UNCALIBRATED = "uncalibrated"
Method string reported whenever no mapping was applied. Matches the literal the existing specialists already emit, so trace consumers see one vocabulary.
That last clause is a real constraint: VLMSpecialist._confidence_for and
GroundingSpecialist._confidence_for already emit method="uncalibrated", calibrated=None
independently. The engine generalises that behaviour rather than introducing a second
vocabulary (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1).
7.2 ConfidenceBreakdown β the schema
class ConfidenceBreakdown(BaseModel):
"""Never an LLM utterance. Always measurable signals (plan section 26)."""
model_config = ConfigDict(extra="forbid")
raw: float = Field(ge=0.0, le=1.0)
calibrated: float | None = Field(default=None, ge=0.0, le=1.0)
method: str = "uncalibrated"
components: dict[str, float] = Field(default_factory=dict)
degraded: bool = False
degradation_reason: str | None = None
@property
def value(self) -> float:
return self.calibrated if self.calibrated is not None else self.raw
core/schemas.py:259-273.
| Field | Type | Note |
|---|---|---|
raw |
float, [0,1] |
the specialist's hand-weighted score β not a probability |
calibrated |
float | None, [0,1] |
None means no mapping was applied β the honesty signal |
method |
str, default "uncalibrated" |
one of the two constants |
components |
dict[str, float] |
diagnostic signals; docs/API_CONTRACT.md:682 β "Diagnostic only β do not compute a confidence from it" |
degraded |
bool |
"Whether this confidence should be trusted" |
degradation_reason |
str | None |
why, when degraded |
value is a Python @property and is NOT serialised. This is the trap
docs/API_CONTRACT.md:571 warns about:
result.confidence.value| NOT a JSON field. It is a Python@propertyonConfidenceBreakdownand is not serialised (verified:model_dump()yields onlycalibrated, components, degradation_reason, degraded, method, raw). To get the number the user should see, readcalibratedif it is non-null, otherwiseraw.
docs/STEP7_BACKEND_CHAIN_REPORT.md Β§D verifies this against the real model: "confidence.value
is not serialised β the six-key set matches the contract's parsed claim exactly."
7.3 The four rules the frontend must follow
docs/API_CONTRACT.md:686-694:
- Display
calibratedwhen it is notnull; otherwise displayraw. - Display
methodnext to the value.temperature_scalingmeans a fitted correction was applied;uncalibratedmeans it was not. - Never present a confidence as a percentage without its method. A raw 0.99 and a calibrated 0.99 do not mean the same thing.
- When
degradedistrue, showdegradation_reason. Confidence that is degraded is not a quality signal.
8. Temperature scaling β the derivation
8.1 Why the transform is needed at all
evidence/confidence.py:5-9:
Specialists produce a raw reliability score from signals they can actually point at (
max_objectness, input-quality autocorrelation,used_head, β¦). That raw score is not a probability: it is a hand-weighted sum, and a hand-weighted sum is not calibrated by construction. This module is the one place that maps raw -> calibrated.
The signals named in docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1 are max_objectness,
top_mean_objectness, score_contrast, used_head, and input-quality autocorrelation. The
plan (Β§26) names per-task signals too: for grounding, *"box confidence + text/image similarity
- augmentation consistency"*; for change, "mean pixel probability + component stability + registration quality"; for optical-SAR, "fusion classifier margin + optical confidence + SAR confidence + cross-modal agreement".
8.2 The formula
logit(z) = log(z / (1 - z))
calibrated = sigmoid(logit(z) / T)
evidence/confidence.py:39-47:
The standard formulation maps a logit through
sigmoid(logit / T). Our raw scores are already in (0, 1), so they are read as probabilities and converted to log-odds first.
def apply(self, raw: float) -> float:
"""Map a raw score in [0, 1] through the fitted temperature."""
return _clamp01(_sigmoid(_logit(float(raw)) / self.temperature))
evidence/confidence.py:165-167. Three operations in order: _logit β divide by T β
_sigmoid β _clamp01.
T |
Effect | Verified |
|---|---|---|
T > 1 |
softens β pulls scores toward the middle | docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3 |
T < 1 |
sharpens β pushes scores away from the middle | docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3; measured: calibrate(0.87, T=0.9772732) β 0.8749186077809417 |
T = 1 |
identity | handled as uncalibrated, Β§8.3 |
The transform is monotonic, so calibration respaces scores without reordering them β
pinned by test_temperature_scaling_is_monotonic (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3).
8.3 _is_effective β why a "fitted" T of exactly 1.0 is reported as uncalibrated
#: How close to 1.0 a fitted temperature must land before it is treated as
#: "does nothing". Bit-for-bit 1.0 is the honest test; this tolerance absorbs
#: float round-trip through JSON without accepting a real transform.
_IDENTITY_TOLERANCE = 1e-6
def _is_effective(temperature: float) -> bool:
"""Does this temperature actually transform anything?
A temperature within `_IDENTITY_TOLERANCE` of 1.0 is the identity map. It is
treated as "not fitted" so the honest report wins: an artifact that carries
T=1.0 has learned nothing, and reporting `method="temperature_scaling"` for
it would assert a correction that was never made.
"""
return abs(temperature - 1.0) > _IDENTITY_TOLERANCE
evidence/confidence.py:90-93, :118-126.
Verified empirically (read-only check against the shipped module):
| Input | _is_effective |
Result |
|---|---|---|
0.9772731820958189 (the fitted T) |
True |
a real transform is applied |
1.0 |
False |
reported as uncalibrated |
Why the tolerance and not != 1.0. The comment gives the reason: "Bit-for-bit 1.0 is the
honest test; this tolerance absorbs float round-trip through JSON without accepting a real
transform." A JSON round-trip of 1.0 can land at 1.0000000000000002; an exact != 1.0
test would then accept a transform that does nothing. 1e-6 is wide enough to absorb the
round-trip and far too narrow to swallow any real fitted temperature.
The identity branch records why (evidence/confidence.py:299-303):
if calibration is not None:
# A real artifact was supplied but it is the identity map. Say so
# rather than dropping the information on the floor.
components["calibration_identity"] = 1.0
components.update(calibration.components)
Measured components for T = 1.0:
{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}.
docs/PHASE13_EVIDENCE_ENGINE.md Β§3.2 frames this as the same failure mode as Β§3.1: "This is
a small case the brief did not name explicitly; it is the same false-claim failure mode as
Β§3.1 and is handled by the same principle."
8.4 _EPS β endpoint clamping so a hard 0.0 / 1.0 cannot become NaN
#: Scores are clamped to this closed interval before the log-odds transform.
#: `logit(0)` and `logit(1)` are infinite; clamping at the boundary keeps the
#: mapping finite and keeps a hard 0.0 / 1.0 from becoming a NaN.
_EPS = 1e-6
def _logit(p: float) -> float:
"""Log-odds of a probability, with the endpoints clamped to stay finite."""
clamped = min(1.0 - _EPS, max(_EPS, p))
return math.log(clamped / (1.0 - clamped))
evidence/confidence.py:80-83, :108-111.
Input z |
Unclamped logit(z) |
With _EPS clamp |
Result |
|---|---|---|---|
0.0 |
log(0) = -inf |
logit(1e-6) β -13.8155 |
finite |
1.0 |
log(1/0) = +inf |
logit(1 - 1e-6) β +13.8155 |
finite |
0.5 |
0.0 |
0.0 |
unchanged |
Pinned by test_endpoint_scores_do_not_produce_nan (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3).
8.5 _sigmoid β numerically stable, and why the branch exists
def _sigmoid(x: float) -> float:
"""Numerically stable logistic function.
The naive form overflows for large negative `x`; the branch below keeps
`exp`'s argument non-positive so the result is finite for every input.
"""
if x >= 0.0:
return 1.0 / (1.0 + math.exp(-x))
z = math.exp(x)
return z / (1.0 + z)
evidence/confidence.py:96-105. For x >= 0 the exp argument is -x <= 0; for x < 0 the
exp argument is x < 0. In both branches math.exp receives a non-positive argument, so
it can never overflow. The two forms are algebraically identical
(1/(1+e^-x) β‘ e^x/(1+e^x)) but only one of them is safe on each side of zero.
def _clamp01(x: float) -> float:
return min(1.0, max(0.0, x))
evidence/confidence.py:114-115. The final guard: sigmoid returns (0, 1) mathematically,
but float rounding can produce exactly 0.0 or 1.0, and _clamp01 keeps the value inside
the schema's ge=0.0, le=1.0 bound so Pydantic never rejects an output of this module.
8.6 TemperatureCalibration β the artifact object
@dataclass(frozen=True)
class TemperatureCalibration:
temperature: float
fitted_on: str | None = None
artifact: str | None = None
n_samples: int | None = None
evidence/confidence.py:129-150.
| Field | Type | Meaning (docstring) |
|---|---|---|
temperature |
float |
"the fitted scalar. T > 1 softens β¦, T < 1 sharpens. Validated on construction." |
fitted_on |
str | None |
"free-text provenance, e.g. "valid" or a split hash. Carried into the breakdown components so a result can be traced back to the artifact that shaped it." |
artifact |
str | None |
"path/identifier of the source file, for the same reason." |
n_samples |
int | None |
"how many validation samples the fit used, when known." |
Range validation (evidence/confidence.py:85-88, :152-159):
_MIN_TEMPERATURE = 1e-3
_MAX_TEMPERATURE = 1e3
def __post_init__(self) -> None:
t = float(self.temperature)
if not math.isfinite(t) or not _MIN_TEMPERATURE <= t <= _MAX_TEMPERATURE:
raise ValueError(
f"temperature must be finite and in "
f"[{_MIN_TEMPERATURE}, {_MAX_TEMPERATURE}], got {self.temperature!r}"
)
object.__setattr__(self, "temperature", t)
A temperature of zero or below is not a calibration; it is a division error. (
evidence/confidence.py:85-88)
The check is math.isfinite(t) and the range β so NaN and Β±inf are both rejected, and
object.__setattr__ is used because the dataclass is frozen=True.
Two methods:
| Method | Returns | Purpose |
|---|---|---|
is_effective() |
bool |
_is_effective(self.temperature) |
apply(raw) |
float |
_clamp01(_sigmoid(_logit(raw) / T)) |
components (property) |
dict[str, float] |
{"temperature": T} plus {"calibration_samples": n} when known |
describe() |
dict[str, Any] |
{method, temperature, effective, fitted_on, artifact, n_samples} |
from_dict() β accepts two shapes (evidence/confidence.py:187-219):
body: Any = payload
if isinstance(payload.get("temperature_scaling"), dict):
body = payload["temperature_scaling"]
elif isinstance(payload.get("calibration"), dict):
body = payload["calibration"]
if not isinstance(body, dict) or "temperature" not in body:
raise ValueError(
"calibration artifact has no 'temperature' field "
f"(source: {source or '<dict>'})"
)
Accepts either the flat form (
{"temperature": 1.4, ...}) or a nested{"temperature_scaling": {...}}form, because the artifact format is owned bytraining/calibration/and this loader should not be the thing that decides it.
fitted_on reads body.get("fitted_on") or body.get("split"); artifact prefers the
source path over body.get("artifact"); n_samples is coerced to int when present.
The shipped artifact uses the nested form β artifacts/calibration_v001.json has a
top-level "temperature_scaling": {"fitted_on": "Val", "n_samples": 16441, "temperature": 0.9772731820958189} block. This is exactly why from_dict accepts it.
from_json() β required semantics (evidence/confidence.py:221-252):
| Condition | required=False (default) |
required=True |
|---|---|---|
| file absent | returns None |
raises FileNotFoundError |
| file present, valid JSON | returns the object | returns the object |
| file present, malformed JSON | raises ValueError |
raises ValueError |
ValueError: the file exists and is malformed (never swallowed: a corrupt artifact silently treated as "no artifact" would hide an operational defect).
This is the same asymmetry as the evidence cap and the config validation: absent is a legitimate state; corrupt is a defect and must be loud.
8.7 calibrate() β the single entry point
def calibrate(
raw: float,
calibration: TemperatureCalibration | None,
*,
extra_components: dict[str, float] | None = None,
degraded: bool = False,
degradation_reason: str | None = None,
) -> ConfidenceBreakdown:
evidence/confidence.py:255-323. This is the single entry point used by
EvidenceEngine.confidence_for.
Non-finite input RAISES (evidence/confidence.py:289-292):
value = float(raw)
if not math.isfinite(value):
raise ValueError(f"raw confidence must be finite, got {raw!r}")
clamped = _clamp01(value)
Refusing is deliberate: a non-finite score is a bug upstream, and clamping it to 0.5 would invent a measurement. (
evidence/confidence.py:284-287)
Note the asymmetry with _logit's endpoint clamping: an infinite input is refused, but a
finite 0.0 or 1.0 is accepted and clamped inside the transform. The distinction is
that 0.0 is a real measurement (a specialist is certain it found nothing) while NaN is
not a measurement at all.
The honest path (evidence/confidence.py:296-311):
if calibration is None or not calibration.is_effective():
components["calibrated_applied"] = 0.0
if calibration is not None:
components["calibration_identity"] = 1.0
components.update(calibration.components)
return ConfidenceBreakdown(
raw=clamped,
calibrated=None,
method=METHOD_UNCALIBRATED,
components=components,
degraded=degraded,
degradation_reason=degradation_reason,
)
The fitted path (evidence/confidence.py:313-323):
components["calibrated_applied"] = 1.0
components.update(calibration.components)
return ConfidenceBreakdown(
raw=clamped,
calibrated=calibration.apply(clamped),
method=METHOD_TEMPERATURE,
components=components,
degraded=degraded,
degradation_reason=degradation_reason,
)
Measured component blocks (read-only verification against the shipped module):
| Case | components |
|---|---|
calibrate(0.87, T=0.9772731820958189) |
{'calibrated_applied': 1.0, 'temperature': 0.9772731820958189, 'calibration_samples': 16441.0} |
calibrate(0.87, T=1.0) |
{'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0} |
calibrate(0.87, None) |
{'calibrated_applied': 0.0} |
The calibrated_applied flag is a float, not a bool, because components is typed
dict[str, float]. 0.0 / 1.0 is the encoding. raw is retained in every case β the
breakdown always carries both the input and the output, so nothing is destroyed by
calibrating.
8.8 calibrate_result() β preserve the specialist's judgement
def calibrate_result(breakdown, calibration) -> ConfidenceBreakdown:
return calibrate(
breakdown.raw,
calibration,
extra_components=dict(breakdown.components),
degraded=breakdown.degraded,
degradation_reason=breakdown.degradation_reason,
)
evidence/confidence.py:326-346.
The specialist's own components and degradation reason are carried through untouched: this function adds the calibration, it does not re-judge the specialist's measurement.
Pinned by test_specialist_degradation_survives_calibration
(docs/PHASE13_EVIDENCE_ENGINE.md Β§3.5): "A degraded=True fallback result stays degraded
after calibration, with its degradation_reason intact."
9. load_calibration() β the ordered candidate search
9.1 The frozen keys it reads
confidence:
temperature_scaling: true
calibration_file: calibration_v001.json
configs/base.yaml:231-233.
if not bool(config.get("confidence.temperature_scaling", False)):
return None
filename = config.get("confidence.calibration_file")
if not filename:
return None
evidence/confidence.py:396-401. It returns None β never a fabricated default β when:
| Condition | Returns |
|---|---|
| the master switch is off | None |
| no filename is configured | None |
| the artifact is absent from every candidate | None |
| the artifact is present but malformed | raises ValueError |
Every one of those is a legitimate deployment state (an HF Space may ship without the fitted artifact), and each degrades to the pass-through. (
evidence/confidence.py:361-364)
9.2 The ordered search
name = str(filename)
candidates: list[Path] = []
if base_dir is not None:
# An explicit base_dir is honoured first, but a REPO-anchored candidate
# is still tried: several call sites in this repo pass "." meaning "the
# repo", which only works when CWD happens to be the repo root.
candidates.append(Path(base_dir) / name)
from core.config import REPO_ROOT # local import: keeps this module cheap
if base_dir is not None:
candidates.append(Path(REPO_ROOT) / name)
candidates.append(Path(REPO_ROOT) / "artifacts" / name)
candidates.append(Path(REPO_ROOT) / "configs" / name)
for candidate in candidates:
if candidate.exists():
return TemperatureCalibration.from_json(candidate, required=True)
# Nothing found: the master switch is on but no artifact ships. Degrade.
return None
evidence/confidence.py:403-424.
flowchart TB
S["load_calibration(config, base_dir=?)"] --> SW{"confidence.temperature_scaling?"}
SW -->|false| N1["return None"]
SW -->|true| FN{"calibration_file set?"}
FN -->|no| N2["return None"]
FN -->|yes| C1{"base_dir given?"}
C1 -->|yes| A["1. base_dir / name"]
A --> A2{"exists?"}
C1 -->|no| B
A2 -->|yes| LOAD["from_json(required=True)"]
A2 -->|no| B["2. REPO_ROOT / name<br/><i>only when base_dir was given</i>"]
B --> B2{"exists?"}
B2 -->|yes| LOAD
B2 -->|no| C["3. REPO_ROOT / artifacts / name<br/><b>the real home</b>"]
C --> C2{"exists?"}
C2 -->|yes| LOAD
C2 -->|no| D["4. REPO_ROOT / configs / name"]
D --> D2{"exists?"}
D2 -->|yes| LOAD
D2 -->|no| N3["return None β honest degradation"]
style C fill:#3fb95022,stroke:#3fb950
style LOAD fill:#1f6feb22,stroke:#1f6feb
The exact candidate list, by input:
base_dir |
candidate 1 | candidate 2 | candidate 3 | candidate 4 |
|---|---|---|---|---|
None |
β | β | REPO_ROOT/artifacts/name |
REPO_ROOT/configs/name |
"." |
./name |
REPO_ROOT/name |
REPO_ROOT/artifacts/name |
REPO_ROOT/configs/name |
"configs" |
configs/name |
REPO_ROOT/name |
REPO_ROOT/artifacts/name |
REPO_ROOT/configs/name |
"artifacts/calibration/confidence_calibration_v001" |
that dir / name | REPO_ROOT/name |
REPO_ROOT/artifacts/name |
REPO_ROOT/configs/name |
The search is ordered and deterministic: the first existing candidate wins, and no candidate existing means
None(honest degradation, not an error). (evidence/confidence.py:388-389)
9.3 The silent-failure shape this fixes
evidence/confidence.py:366-386 documents the defect precisely:
base_diris resolved in this order (added 2026-09-22):
- an explicit
base_dirargument, when the caller passes one;configs/under the repository root, but only if the artifact is actually there β it never is by default, so this branch exists only to keep an existing caller that parks the artifact beside the config working;<repo root>/artifacts/β the artifact's real home, and the locationscripts/fit_calibration.pywrites to.Step 3 is the important one and it is why this function changed. The frozen config names a bare filename (
calibration_v001.json), while the docstring examples in this repository's own docs passbase_dir="."ANDbase_dir="configs"β two different directories, neither of which isartifacts/. A deployment following either example would silently resolve to a nonexistent path, degrade tomethod="uncalibrated", and report a fitted artifact as never having been fitted. That is precisely the silent-failure shape this project's provenance discipline exists to prevent, so the search now anchors to the repository root using the sameREPO_ROOTthatapp/serving.pyuses for the change checkpoint.
The failure shape, stated as a table:
| Step | Pre-fix behaviour | Post-fix behaviour |
|---|---|---|
caller passes base_dir="." from a non-repo CWD |
resolves to $CWD/calibration_v001.json β absent β None β uncalibrated |
candidate 1 fails, candidate 3 succeeds β the artifact loads |
caller passes base_dir="configs" |
resolves to configs/calibration_v001.json β absent β None β uncalibrated |
candidate 1 fails, candidate 3 succeeds β the artifact loads |
| caller passes nothing | resolved to $CWD/... (base_dir defaulted to Path.cwd() pre-fix) |
candidate 3 succeeds β the artifact loads |
| the artifact is genuinely absent | None |
None (unchanged) |
| the artifact is present but corrupt | depends on the path resolved | raises β never swallowed |
The critical property is the symptom: the pre-fix failure was silent and indistinguishable from "no artifact was ever fitted". A fitted model reported itself as unfitted. That is the same class of defect as a fabricated confidence number β it makes the system lie about its own state.
Verified against the shipped repo (read-only execution):
| Call | Resolved artifact |
|---|---|
load_calibration(config) |
β¦/satquery-ai/artifacts/calibration_v001.json |
load_calibration(config, base_dir="configs") |
β¦/satquery-ai/artifacts/calibration_v001.json |
load_calibration(config, base_dir=".") |
β¦/satquery-ai/artifacts/calibration_v001.json |
All three resolve to the same file via candidate 3. The artifact's own
consumer_contract.resolution field recommends load_calibration(config, base_dir='configs')
β which now works, and works for the reason the fix was made.
The older record, and its status. docs/PHASE13_EVIDENCE_ENGINE.md Β§4.2 was written
before the fix and describes the pre-fix behaviour (base_dir defaults to Path.cwd();
"the controller must pass base_dir explicitly"; "Current behaviour on the real config, for
the record: load_calibration returns None, because no artifact has been fitted yet").
Two things changed since: the search was repo-anchored (2026-09-22), and
artifacts/calibration_v001.json now exists. The current measured behaviour is the table
above. The PHASE13 Β§4.1 expectation of
artifacts/calibration/confidence_calibration_v001/calibration_v001.json did not hold β
the artifact lives at artifacts/calibration_v001.json β which is fine because the bare
filename plus the artifacts/ candidate resolves it.
10. The MEASURED calibration result β a negative result, reported as one
10.1 The artifact
artifacts/calibration_v001.json is the fitted artifact. Its top-level keys:
consumer_contract, created_utc, fit_diagnostics, metrics, provenance,
reliability_diagram, schema, scope, temperature_scaling, type_mask_applied.
fit_diagnostics:
| Field | Value |
|---|---|
temperature |
0.9772731820958189 |
log_temperature |
-0.022989052824434128 |
iterations |
200 |
n_samples |
16441 |
nll_before |
0.6897411093435998 |
nll_after |
0.689630751845387 |
nll_improvement |
0.0001103574982127542 |
effective |
true |
hit_bound |
false |
Note log_temperature = log(0.9772731820958189) = -0.022989β¦ β the fitter optimised in
log-space, and hit_bound: false means the optimum was interior, not at a clamp.
metrics:
| Field | Value |
|---|---|
ece_before |
0.013755 |
ece_after |
0.014929 |
ece_improvement |
-0.001174 (WORSE) |
nll_before |
0.689741 |
nll_after |
0.689631 |
nll_improvement |
0.00011 |
n_bins |
15 |
n_classes |
19 |
n_samples |
16441 |
provenance:
| Field | Value |
|---|---|
checkpoint_path |
artifacts/change_vqa/run/head.pt |
checkpoint_sha256 |
cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a |
config_hash |
78f1e3700da15aa1 |
dataset_id |
cdvqa |
feature_spec |
change_feat_v1 |
fitted_on |
Val |
held_out_splits_excluded |
["Test", "Test2"] |
method |
temperature_scaling |
objective |
mean_negative_log_likelihood |
optimizer |
golden_section_on_log_temperature |
space |
multiclass_logits |
iterations |
200 |
n_classes |
19 |
n_samples |
16441 |
scope β what this artifact calibrates, and what it does not:
This temperature calibrates the R-02 change-VQA head's answer confidence. Other specialists emit their own raw scores and are unaffected.
type_mask_applied: false.
10.2 The result, stated exactly
The raw softmax was already near-calibrated (ECE 0.013755) and temperature scaling made ECE very slightly worse (0.014929) while improving NLL marginally. This is a measurement, not a quality judgment, and it must not be described as scaling being "more accurate". β
docs/STEP7_BACKEND_CHAIN_REPORT.mdΒ§7
docs/FINAL_DELIVERY_TODO.md:104 records it in the measured-metrics table as
0.013755 β 0.014929 (worse) with status "measured (not an improvement)", sourced to
artifacts/calibration_v001.json:26-35. docs/PHASE19_FINAL_HARDENING.md:377 records the
same: "Not claimed as an improvement. β¦ Nothing is calibrated in a deployed path".
| Metric | Before | After | Ξ | Reading |
|---|---|---|---|---|
| ECE (15 bins) | 0.013755 | 0.014929 | β0.001174 | WORSE |
| NLL | 0.689741 | 0.689631 | +0.000110 | marginally better |
Why both numbers matter, and neither alone. The fitter's objective was
mean_negative_log_likelihood, and NLL did improve β that is why T = 0.9773β¦ was
selected at all. But the metric a consumer reads as "is this probability trustworthy" is ECE,
and ECE got worse. Reporting only NLL would present a fitted artifact as a success;
reporting only ECE would hide why the value was chosen. Both are reported, and the headline is
the ECE regression.
Why the reliability diagram is not a single score. The artifact's own note:
Equal-width bins over predicted-class confidence. ECE is bin-count sensitive and is not an aggregate score. β
artifacts/calibration_v001.json,reliability_diagram.note
The diagram is a 15-bin list. The first two bins ([0.0, 0.0667) and [0.0667, 0.1333))
have count: 0 and null accuracy/confidence/gap β the model never predicts that low. The
remaining thirteen bins are populated; the largest is [0.9333, 1.0] with count: 2814,
accuracy: 0.969794, confidence: 0.966895, gap: 0.002899. The largest |gap| is
-0.027817 at [0.6667, 0.7333) (count: 1494).
Bin [lo, hi) |
count | accuracy | confidence | gap |
|---|---|---|---|---|
| 0.0000β0.0667 | 0 | β | β | β |
| 0.0667β0.1333 | 0 | β | β | β |
| 0.1333β0.2000 | 8 | 0.25 | 0.188409 | +0.061591 |
| 0.2000β0.2667 | 437 | 0.283753 | 0.241931 | +0.041822 |
| 0.2667β0.3333 | 736 | 0.290761 | 0.299846 | β0.009085 |
| 0.3333β0.4000 | 626 | 0.386581 | 0.367128 | +0.019454 |
| 0.4000β0.4667 | 690 | 0.450725 | 0.434665 | +0.016060 |
| 0.4667β0.5333 | 1380 | 0.534783 | 0.504472 | +0.030310 |
| 0.5333β0.6000 | 1521 | 0.558185 | 0.565780 | β0.007595 |
| 0.6000β0.6667 | 1395 | 0.624373 | 0.633333 | β0.008961 |
| 0.6667β0.7333 | 1494 | 0.672691 | 0.700507 | β0.027817 |
| 0.7333β0.8000 | 1611 | 0.742396 | 0.766718 | β0.024322 |
| 0.8000β0.8667 | 1701 | 0.833039 | 0.834257 | β0.001218 |
| 0.8667β0.9333 | 2028 | 0.892998 | 0.903143 | β0.010145 |
| 0.9333β1.0000 | 2814 | 0.969794 | 0.966895 | +0.002899 |
(values verbatim from artifacts/calibration_v001.json; the gap column is the artifact's
own accuracy β confidence)
10.3 A structural note on the transform, recorded rather than glossed
The artifact's consumer_contract states:
"applied_as": "sigmoid(logit(z) / T) for a scalar z; softmax(logits / T) for a distribution",
"class": "TemperatureCalibration",
"module": "evidence.confidence",
"read_keys": ["temperature", "fitted_on|split", "artifact", "n_samples"],
"resolution": "load_calibration(config, base_dir='configs')"
The artifact's provenance records space: "multiclass_logits" and n_classes: 19 β the fit
happened over a 19-class softmax. The consumer class
(evidence.confidence.TemperatureCalibration.apply) implements the scalar form:
_clamp01(_sigmoid(_logit(raw) / T)). The class reads exactly the four keys the artifact
declares (temperature, fitted_on|split, artifact, n_samples) and reads nothing else β
the multiclass branch is a documented capability of the artifact format, not a code path in
evidence/confidence.py. Both facts are stated here so the boundary is visible; no claim is
made about which form a deployed path exercises, because
docs/PHASE19_FINAL_HARDENING.md:377 records that "Nothing is calibrated in a deployed
path".
10.4 The deployment state
| Question | Answer | Source |
|---|---|---|
| Does a fitted artifact exist? | yes β artifacts/calibration_v001.json |
filesystem |
Does load_calibration(config) find it? |
yes β via candidate 3 | verified read-only |
Is T effective (non-identity)? |
yes β _is_effective(0.9772731820958189) == True |
verified read-only |
| Does it improve ECE? | no β 0.013755 β 0.014929 | artifact metrics |
| Is it wired into a served path? | OPEN / not claimed β "Nothing is calibrated in a deployed path" |
docs/PHASE19_FINAL_HARDENING.md:377 |
| Which specialist does it cover? | change_vqa only |
artifact scope |
The artifact's config_hash is 78f1e3700da15aa1 β the same frozen hash as
configs/base.yaml (Β§07). The artifact is keyed to the config
that produced it, which is exactly the provenance link the config hash exists to provide.
11. What is deliberately absent from the confidence system
11.1 No LLM path
NO LLM PATH EXISTS HERE. There is deliberately no function that turns text into a number. Confidence is a measurement (freeze section 5: "No LLM-generated confidence"). The only inputs accepted are a float the specialist computed and a calibration artifact fitted on validation data. β
evidence/confidence.py:32-37
docs/ARCHITECTURE_FREEZE.md Β§5 lists it as a non-negotiable: "No LLM-generated coordinates.
No LLM-generated confidence." And the plan Β§26 opens with it: "No LLM-generated confidence."
11.2 No temperature fitter
A fitter is NOT included: Phase 13 fits T on validation data, and inventing one here would be exactly the kind of unfounded number this module refuses to emit.
TemperatureCalibrationis the consumer of a fittedT; the artifact loader reads whatevertraining/calibration/produces. βevidence/confidence.py:57-61
The fitter lives in training/calibration/ (__init__.py, artifact.py, evaluate.py,
fitter.py) and produced the artifact in Β§10. evidence/confidence.py deliberately does not
contain one β the separation is what makes "the number came from a fit" a checkable claim.
11.3 No retrieval path
There is no artifact-serving endpoint in v1 (F-16, Β§2.3). This is why artifact_ref is
permanently null and why the confidence module's artifact field carries a path for
provenance, not a client-facing reference. TemperatureCalibration.artifact is used in
describe() for operator diagnostics; it is not serialised into a client response.
Part D β The execution trace and the eight events
12. ExecutionTrace β observable facts only
class ExecutionTrace(BaseModel):
model_config = ConfigDict(extra="forbid")
run_id: str = Field(default_factory=lambda: _new_id("run"))
schema_version: str = SCHEMA_VERSION
task: Task | None = None
query: str | None = None
inputs: list[str] = Field(default_factory=list)
modalities: list[Modality] = Field(default_factory=list)
intent: Intent | None = None
validation: dict[str, Any] = Field(default_factory=dict)
workflow: list[str] = Field(default_factory=list)
steps: list[TraceStep] = Field(default_factory=list)
selected_models: list[ModelRef] = Field(default_factory=list)
parameters: dict[str, Any] = Field(default_factory=dict)
outputs: list[str] = Field(default_factory=list)
confidence: ConfidenceBreakdown | None = None
timings: dict[str, float] = Field(default_factory=dict)
fallbacks: list[str] = Field(default_factory=list)
errors: list[dict[str, Any]] = Field(default_factory=list)
contradiction: bool = False
config_hash: str | None = None
started_at: str = Field(default_factory=_utcnow)
finished_at: str | None = None
core/schemas.py:296-319. The section header above it reads:
# Execution trace (observable facts only β never chain-of-thought)
(core/schemas.py:277).
| Field | Type | Carries |
|---|---|---|
run_id |
str |
run_<12 hex>, generated per run |
schema_version |
str |
SCHEMA_VERSION = "1.0" |
task |
Task | None |
the routed task |
query |
str | None |
the user's question, verbatim |
inputs |
list[str] |
asset references |
modalities |
list[Modality] |
inferred modalities |
intent |
Intent | None |
the router's advisory output |
validation |
dict[str, Any] |
input-validation facts (input_count, format, modality in the plan Β§27 shape) |
workflow |
list[str] |
the planned step names |
steps |
list[TraceStep] |
the state machine's visits |
selected_models |
list[ModelRef] |
which models ran |
parameters |
dict[str, Any] |
parameters used |
outputs |
list[str] |
output kinds (bbox, answer, β¦) |
confidence |
ConfidenceBreakdown | None |
the confidence breakdown |
timings |
dict[str, float] |
measured durations |
fallbacks |
list[str] |
which fallbacks fired |
errors |
list[dict[str, Any]] |
error records |
contradiction |
bool |
whether a contradiction was detected |
config_hash |
str | None |
the frozen config identity β 78f1e3700da15aa1 for this revision |
started_at / finished_at |
str |
ISO-8601 UTC |
TraceStep (core/schemas.py:279-285):
class TraceStep(BaseModel):
model_config = ConfigDict(extra="forbid")
state: ControllerState
started_at: str = Field(default_factory=_utcnow)
duration_ms: float | None = None
detail: dict[str, Any] = Field(default_factory=dict)
ModelRef (core/schemas.py:288-293):
class ModelRef(BaseModel):
model_config = ConfigDict(extra="forbid")
name: str
revision: str | None = None
role: str | None = None
ControllerState β the nine states (core/schemas.py:79-88):
| # | State | Meaning |
|---|---|---|
| 1 | RECEIVE |
the request arrived |
| 2 | PARSE |
the query was parsed |
| 3 | VALIDATE |
inputs validated |
| 4 | PLAN |
the workflow was chosen |
| 5 | PREPROCESS |
tiling / normalisation |
| 6 | EXECUTE |
specialists ran |
| 7 | AGGREGATE |
evidence was aggregated |
| 8 | VERIFY |
confidence was computed |
| 9 | RESPOND |
the envelope was assembled |
configs/base.yaml Β§agent.states declares exactly this list, in this order β so the
config and the enum cannot drift.
12.1 There is no field for reasoning
ExecutionTrace has no reasoning, thought, rationale, or chain_of_thought field,
and model_config = ConfigDict(extra="forbid") means one cannot be smuggled in. This is a
structural guarantee, not a convention. docs/ARCHITECTURE_FREEZE.md Β§5:
Every result carries an observable execution trace. No chain-of-thought.
The plan Β§27 says the same at the end of its trace example: "No chain-of-thought. Only
observable execution facts." The summary() methods in the evidence layer obey the same
rule β EvidenceCollection.summary()'s docstring: "Observable trace facts. No chain-of-thought,
no interpretation."
13. The eight execution events
The frontend's event protocol is the integration seam between the backend and the instrument
UI. It is declared in frontend/assets/js/core.js:616-620:
SQ.EVENT_NAMES = [
'QUERY_RECEIVED', 'QUERY_UNDERSTOOD', 'ROUTE_SELECTED',
'SPECIALIST_STARTED', 'SPECIALIST_COMPLETED',
'EVIDENCE_GENERATED', 'CONFIDENCE_COMPUTED', 'RESULT_ASSEMBLED'
];
| # | Event | What it reports | Drives the stage |
|---|---|---|---|
| 1 | QUERY_RECEIVED |
the query text, the scene seed, the AOI, the GSD, the date pair | QUERY |
| 2 | QUERY_UNDERSTOOD |
the routed task, the parsed slots, the policy id | UNDERSTAND |
| 3 | ROUTE_SELECTED |
the task, the specialist list, the rule trace, the policy version | ROUTE |
| 4 | SPECIALIST_STARTED |
which component, its index/ordinal, its stage, its model, its device | ANALYZE / GROUND |
| 5 | SPECIALIST_COMPLETED |
which component, its duration, its output kind | (advances the active marker) |
| 6 | EVIDENCE_GENERATED |
the evidence regions, the count, the threshold, the registration RMSE | EVIDENCE |
| 7 | CONFIDENCE_COMPUTED |
the reported/calibrated values, the method, the reliability bins | CONFIDENCE |
| 8 | RESULT_ASSEMBLED |
the answer text, the task, the specialists, the evidence ids, the confidence, the provenance | ANSWER |
The stage rail is a separate, coarser list (frontend/assets/js/core.js:605-614):
SQ.STAGES = [
{ id: 'QUERY', label: 'QUERY' },
{ id: 'UNDERSTAND', label: 'UNDERSTAND' },
{ id: 'ROUTE', label: 'ROUTE' },
{ id: 'ANALYZE', label: 'ANALYZE' },
{ id: 'GROUND', label: 'GROUND' },
{ id: 'EVIDENCE', label: 'EVIDENCE' },
{ id: 'CONFIDENCE', label: 'CONFIDENCE' },
{ id: 'ANSWER', label: 'ANSWER' }
];
Eight stages for eight events β but not a 1:1 mapping: events 4 and 5 share the
ANALYZE/GROUND stages, because a specialist's start and completion are two events about
one stage visit.
13.1 The run engine is deliberately dumb
frontend/assets/js/core.js:8-11:
The run engine is deliberately dumb: it renders whatever events it receives. Nothing about the visuals depends on the events being synthetic. Swapping the mock driver for a websocket / SSE feed of the same event names is the entire integration surface.
/**
* THE INTEGRATION SEAM.
* Feed real execution events here β same names, same payload shapes β and
* every visual state in the prototype updates identically.
*/
ingest: function (type, payload) { ... }
frontend/assets/js/core.js:737-742. ingest() is a switch over the eight names, each
case updating state and calling emit(). emit() records
{type, t, seq, payload} where t = performance.now() - t0 and seq is a monotonic counter
(frontend/assets/js/core.js:709-718).
Two drivers exist, and they are not interchangeable:
| Driver | When | Payload honesty |
|---|---|---|
startMock(query, scene) |
preview only β when no file is selected | synthetic; marked is-mock |
runLive(query) in frontend/assets/js/mission.js |
the production path | payloads come from the real network response |
docs/FINAL_DELIVERY_TODO.md:73 records the distinction: "runMock only when no file is
selected; emits empty payloads, marked is-mock; not in production path." The live driver
emits all eight events around real network calls (docs/FINAL_DELIVERY_TODO.md:71):
"runLive emits 8 events around real network calls; markState uses real values."
13.2 markState β the trace is driven by events, never by a timer
frontend/assets/js/mission.js:404-425:
function markState(id, note, isLive) {
var n = traceNodes[id];
if (!n) return;
n.node.classList.toggle('is-mock', !isLive);
n.tm.textContent = note !== undefined ? note : (STATE_NOTE[id] || '');
/* Advance the trace to the furthest state reached. This is driven by the
same events that carry the real result, so the bar and the node states
move only when the run actually reaches a stage β never on a timer. The
fill spans from the left edge to the centre of the current node. */
var idx = STATES.indexOf(id);
if (idx > traceProgress) traceProgress = idx;
...
if (traceFill) {
traceFill.style.width = (((traceProgress + 0.5) / STATES.length) * 100) + '%';
}
}
The design property in the comment is the important one: the progress bar is a function of reached states, not of elapsed time. A run that stalls shows a stalled bar.
13.3 The measured 94.4444 % trace fill
docs/FINAL_DELIVERY_TODO.md:72:
| Area | Status | Evidence |
|---|---|---|
| Execution trace (progress bar) | VERIFIED | .trace__fill width is set from real event count (measured 94.4444 % live, 2026-09-25) |
The number is derivable from the formula, which is why it is a real measurement rather
than a coincidence. The live path reaches RESPOND, the ninth of nine ControllerState
values:
width = ((traceProgress + 0.5) / STATES.length) * 100
= ((8 + 0.5) / 9) * 100
= (8.5 / 9) * 100
= 94.4444β¦ %
STATES is the nine-state ControllerState list (Β§12); traceProgress = 8 is the
zero-based index of RESPOND. The + 0.5 is what makes the bar span to the centre of the
current node rather than to its left edge β so a fully-completed run stops at 94.4444 %, not
100 %, because the last node's centre is half a node-width short of the right edge. A bar
that read 100 % on a nine-state rail would be reporting something the run never did.
14. Worked example β one item's full journey
The following traces a single grounding claim from a specialist's Box to a client-facing
evidence item. Every step is a real code path.
sequenceDiagram
participant G as GroundingSpecialist
participant E as EvidenceEngine
participant C as confidence.calibrate
participant T as ExecutionTrace
participant F as Frontend
G->>G: compute Box(x1,y1,x2,y2, score, label, coordinate_system)
G->>E: SpecialistResult.evidence = [Evidence(type=bounding_box, ...)]
Note over E: aggregate([grounding_result])
E->>E: collect β [ev_a1b2c3d4e5f6]
E->>E: _identity_key β (bounding_box, grounding, normalized_0_1, (...), 0.87)
E->>E: _deduplicate β no collision
E->>E: sorted(_sort_key)
E->>E: _record_agreement β single specialist, no annotation
E->>E: _renumber β evidence_001
E-->>T: EvidenceCollection(items=[evidence_001], sources=['grounding'], ...)
E->>C: confidence_for(grounding_result)
C->>C: _is_effective(0.9772731820958189) == True
C->>C: _logit(0.87) / 0.9772732 β _sigmoid β _clamp01
C-->>T: ConfidenceBreakdown(raw=0.87, calibrated=0.8749186β¦, method="temperature_scaling")
T->>F: emit EVIDENCE_GENERATED {regions, count, threshold, registrationRMSE}
T->>F: emit CONFIDENCE_COMPUTED {calibrated, method, bins, observed}
T->>F: emit RESULT_ASSEMBLED {text, evidenceIds, confidence, provenance}
The state of the item at each step:
| Step | evidence_id |
type |
source_specialist |
coordinate_system |
coordinates |
score |
artifact_ref |
payload |
|---|---|---|---|---|---|---|---|---|
| specialist output | ev_a1b2c3d4e5f6 |
bounding_box |
grounding |
normalized_0_1 |
[0.21,0.33,0.47,0.61] |
0.87 |
null |
{"label": "water"} |
| after dedup | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged |
| after sort | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged |
| after agreement | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | unchanged | {"label": "water"} (no corroboration β single specialist) |
| after renumber | evidence_001 |
β¦ | β¦ | β¦ | β¦ | β¦ | β¦ | β¦ |
And the source object is untouched. After aggregate returns, the original
SpecialistResult.evidence[0] still has evidence_id == "ev_a1b2c3d4e5f6". This is the purity
contract (Β§5.4) made concrete: _renumber calls model_copy(update={"evidence_id": ...}),
which produces a new object and leaves the source alone.
Verified confidence values (read-only execution against the shipped module):
| Input | raw |
calibrated |
method |
|---|---|---|---|
calibrate(0.87, T=0.9772731820958189) |
0.87 |
0.8749186077809417 |
temperature_scaling |
calibrate(0.87, T=1.0) |
0.87 |
None |
uncalibrated |
calibrate(0.0, None) |
0.0 |
None |
uncalibrated |
Part E β Honest boundaries
15. What is NOT RUN, OPEN, BLOCKED or REJECTED for this topic
| Item | Status | Note |
|---|---|---|
| Calibration improves ECE | REJECTED |
measured 0.013755 β 0.014929 (worse). The artifact is retained because it is in the frozen config, not because it helps. |
| Calibration applied in a served path | OPEN |
docs/PHASE19_FINAL_HARDENING.md:377 β "Nothing is calibrated in a deployed path" |
Calibration covers specialists other than change_vqa |
NOT RUN |
the artifact's scope covers change_vqa only; every other specialist's confidence is uncalibrated |
| A second fitted artifact for the other specialists | NOT RUN |
no artifact exists |
EvidenceEngine wired into core/controller.py |
IMPLEMENTED (Phase 14) |
docs/PHASE13_EVIDENCE_ENGINE.md Β§8 recorded it as not-yet-wired at Phase 13; the live run path emits the eight events |
| Artifact rendering (crops, masks, change maps served to a client) | OPEN |
"artifact_ref and the evidence_from_* methods provide the hooks; rendering crops, masks and change maps is controller/GUI territory." (docs/PHASE13_EVIDENCE_ENGINE.md Β§8) |
| Artifact retrieval endpoint | REJECTED for v1 |
F-16: no artifact-serving endpoint exists, so artifact_ref is permanently null |
contributing_specialists field on Evidence |
REJECTED |
proposed and not adopted; core/schemas.py unchanged (docs/PHASE13_EVIDENCE_ENGINE.md Β§5) |
Change-degradation clause in SpecialistResult |
RESOLVED (removed, F-16c) |
the narrower replacement rule is REJECTED (core/schemas.py:362-395) |
payload["corroborated_by"] exercised by a live multi-specialist run |
UNKNOWN β not established from the available evidence |
the annotation is pinned by unit tests; no live run was found that produced two specialists making an identical claim |
The multiclass (softmax(logits/T)) form of the calibration applied in code |
UNKNOWN β not established from the available evidence |
the artifact documents the capability; evidence/confidence.py implements the scalar form (Β§10.3) |
Whether _sort_key's coordinate tie-break is exercised with non-equal coordinates in production |
UNKNOWN β not established from the available evidence |
pinned by unit test test_equal_scores_order_deterministically_by_coordinates |
The exact count of live runs that populated dropped_over_limit > 0 |
UNKNOWN β not established from the available evidence |
no run record was found |
evidence.max_items = 32 ever being the binding constraint in a live run |
UNKNOWN β not established from the available evidence |
the cap is config-declared; no live run's total_before_limit was found |
16. Where the evidence lives
| Claim | Source |
|---|---|
Evidence fields and the spatial validator |
core/schemas.py:216-253 |
artifact_ref is permanently null (F-16) |
core/schemas.py:225-235; docs/API_CONTRACT.md Β§2.4; docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5 |
| F-16c and the removed change clause | core/schemas.py:362-395; docs/DEPLOYMENT_ARCHITECTURE.md:855-866 |
EvidenceType β all 11 members |
core/schemas.py:65-76 |
CoordinateSystem β 3 values |
core/schemas.py:57-62 |
| Purity contract | evidence/engine.py:38-50 |
_identity_key and its exclusions |
evidence/engine.py:125-142; :52-68 |
_claim_key |
evidence/engine.py:145-162 |
_sort_key and the four-level rationale |
evidence/engine.py:165-190 |
_deduplicate and the payload merge |
evidence/engine.py:352-389 |
_record_agreement and AGREEMENT_KEY |
evidence/engine.py:391-425; :234-238 |
_renumber and ID_PREFIX |
evidence/engine.py:427-430; :104-107 |
| The cap and the five accounting fields | evidence/engine.py:193-217; :337-350 |
evidence_digest |
evidence/engine.py:680-704 |
evidence_type_for |
evidence/engine.py:434-458 |
evidence_from_box/region/geospatial |
evidence/engine.py:462-572 |
confidence_for and the degradation-wins rule |
evidence/engine.py:576-637 |
| The honesty rule | evidence/confidence.py:12-30 |
logit/sigmoid/_EPS/_clamp01 |
evidence/confidence.py:80-115 |
_is_effective and _IDENTITY_TOLERANCE |
evidence/confidence.py:90-93, :118-126 |
TemperatureCalibration and its range guard |
evidence/confidence.py:129-252 |
calibrate / calibrate_result |
evidence/confidence.py:255-346 |
load_calibration and the ordered search |
evidence/confidence.py:349-424 |
| The 77-test phase record | docs/PHASE13_EVIDENCE_ENGINE.md |
| The measured calibration result | artifacts/calibration_v001.json; docs/STEP7_BACKEND_CHAIN_REPORT.md Β§7; docs/FINAL_DELIVERY_TODO.md:104 |
| The eight events | frontend/assets/js/core.js:616-620 |
| The stage rail | frontend/assets/js/core.js:605-614 |
markState and the fill formula |
frontend/assets/js/mission.js:404-425 |
| The 94.4444 % measurement | docs/FINAL_DELIVERY_TODO.md:72 |
ExecutionTrace / TraceStep / ModelRef |
core/schemas.py:279-319 |
ControllerState β nine states |
core/schemas.py:79-88; configs/base.yaml Β§agent.states |
evidence.max_items: 32 |
configs/base.yaml Β§evidence |
confidence.* frozen keys |
configs/base.yaml Β§confidence |
| Layer verbs | docs/ARCHITECTURE_FREEZE.md Β§5 |
Next: 07 Configuration freeze β the registry, the enforced
invariants, and the hash 78f1e3700da15aa1 that every artifact in this chapter is keyed to.