SatQuery / docs /architecture /06-evidence-and-confidence.md
thundercode's picture
release: add docs/architecture/06-evidence-and-confidence.md
928a5aa verified
|
Raw History Blame Contribute Delete
108 kB

06 β€” Evidence and Confidence

Parent: Architecture hub Β· Sibling chapters: 01 System overview Β· 05 Specialists Β· 07 Configuration freeze

Primary sources read for this chapter (all under C:/Users/anish/satquery-ai/):

Source What it establishes here
evidence/engine.py (714 lines) the aggregation pipeline, purity contract, identity/sort/claim keys, _deduplicate, _record_agreement, _renumber, evidence_type_for, evidence_from_box/region/geospatial, confidence_for, evidence_digest
evidence/confidence.py (434 lines) the honesty rule, TemperatureCalibration, _is_effective, _logit/_sigmoid, _EPS, calibrate, calibrate_result, load_calibration and its ordered candidate search
core/schemas.py (462 lines) Evidence, EvidenceType, ConfidenceBreakdown, ExecutionTrace, TraceStep, ModelRef, SpecialistResult, CoordinateSystem, and every validator
configs/base.yaml (Β§evidence, Β§confidence) evidence.max_items: 32, confidence.temperature_scaling: true, confidence.calibration_file: calibration_v001.json
artifacts/calibration_v001.json the measured fitted temperature and the reliability diagram
frontend/assets/js/core.js SQ.EVENT_NAMES β€” the eight execution events
frontend/assets/js/mission.js markState() and the .trace__fill width formula
docs/ARCHITECTURE_FREEZE.md Β§1, Β§3, Β§5 the layer verbs and the frozen non-negotiables
docs/PHASE13_EVIDENCE_ENGINE.md the phase record for this package (77 tests)
docs/API_CONTRACT.md Β§2.4, Β§3.3, Β§4 the client-facing shape of evidence and confidence
docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5 the F-16 owner ruling on artifact refs
docs/STEP7_BACKEND_CHAIN_REPORT.md Β§7, Β§8 the measured calibration and evidence integration results

1. Where this subsystem sits

docs/ARCHITECTURE_FREEZE.md Β§5 gives every layer exactly one verb:

Router understands; policy engine decides; specialists compute; VLM explains; evidence engine proves.

Two consecutive stages implement that last verb:

flowchart LR
  SPEC["Specialists<br/>compute"] -->|"SpecialistResult.evidence"| EV["Evidence engine<br/>aggregate()<br/>collect β†’ dedup β†’ sort β†’<br/>annotate β†’ renumber β†’ cap"]
  EV -->|"EvidenceCollection"| CONF["Confidence<br/>confidence_for() / calibrate()"]
  CONF -->|"ConfidenceBreakdown"| RES["ResultEnvelope<br/>+ ExecutionTrace"]
  style EV fill:#1f6feb22,stroke:#1f6feb
  style CONF fill:#1f6feb22,stroke:#1f6feb

The freeze Β§1 pipeline diagram places them adjacently:

   +---------+---------+---------+
   |         |         |         |
 VQA/CAP  GROUNDING  CHANGE   OPTICAL-SAR
   |         |         |         |
   +---------+---------+---------+
                 |
                 v
          Evidence engine          <- this chapter, part A
                 |
                 v
       Confidence calibration      <- this chapter, part B
                 |
                 v
          Result normaliser  ->  JSON + trace + PDF

Division of labour, stated precisely. evidence/engine.py does not re-derive any specialist's claim. The module docstring is explicit (evidence/engine.py:10-14):

This module is the "proves" step. Specialists each emit Evidence for what they computed (see GroundingSpecialist._build_evidence). This engine does not re-derive any of that. Its job is aggregation:

collect across specialists -> order -> deduplicate -> renumber -> bound

Status of this subsystem.

Component Status Basis
Evidence aggregation IMPLEMENTED + VERIFIED evidence/engine.py; 77 tests in tests/unit/test_evidence_engine.py (docs/PHASE13_EVIDENCE_ENGINE.md Β§7)
Dedup / ordering / renumber / cap IMPLEMENTED + VERIFIED pinned by named tests (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.2–§2.5)
Confidence honesty rule (uncalibrated pass-through) IMPLEMENTED + VERIFIED evidence/confidence.py:296-311
Temperature-scaling transform IMPLEMENTED + VERIFIED monotonicity + endpoint-safety tests (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3)
Calibration artifact fitted MEASURED artifacts/calibration_v001.json
Calibration improves ECE REJECTED β€” it does not ECE 0.013755 β†’ 0.014929 (worse); see Β§7
Evidence engine wired into the controller IMPLEMENTED (Phase 14) docs/PHASE13_EVIDENCE_ENGINE.md Β§8 recorded it as "not yet wired"; the live run path emits the eight events (Β§9)
Artifact rendering / retrieval OPEN β€” deliberately null in v1 F-16 ruling, Β§3.2

Part A β€” The evidence schema

2. Evidence β€” the canonical record

Evidence is defined once, in core/schemas.py:216-253. Every specialist emits it; nothing invents a second shape (docs/ARCHITECTURE_FREEZE.md Β§3: "Single schema for every specialist. No specialist invents its own shape.").

class Evidence(BaseModel):
    model_config = ConfigDict(extra="forbid")

    evidence_id: str = Field(default_factory=lambda: _new_id("ev"))
    type: EvidenceType
    source_specialist: str
    coordinate_system: CoordinateSystem | None = None
    coordinates: list[float] | None = None
    score: float | None = Field(default=None, ge=0.0, le=1.0)
    artifact_ref: str | None = Field(
        default=None,
        description=(
            "Reference to an externally retrievable artifact. NEVER a "
            "filesystem path (F-16, owner ruling 2026-09-23): v1 exposes no "
            "artifact-serving endpoint, so this is null unless a deployment "
            "supplies a client-fetchable reference. An artifact may still be "
            "written server-side where configured; being written is not the "
            "same as being retrievable."
        ),
    )
    payload: dict[str, Any] = Field(default_factory=dict)

2.1 Every field, exhaustively

Field Type Required Default Meaning Note
evidence_id str no _new_id("ev") β†’ ev_<12 hex> the item's identity run-local, engine-assigned after aggregation (Β§5.8); a uuid before that
type EvidenceType yes β€” the kind of proof closed 11-member enum, Β§3
source_specialist str yes β€” which specialist made this claim a single string, not a set β€” this is load-bearing, Β§5.7
coordinate_system CoordinateSystem | None no None the frame the coordinates live in mandatory in practice for spatial types β€” see the validator, Β§4
coordinates list[float] | None no None the geometry a flat list; a box is [x1, y1, x2, y2]
score float | None no None reliability, bounded ge=0.0, le=1.0 β€” Pydantic rejects out-of-range
artifact_ref str | None no None a retrievable artifact reference always null in v1 β€” Β§2.3
payload dict[str, Any] no {} observations that support the claim without being it where corroborated_by is written, Β§5.7

2.2 What Evidence deliberately does not have

These absences are load-bearing; each one is why a downstream design decision exists.

Absent field Why it is absent Consequence
any ordering field (index, rank, order) the engine owns order; a specialist must not pre-empt it order is defined by _sort_key, Β§5.6
value the field is called score docs/PHASE19_FINAL_HARDENING.md Β§3.6 records that the docs once said value and that was one of nine validated defects
source the field is called source_specialist same defect class; a frontend using source renders nothing
summary not in the schema same defect class
contributing_specialists: list[str] proposed and NOT adopted see Β§5.7 and docs/PHASE13_EVIDENCE_ENGINE.md Β§5
label at top level labels live in payload evidence_from_box writes payload["label"]
a retrievable boolean retrieval is not a v1 capability artifact_ref being null is the signal, Β§2.3

model_config = ConfigDict(extra="forbid") is what makes all of the above enforceable: a response carrying value instead of score is a 422, not a silently-ignored key. The conformance test asserts this directly β€” docs/STEP7_BACKEND_CHAIN_REPORT.md Β§D records "extra="forbid" on all eight client-facing models; GeoMetadata is the one documented extra="allow" exception."

2.3 artifact_ref is permanently null in v1 β€” owner ruling F-16

This is the single most mis-documented field in the system, and the ruling is recorded in three places: the field's own description (core/schemas.py:225-235), the client contract (docs/API_CONTRACT.md Β§2.4 "Artifact refs β€” null in v1, and why"), and the audit (docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5).

The ruling (F-16, owner, 2026-09-23): never expose filesystem paths.

docs/API_CONTRACT.md:586-602 states the contract in full:

Every artifact_ref and change_map in a v1 response is null. This is a deliberate contract, not a missing value.

F-16 (owner ruling 2026-09-23): never expose filesystem paths. The specialists do render their artifacts β€” the change map and the optical/SAR views are written server-side β€” but their location is an operator fact, not a client-facing one. A response that carried the server's path would disclose the deployment's directory layout to an unauthenticated caller, and nothing the frontend can do requires it.

No artifact:// URI is fabricated in its place. v1 has no artifact-serving endpoint, so a URI would be a promise the service cannot keep β€” strictly worse than null, because the frontend would build a link that 404s.

The docstring's own phrasing is the crispest statement of the principle:

being written is not the same as being retrievable.

flowchart TB
  subgraph Server["Server side (operator facts)"]
    R["specialist renders change map / views"]
    W["file written where configured<br/>change.artifact_dir"]
    R --> W
  end
  subgraph Client["Client-facing contract (v1)"]
    N["artifact_ref = null"]
    P["payload statistics<br/>total_change_pixels<br/>n_components_kept<br/>threshold"]
    WARN["warnings[] entry:<br/>artifact NOT retrievable"]
  end
  W -.->|"deliberately NOT exposed"| X["filesystem path<br/>(never sent)"]
  R --> N
  R --> P
  R --> WARN
  style X fill:#f8514922,stroke:#f85149
  style N fill:#3fb95022,stroke:#3fb950

What replaces the ref (docs/API_CONTRACT.md:604-610):

Removed Replaced by
change_map path null, plus the change statistics in the CHANGE_MAP evidence's payload (total_change_pixels, n_components_kept, threshold)
view artifact_ref path null, plus payload.rendered / payload.retrievable / payload.retrieval
β€” an explicit warnings[] entry saying the artifact is NOT retrievable

Two live carriers, and the half-implementation that followed. The audit records that the ruling was initially applied to only one carrier (docs/DEPLOYMENT_ARCHITECTURE.md:361):

Two live carriers β€” change and croma β€” and both had to be brought into line: fixing only the croma carrier left the ruling half implemented.

docs/STATUS.md:21 names the lesson: "The finding worth carrying forward: the F-16 ruling was HALF IMPLEMENTED, and the suite was green."

F-16c β€” the consequence the fix introduced, measured and left open. The ruling removed a proxy and thereby changed a degraded signal. The SpecialistResult validator used to read (core/schemas.py:362-372, retained as a comment):

if self.task is Task.CHANGE and self.change_map is None and not self.regions:
    ... "change analysis produced no spatial output" ... degraded = True

change_map was a proxy for "a map was produced". With the ref permanently null, the clause collapsed to not regions, and a successful no-change analysis began reporting degraded: true. The clause was therefore removed under ruling F-16c (core/schemas.py:362-395). The comment left in its place is worth quoting because it states the reasoning better than a summary could:

That conflated two different things, and the conflation WAS the defect. degraded means "the analysis could not be fully performed": no trained detector, or co-registration too poor to support a spatial claim. Both are set by the specialist itself (specialists/change/specialist.py:345 and :365) and neither is this validator's to invent. "No change was detected" is a NORMAL, successful outcome.

And, on why no replacement clause was added:

No replacement clause is added, deliberately. A narrower "empty answer => degraded" rule was tried and backed out: it is not what the ruling asked for, it invented a semantic the specialist already owns, and it made a pre-existing, unrelated fixture (test_change_with_regions_is_not_degraded, a result with regions and no answer) fail. A fix that forces edits to tests it has nothing to do with is signalling over-reach, not diligence. CHANGE is the one task whose degraded flag is now set entirely by its specialist.

Status: F-16 CLOSED (both carriers conform). F-16c RESOLVED by clause removal β€” the narrower replacement rule is REJECTED. mask_ref is not an exemption: the audit corrected an earlier working note that had listed it alongside artifact_ref (docs/DEPLOYMENT_ARCHITECTURE.md:868-871).

3. EvidenceType β€” the closed 11-member vocabulary

class EvidenceType(str, Enum):
    IMAGE_CROP = "image_crop"
    TILE = "tile"
    BOUNDING_BOX = "bounding_box"
    MASK = "mask"
    CHANGE_MAP = "change_map"
    OPTICAL_VIEW = "optical_view"
    SAR_VIEW = "sar_view"
    JOINT_FEATURE_REGION = "joint_feature_region"
    STATISTIC = "statistic"
    GEOLOCATION = "geolocation"
    AVAILABILITY_MASK = "availability_mask"   # C-1: modality trust evidence

core/schemas.py:65-76. All eleven members, with what each one asserts:

# Member Wire value Asserts Typically emitted by
1 IMAGE_CROP image_crop a rectangular raster crop exists preprocessing / tiling
2 TILE tile one tile of the tiling policy was examined preprocessing/tiling.py
3 BOUNDING_BOX bounding_box a rectangle locates a referent grounding, change
4 MASK mask a per-pixel region exists grounding, change
5 CHANGE_MAP change_map a bitemporal difference map exists change specialist
6 OPTICAL_VIEW optical_view the optical rendering of an optical/SAR pair optical-SAR specialist
7 SAR_VIEW sar_view the SAR rendering of an optical/SAR pair optical-SAR specialist
8 JOINT_FEATURE_REGION joint_feature_region a region defined in CROMA's joint embedding space optical-SAR specialist
9 STATISTIC statistic a measured scalar or count, with no geometry VQA, caption, unsupported, and any geometry-less region
10 GEOLOCATION geolocation the raster is georeferenced and where it sits geospatial layer
11 AVAILABILITY_MASK availability_mask which modality channels were actually available sensor adapter (C-1)

3.1 availability_mask β€” the member the freeze prose omits

docs/ARCHITECTURE_FREEZE.md Β§3's prose list of evidence types stops at ten members and does not name availability_mask. The enum member is nevertheless real, and the engine says so explicitly (evidence/engine.py:441-447):

The members are exactly those of core.schemas.EvidenceType. The freeze's section 3 prose list omits availability_mask; that member is real (C-1: modality trust evidence) and is included here because a missing member would otherwise force a wrong fallback.

Its existence follows from finding C-1: the channel-availability mask is a first-class fusion input and is never a CROMA input (core/schemas.py:7; docs/PHASE0_CONTRACT_VALIDATION.md Β§1.2). Because the mask is a first-class input, it is also a first-class observation β€” "these twelve optical channels were present and these two SAR channels were not" is a fact a consumer can audit. docs/API_CONTRACT.md:579 lists all eleven values for the frontend.

3.2 Which types are spatial

The engine's own validator names the spatial set (core/schemas.py:240-247) β€” this is the authoritative list, not a prose paraphrase:

spatial = {
    EvidenceType.BOUNDING_BOX,
    EvidenceType.MASK,
    EvidenceType.CHANGE_MAP,
    EvidenceType.TILE,
    EvidenceType.IMAGE_CROP,
    EvidenceType.JOINT_FEATURE_REGION,
}
Type Spatial? Why
BOUNDING_BOX yes a box is meaningless without a frame
MASK yes pixels need a frame
CHANGE_MAP yes a map is a raster
TILE yes a tile is a rectangle of a raster
IMAGE_CROP yes a crop is a rectangle of a raster
JOINT_FEATURE_REGION yes a region in a feature grid still needs a frame
STATISTIC no a scalar has no geometry
GEOLOCATION no (by this validator) geo bounds are already self-describing via payload["crs"]; evidence_from_geospatial sets CoordinateSystem.GEO when bounds are given
OPTICAL_VIEW no (by this validator) a whole-frame view is not a sub-region claim
SAR_VIEW no (by this validator) as above
AVAILABILITY_MASK no (by this validator) a channel-presence vector is not spatial

4. The validator that refuses coordinates without a frame

@model_validator(mode="after")
def _spatial_needs_crs(self) -> "Evidence":
    spatial = {
        EvidenceType.BOUNDING_BOX,
        EvidenceType.MASK,
        EvidenceType.CHANGE_MAP,
        EvidenceType.TILE,
        EvidenceType.IMAGE_CROP,
        EvidenceType.JOINT_FEATURE_REGION,
    }
    if self.type in spatial and self.coordinates and self.coordinate_system is None:
        raise ValueError(
            f"evidence type '{self.type.value}' carries coordinates "
            "but no coordinate_system"
        )
    return self

core/schemas.py:238-253. Three things about this are deliberate:

  1. It is a model_validator(mode="after"), not a field_validator. The rule is cross-field β€” it depends on type, coordinates and coordinate_system together, so it cannot be expressed on any single field.

  2. The trigger is self.coordinates, not is not None. An empty list is falsy and does not trip the rule; a non-empty list does. A spatial item with no coordinates at all is legal (a MASK whose geometry lives only in a mask_ref).

  3. It raises rather than defaulting. Defaulting to normalized_0_1 would silently mislabel pixel or geo coordinates as normalized, which is the failure mode the finding exists to prevent. docs/ARCHITECTURE_FREEZE.md Β§3 puts the requirement bluntly:

    Every spatial object carries coordinate_system ∈ {normalized_0_1, pixel, geo}.

docs/ARCHITECTURE_FREEZE.md Β§2.3 states the origin: internal box coordinates are normalised 0–1, always with an explicit coordinate_system field (finding C-5).

4.1 CoordinateSystem β€” three values, exact spellings

class CoordinateSystem(str, Enum):
    """Never omit this. A bare box is meaningless without it. (C-5)"""

    NORMALIZED_0_1 = "normalized_0_1"
    PIXEL = "pixel"
    GEO = "geo"

core/schemas.py:57-62.

Member Wire value Frame
NORMALIZED_0_1 normalized_0_1 fractions of the frame, [0, 1]
PIXEL pixel integer-ish pixel indices
GEO geo a projected/geographic CRS named in payload["crs"]

Documentation hazard, measured. docs/PHASE19_FINAL_HARDENING.md Β§3.6 records that the docs once spelled these normalized / geographic. The real values are normalized_0_1 / geo. "A frontend using the documented names would send values the server rejects with 422." docs/STEP7_BACKEND_CHAIN_REPORT.md Β§D confirms the live enum is exactly {normalized_0_1, pixel, geo}.

Evidence's coordinate_system is CoordinateSystem | None β€” the field is optional in type but effectively mandatory in practice, because the validator above enforces it for every spatial type that carries coordinates.


Part B β€” The aggregation pipeline

5. EvidenceEngine.aggregate() β€” collect β†’ dedup β†’ sort β†’ annotate β†’ renumber β†’ cap

5.1 Why the module exists: three problems a specialist cannot solve

The module docstring enumerates them (evidence/engine.py:18-36); each is real and none is fixable inside a single specialist:

# Problem Why a specialist cannot fix it
1 Stable identity Evidence.evidence_id defaults to a random uuid β€” fine within one result, useless once four specialists' evidence is merged into one trace. "Nothing can be cited."
2 Order Evidence has no ordering field, so the collection's order follows specialist completion order β€” the same inputs produce different JSON on different runs. "Pure functions must not do that, and this whole system's reproducibility argument rests on it."
3 Duplication the VQA specialist emits a STATISTIC carrying its answer, and that answer travels in SpecialistResult.answer; two specialists that georeference the same asset emit the same GEOLOCATION. "Neither is wrong, and neither knows about the other."

5.2 The pipeline, in one place

flowchart TB
  A["results: SpecialistResult<br/>or Sequence[SpecialistResult]<br/>or evidence: Iterable[Evidence]"] --> B["collect<br/>extend(result.evidence)"]
  B --> C["sources = sorted({source_specialist})<br/><i>recorded BEFORE the cap</i>"]
  C --> D["_deduplicate<br/>identity = _identity_key<br/>payloads MERGED on collision"]
  D --> E["sorted(key=_sort_key)<br/>type ↑ Β· specialist ↑ Β·<br/>score ↓ Β· coordinates ↑"]
  E --> F["_record_agreement<br/>_claim_key groups β†’<br/>payload['corroborated_by']"]
  F --> G["annotated[:max_items]<br/>dropped_over_limit = total βˆ’ len(capped)"]
  G --> H["_renumber β†’ evidence_001…"]
  H --> I["EvidenceCollection"]
  style D fill:#1f6feb22,stroke:#1f6feb
  style F fill:#1f6feb22,stroke:#1f6feb

The implementation, verbatim (evidence/engine.py:324-350):

def _canonicalise(self, raw: Iterable[Evidence]) -> EvidenceCollection:
    """Dedup -> sort -> renumber -> cap. The whole pipeline, in one place."""
    items = list(raw)
    sources = sorted({item.source_specialist for item in items})

    deduped, dropped_duplicates = self._deduplicate(items)
    ordered = sorted(deduped, key=_sort_key)
    # Two specialists can make the same claim, and `deduplicate` keeps both
    # (different `source_specialist` -> different key). Record the agreement
    # now that ordering is fixed, so the annotation is part of the pure
    # pipeline rather than a post-hoc edit a caller might forget.
    annotated = self._record_agreement(ordered)

    total = len(annotated)
    capped = annotated[: self.max_items]
    dropped_over_limit = total - len(capped)

    return EvidenceCollection(
        items=[
            self._renumber(item, index) for index, item in enumerate(capped, start=1)
        ],
        sources=sources,
        dropped_duplicates=dropped_duplicates,
        dropped_over_limit=dropped_over_limit,
        total_before_limit=total,
        truncated=dropped_over_limit > 0,
    )

Note the ordering choice: annotate comes after sort but before cap. Annotating before sorting would be equivalent (the annotation is a payload key, which is not in _sort_key), but the docstring's stated reason is that the annotation must be inside the pure pipeline rather than "a post-hoc edit a caller might forget".

5.3 The entry point and its argument rules

def aggregate(
    self,
    results: SpecialistResult | Sequence[SpecialistResult] | None = None,
    *,
    evidence: Iterable[Evidence] | None = None,
) -> EvidenceCollection:

evidence/engine.py:280-320. Two argument invariants, both raising:

Condition Behaviour Reason given in the docstring
both results and evidence are None ValueError("aggregate() requires results= or evidence=") there is nothing to aggregate
both are supplied ValueError("aggregate() accepts results= or evidence=, not both") ambiguity about precedence
empty list returns an empty EvidenceCollection "Never raises on empty input."

Only result.evidence is read. The docstring explains why reading both forms would double-count (evidence/engine.py:292-297):

Specialists state their evidence twice β€” in SpecialistResult.evidence and via Specialist.produce_evidence, which currently delegates to the same list β€” so only result.evidence is read. Reading both would double-count every item and inflate the duplicate count.

5.4 The purity contract

"""THE PURITY CONTRACT
-------------------
`aggregate` is pure and deterministic:

    * source results are never mutated -- `model_copy` is used to rebuild the
      collection rather than editing `Evidence.evidence_id` in place;
    * ids are assigned from the sorted position, not from input order;
    * dedup is order-insensitive by construction (the key is built from the
      sorted specialist list);
    * no clock, no RNG, no I/O.

Same inputs -> byte-identical output. `tests/unit/test_evidence_engine.py`
pins this.
"""

evidence/engine.py:38-50.

Property Mechanism Pinned by
source non-mutation model_copy(update={...}) at every write β€” _renumber (:430), payload merge (:385), agreement (:424) "the original keeps its uuid and the aggregated copy is a different object" (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.1)
ids from sorted position enumerate(capped, start=1) after sorted(...) Β§5.8
order-insensitive dedup key content-only; _record_agreement sorts names Β§5.5, Β§5.7
no clock / RNG / I/O the module imports only hashlib, collections.abc, dataclasses, typing evidence/engine.py:80-102

The dedup merge rebuilds rather than mutates, and the comment says why (evidence/engine.py:379-382):

Survivor's own keys win; the loser only fills gaps. Mutating the survivor in place would be fine here because merged holds the same object, but rebuilding keeps the "never mutate a source result" contract true even if the caller kept a reference.

5.5 Deduplication β€” identity is defined, not assumed

The module docstring (evidence/engine.py:52-68) states the principle:

Evidence.evidence_id is a uuid default, so it cannot be part of an identity key β€” two structurally identical items from two runs would never dedup. The key is therefore the content of the observation.

def _identity_key(item: Evidence) -> tuple[Any, ...]:
    coords = (
        tuple(_round(c) for c in item.coordinates)
        if item.coordinates is not None
        else None
    )
    return (
        item.type.value,
        item.source_specialist,
        item.coordinate_system.value if item.coordinate_system else None,
        coords,
        _round(item.score),
    )

evidence/engine.py:125-142.

Key component Included? Rationale
type yes a box and a mask are different claims even at the same coordinates
source_specialist yes keeps two specialists' identical claims as two items (Β§5.7)
coordinate_system yes the same four numbers in pixel and in geo are different claims
rounded coordinates yes disagreeing geometry β‡’ two claims; suppressing it would be a silent contradiction
rounded score yes different reliability is a different observation
evidence_id NO "including it would mean dedup never fires in production"
payload NO "a detail of the claim, not a different claim"
artifact_ref NO "two renderings of one claim are still one claim"

The docstring on the exclusions is the clearest statement of intent (evidence/engine.py:60-68):

payload and artifact_ref are deliberately EXCLUDED. Two items that agree on the same geolocation, one carrying a crs in its payload and one not, have made the same claim about the world; deduplicating them is correct, and the surviving item's payload is merged with the discarded one's so the crs is not lost. Conversely two items that share a type and score but disagree on coordinates are two different claims and are both kept β€” which matters, because suppressing a spatial disagreement would be exactly the silent contradiction the freeze forbids.

The rounding constant.

#: Rounding applied before an evidence item is turned into a dedup key. Six
#: decimals on normalized coordinates is ~1e-6 of the frame, far below any
#: meaningful spatial difference, but enough to absorb float noise from two
#: specialists rounding the same value two different ways.
_KEY_PRECISION = 6

def _round(value: float | None) -> float | None:
    return None if value is None else round(float(value), _KEY_PRECISION)

evidence/engine.py:114-122. _round is None-safe: a None coordinate stays None and does not become 0.0 in the identity key (unlike the sort key, Β§5.6, where a missing coordinate becomes 0.0 because a total order needs a value).

Payloads are MERGED on collision.

dropped += 1
# Survivor's own keys win; the loser only fills gaps.
payload = {**item.payload, **existing.payload}
if payload != existing.payload:
    updated = existing.model_copy(update={"payload": payload})
    seen[key] = updated
    merged[merged.index(existing)] = updated

evidence/engine.py:378-387. The spread order is {**loser, **survivor} β€” because Python dict unpacking is last-wins, existing.payload (the survivor's) overrides the loser's on key collision. This is the "survivor's own keys win; the loser only fills gaps" rule expressed directly in the merge order.

Behaviour Result
survivor had crs, loser did not survivor's crs survives
loser had crs, survivor did not crs is filled in from the loser β€” the information is not lost
both had crs with different values survivor's wins; the loser's value is dropped
no payload difference no model_copy is performed (the if payload != existing.payload guard)

docs/PHASE13_EVIDENCE_ENGINE.md Β§2.3 names the test that pins this: test_deduplication_ignores_payload_differences_and_merges_them.

What dedup does not merge. Because the key includes source_specialist, "only same-specialist repeats are merged here; two specialists agreeing is handled by _record_agreement, which preserves the second specialist's identity rather than discarding it" (evidence/engine.py:356-361).

5.6 The total sort order β€” and why each level exists

def _sort_key(item: Evidence) -> tuple[Any, ...]:
    if item.score is None:
        score_key: tuple[int, float] = (1, 0.0)
    else:
        score_key = (0, -float(item.score))
    coords = (
        tuple(_round(c) or 0.0 for c in item.coordinates)
        if item.coordinates is not None
        else ()
    )
    return (item.type.value, item.source_specialist, score_key, coords)

evidence/engine.py:165-190.

# Key Direction Why it exists (verbatim from the docstring)
1 type ascending "groups like with like, so a reader sees all the boxes together"
2 source_specialist ascending "stable and meaningful, unlike the uuid"
3 score descending "within a type, the strongest claim leads" β€” implemented as -score so the tuple stays uniformly ascending
4 coordinates ascending "the tie-breaker that makes the order total. Without it two items of the same type, specialist and score would sort by Python's stable-sort insertion order, which reintroduces input-order dependence."

docs/PHASE13_EVIDENCE_ENGINE.md Β§2.2 restates key 4 and names the pinning test:

Key 4 is not decoration. Without it, two items of the same type, specialist and score would sort by Python's stable-sort insertion order β€” which reintroduces exactly the input-order dependence the sort exists to remove. Pinned by test_equal_scores_order_deterministically_by_coordinates.

Unscored items sort last. "Items with no score sort AFTER items with a score: a missing score is not evidence of strength." (evidence/engine.py:178-179). The mechanism is a (0, …) / (1, …) discriminator in score_key: a scored item leads with 0, an unscored item leads with 1, so every scored item precedes every unscored item within the same (type, specialist) group.

flowchart LR
  subgraph G1["type=bounding_box, specialist=grounding"]
    direction TB
    A1["score 0.91"] --> A2["score 0.62"] --> A3["score None<br/><i>unscored last</i>"]
  end
  subgraph G2["type=statistic, specialist=vqa"]
    direction TB
    B1["score 0.74"]
  end
  G1 --> G2
  style A3 fill:#d2992222,stroke:#d29922

5.7 _record_agreement β€” corroboration, and why nothing is ever removed

The problem. source_specialist is a single str, not a set (the schema is frozen), so two specialists making the same claim are two different items under _identity_key and both survive. The docstring calls this "the correct conservative behaviour" (evidence/engine.py:70-77):

That is the correct conservative behaviour: collapsing them would require the engine to pick a winner, and the engine has no basis for that. What it does instead is keep both and record the disagreement-free agreement in the collection's sources and in each item's payload via _record_agreement.

The claim key β€” deliberately distinct from the identity key.

def _claim_key(item: Evidence) -> tuple[Any, ...]:
    """Identity of the *claim*, ignoring which specialist made it.

    Used only to detect corroboration. Deliberately distinct from
    `_identity_key`: that one answers "is this the same item", this one answers
    "are these two specialists saying the same thing about the world".
    """
    coords = (
        tuple(_round(c) for c in item.coordinates)
        if item.coordinates is not None
        else None
    )
    return (
        item.type.value,
        item.coordinate_system.value if item.coordinate_system else None,
        coords,
        _round(item.score),
    )

evidence/engine.py:145-162. The only difference from _identity_key is the absent source_specialist component.

flowchart TB
  I["_identity_key =<br/>(type, specialist, crs, coords, score)"] --> IQ{"same item?"}
  C["_claim_key =<br/>(type, crs, coords, score)"] --> CQ{"same claim about the world?"}
  IQ -->|"yes β†’ merge payloads"| DEDUP["_deduplicate"]
  CQ -->|"yes, β‰₯2 distinct specialists β†’ annotate"| AGR["_record_agreement"]
  style I fill:#1f6feb22,stroke:#1f6feb
  style C fill:#3fb95022,stroke:#3fb950

The annotation.

AGREEMENT_KEY = "corroborated_by"

...
out = list(items)
for indexes in groups.values():
    if len(indexes) < 2:
        continue
    names = sorted({items[i].source_specialist for i in indexes})
    if len(names) < 2:
        # Same specialist restating itself across the collection. The
        # identity key would normally have merged these, so reaching
        # here means the payloads differed; it is not corroboration.
        continue
    for i in indexes:
        item = out[i]
        others = [n for n in names if n != item.source_specialist]
        payload = {**item.payload, EvidenceCollection.AGREEMENT_KEY: others}
        out[i] = item.model_copy(update={"payload": payload})
return out

evidence/engine.py:391-425.

Why the payload, and not a schema field. The constant's own comment states it (evidence/engine.py:234-238):

Payload key under which _record_agreement stashes the specialists that made the identical claim. A single string field cannot hold a set, so the agreement is recorded in the payload β€” which is exactly what the payload is for: observations that support the evidence but are not the claim.

The contributing_specialists: list[str] field that would have made this first-class was proposed and rejected β€” docs/PHASE13_EVIDENCE_ENGINE.md Β§5 records the full ARCHITECTURE CHANGE entry. The three reasons:

# Reason
1 Adding a field is a frozen-contract change, and the freeze rule requires evidence that the current design is insufficient. No such evidence exists: identity comes from evidence_id; provenance is already served by EvidenceCollection.sources + Evidence.source_specialist; corroboration is recorded in payload["corroborated_by"].
2 Because source_specialist is a single string, two specialists making the identical claim are two items and both survive. Collapsing them would require the engine to pick a winner, and the engine has no basis for that.
3 Reversing this later is cheap: the engine already computes _claim_key(), which groups items by claim ignoring the specialist. Adopting the field would mean emitting one item per group instead of N.

Decision: proposal REJECTED; core/schemas.py unchanged. Adopted interim = source_specialist (single) + EvidenceCollection.sources (run-level provenance) + payload["corroborated_by"] (per-claim corroboration). "Revisit only if a downstream consumer needs corroboration as a first-class, queryable field."

Nothing is ever removed. The docstring is unambiguous (evidence/engine.py:401-403):

Nothing is removed: dropping a member would throw away a specialist's attribution, and the freeze does not permit silently discarding a specialist's evidence.

The two guards inside the loop are worth naming because they are the two ways a non-corroboration could masquerade as one:

Guard Condition Why
group size len(indexes) < 2 β†’ skip a lone item is not agreement
distinct specialists len(names) < 2 β†’ skip the same specialist restating itself is not corroboration; if this is reached, the payloads differed (otherwise _identity_key would have merged them)

corroborated_by is always the other specialists (others = [n for n in names if n != item.source_specialist]), so an item never lists itself.

Excluded from the digest. "The corroboration annotation in the payload IS excluded, because it is a derived observation about the collection, not part of the claim." (evidence/engine.py:692-693) β€” consistent with payload being excluded from _identity_key.

5.8 Renumbering to evidence_001…

ID_PREFIX = "evidence"

@staticmethod
def _renumber(item: Evidence, index: int) -> Evidence:
    """Give one item its canonical id, leaving the source untouched."""
    return item.model_copy(update={"evidence_id": f"{ID_PREFIX}_{index:03d}"})

evidence/engine.py:104-107, :427-430.

Property Value Reason
format evidence_001 … zero-padded to three digits
why three digits "so lexical sort matches numeric sort up to 999 items -- well past the evidence.max_items bound of 32" (:104-107)
assignment basis sorted position, not input order enumerate(capped, start=1) after sorted
scope run-local "Ids restart at 001 for each aggregate call β€” they are run-local citations, not global identities." (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.4)
source untouched model_copy the source keeps its ev_<uuid>

Why this matters at all (evidence/engine.py:21-25):

Evidence.evidence_id defaults to a random uuid, which is fine within one result but useless when the controller merges four specialists' evidence into the trace of a single run. Nothing can be cited. The engine renumbers to evidence_001, evidence_002, … so a downstream artefact (report, UI, audit) can reference one item deterministically.

5.9 The cap, and loss accounting that is never silent

DEFAULT_MAX_ITEMS = 32

evidence/engine.py:109-112, mirroring configs/base.yaml:

evidence:
  max_items: 32
  coordinate_system_default: normalized_0_1

EvidenceEngine.__init__ refuses a non-positive cap (evidence/engine.py:267-276):

if max_items < 1:
    raise ValueError(f"max_items must be >= 1, got {max_items}")

a cap of zero would silently discard every piece of evidence, which is a policy decision this engine has no business making on its own. (evidence/engine.py:258-261)

The accounting fields, all five:

Field Type Meaning Non-zero means
sources list[str] sorted specialist names that contributed anything, BEFORE the cap provenance β€” "Provenance is a property of the run, not of the surviving items."
dropped_duplicates int items merged into an existing claim normal and healthy β€” two specialists agreed
dropped_over_limit int items the cap discarded a warning β€” a specialist's evidence did not survive
total_before_limit int deduplicated count prior to capping the denominator for the two drop counts
truncated bool dropped_over_limit > 0 a boolean summary so a trace can branch without recomputing

evidence/engine.py:196-217; restated in docs/PHASE13_EVIDENCE_ENGINE.md Β§2.5.

Losses are counted, never silent.

And on why sources is recorded pre-cap:

sources is recorded pre-cap deliberately: provenance is a property of the run, not of the surviving items. A specialist whose evidence did not fit under the cap still ran, and the trace must still say so. (docs/PHASE13_EVIDENCE_ENGINE.md Β§2.5)

This is the same principle as core/config.py's refusal to silently truncate and evidence/confidence.py's refusal to fabricate a number: a bounded output must declare what the bound cost.

5.10 EvidenceCollection β€” the returned container

@dataclass(frozen=True)
class EvidenceCollection:
    items: list[Evidence] = field(default_factory=list)
    sources: list[str] = field(default_factory=list)
    dropped_duplicates: int = 0
    dropped_over_limit: int = 0
    total_before_limit: int = 0
    truncated: bool = False

evidence/engine.py:193-250. It is frozen=True β€” a dataclass decorator that blocks attribute rebinding on the collection, complementing the model_copy discipline that keeps the individual Evidence items unmutated.

Method Returns Purpose
__len__ int len(items)
__iter__ iterator iterate the items directly
ids() list[str] [item.evidence_id for item in self.items]
by_type(t) list[Evidence] filter by EvidenceType (identity comparison: item.type is evidence_type)
by_specialist(name) list[Evidence] filter by source_specialist
summary() dict[str, Any] the observable trace facts
AGREEMENT_KEY "corroborated_by" the payload key for corroboration

summary() β€” the trace-facing view.

def summary(self) -> dict[str, Any]:
    """Observable trace facts. No chain-of-thought, no interpretation."""
    return {
        "returned": len(self.items),
        "total_before_limit": self.total_before_limit,
        "dropped_duplicates": self.dropped_duplicates,
        "dropped_over_limit": self.dropped_over_limit,
        "truncated": self.truncated,
        "sources": list(self.sources),
        "types": sorted({item.type.value for item in self.items}),
    }

evidence/engine.py:240-250. Note types is a sorted set of the type values actually present in the surviving items β€” so a consumer can see which evidence kinds survived the cap without walking the list. The docstring's first line is a constraint, not a description: "No chain-of-thought, no interpretation." This mirrors docs/ARCHITECTURE_FREEZE.md Β§5's "Every result carries an observable execution trace. No chain-of-thought."

5.11 evidence_digest() β€” reproducibility as an assertion

def evidence_digest(collection: EvidenceCollection | Sequence[Evidence]) -> str:
    items = (
        collection.items
        if isinstance(collection, EvidenceCollection)
        else list(collection)
    )
    hasher = hashlib.sha256()
    for item in items:
        hasher.update(repr(_identity_key(item)).encode("utf-8"))
        hasher.update(b"\n")
    return hasher.hexdigest()

evidence/engine.py:680-704.

Design choice Consequence
built from _identity_key, in order changes when the claims change
not from evidence_id does not change when a specialist restates the same claim under a new uuid
not from payload payload is a detail of the claim, not the claim
repr(...) + b"\n" separator unambiguous framing; no two distinct sequences can hash the same
corroboration annotation excluded "it is a derived observation about the collection, not part of the claim"

The docstring states the contract it enables (evidence/engine.py:687-690):

That is what makes it useful as a reproducibility assertion:

aggregate(inputs_a) is reproducible iff digest(a) == digest(b)

against the same input, even across processes where the uuid defaults differ.

5.12 The canonical vocabulary: evidence_type_for

@staticmethod
def evidence_type_for(result: SpecialistResult) -> EvidenceType:
    from core.schemas import Task

    mapping = {
        Task.GROUNDING: EvidenceType.BOUNDING_BOX,
        Task.CHANGE: EvidenceType.CHANGE_MAP,
        Task.OPTICAL_SAR: EvidenceType.JOINT_FEATURE_REGION,
        Task.VQA: EvidenceType.STATISTIC,
        Task.CAPTION: EvidenceType.STATISTIC,
        Task.UNSUPPORTED: EvidenceType.STATISTIC,
    }
    return mapping.get(result.task, EvidenceType.STATISTIC)

evidence/engine.py:434-458.

This is the ONE place that knows the mapping from specialist task to evidence vocabulary, so no specialist needs to branch on it. Note it returns the primary type; a specialist's own produce_evidence emits its real per-artefact types, and those are preserved by aggregate.

Task Primary EvidenceType Why
grounding BOUNDING_BOX the task is localisation
change CHANGE_MAP the task is a change map
optical_sar JOINT_FEATURE_REGION the claim lives in CROMA's joint embedding space
vqa STATISTIC an answer is a measured scalar, not geometry
caption STATISTIC as above
unsupported STATISTIC there is no claim to localise
(fallback) STATISTIC .get(..., STATISTIC) β€” an unknown task degrades to a non-spatial claim rather than a wrong spatial one

change_vqa is absent from the mapping. Task.CHANGE_VQA exists in the enum (core/schemas.py:46) but has no row, so it falls through to STATISTIC β€” which is correct: change-VQA returns a short answer, not a map. The docstring's note that this returns the primary type matters here: a change-VQA result may still carry CHANGE_MAP evidence emitted by the detector it shares, and aggregate preserves that.

5.13 The three evidence_from_* constructors

These exist so the controller does not hand-build Evidence objects β€” the module docstring calls them "Convenience for the controller when it needs evidence for a box that a specialist emitted but did not itself describe."

evidence_from_box

def evidence_from_box(self, box, *, source_specialist, payload=None, asset_ref=None) -> Evidence:
    return Evidence(
        type=EvidenceType.BOUNDING_BOX,
        source_specialist=source_specialist,
        coordinate_system=box.coordinate_system,
        coordinates=[box.x1, box.y1, box.x2, box.y2],
        score=box.score,
        artifact_ref=asset_ref,
        payload={"label": box.label, **(payload or {})},
    )

evidence/engine.py:462-485.

The coordinate_system is copied from the box, never assumed β€” core.schemas requires it and Evidence rejects spatial coordinates without one.

The coordinates are flattened to [x1, y1, x2, y2] β€” matching the flat-Box shape that docs/PHASE19_FINAL_HARDENING.md Β§3.6 records as a corrected documentation defect.

evidence_from_region β€” a three-way branch

evidence/engine.py:487-540. A region may carry a mask, a box, both, or neither, and each case maps to a different evidence type:

flowchart TB
  R["Region"] --> M{"region.mask_ref?"}
  M -->|yes| MASK["EvidenceType.MASK<br/>artifact_ref = asset_ref or mask_ref<br/>coords = box if box else None"]
  M -->|no| B{"region.box?"}
  B -->|yes| BOX["EvidenceType.BOUNDING_BOX<br/>coords from box<br/>score = region.score or box.score"]
  B -->|no| STAT["EvidenceType.STATISTIC<br/>no coordinates<br/>score = region.score"]
  style MASK fill:#1f6feb22,stroke:#1f6feb
  style BOX fill:#1f6feb22,stroke:#1f6feb
  style STAT fill:#d2992222,stroke:#d29922

A region carrying a mask reference becomes MASK evidence; one with only a box becomes BOUNDING_BOX evidence. A region with neither is reported as a STATISTIC rather than a spatial claim, because there is no geometry to prove. (evidence/engine.py:496-500)

Two details worth naming:

  • The MASK branch's coordinates are None when region.box is None β€” a mask without a bounding box is legal, and the validator (Β§4) only fires when coordinates are present.
  • The BOUNDING_BOX branch reads coordinate_system from region.box.coordinate_system (the box's own frame), not from region.coordinate_system. The frame of the geometry is the frame of the geometry.

evidence_from_geospatial β€” may return None

def evidence_from_geospatial(self, geo, *, source_specialist, bounds=None, score=None) -> Evidence | None:
    if not geo.has_crs and not geo.crs:
        return None
    return Evidence(
        type=EvidenceType.GEOLOCATION,
        source_specialist=source_specialist,
        coordinate_system=CoordinateSystem.GEO if bounds else None,
        coordinates=list(bounds) if bounds else None,
        score=score,
        payload={
            "crs": geo.crs,
            "bounds": list(geo.bounds) if geo.bounds else None,
            "width": geo.width,
            "height": geo.height,
            "is_georeferenced": geo.is_georeferenced,
        },
    )

evidence/engine.py:542-572.

Returns None when there is no CRS to report. Emitting a GEOLOCATION item without a CRS would assert a placement the data does not support, which is the same failure mode the grounding degenerate-box guard exists to prevent.

This is the honesty rule applied to geometry: a claim that cannot be supported is not emitted, rather than emitted with a placeholder. Note the guard is not geo.has_crs and not geo.crs β€” either signal is sufficient, so a raster that reports a CRS string without the has_crs flag still produces evidence.

6. confidence_for() β€” calibrating through the engine's artifact

def confidence_for(
    self,
    result: SpecialistResult | ConfidenceBreakdown,
    *,
    extra_components: dict[str, float] | None = None,
    degraded: bool | None = None,
    degradation_reason: str | None = None,
) -> ConfidenceBreakdown:

evidence/engine.py:576-637. It accepts either form so "a caller that already has the pieces does not rebuild a SpecialistResult to use it."

The degradation rule β€” the specialist's verdict WINS.

The specialist's own degradation verdict WINS unless explicitly overridden: it knows things the engine does not (a zero-shot fallback, a failed input-quality gate), and overwriting degraded=False here would launder a degraded result into a confident-looking one. (evidence/engine.py:589-592)

Situation Result
degraded=None, specialist said degraded=True stays True
degraded=None, specialist said degraded=False stays False
degraded=True explicitly True, and if the specialist had not flagged it, the reason is sourced from result.warnings[0] β€” "so the reason is a fact rather than a restatement of the boolean" (:618-622)
degraded=True and no reason and no warnings falls back to the literal string "aggregate degraded" (:628-629)
degradation_reason explicitly passed that string is used verbatim

calibrate_result() (evidence/engine.py:639-643) is the narrower entry point:

def calibrate_result(self, result: SpecialistResult) -> ConfidenceBreakdown:
    """Calibrate a result's existing breakdown, preserving its provenance."""
    return calibrate_result(result.confidence, self.calibration)

from_config() (evidence/engine.py:647-665):

@classmethod
def from_config(cls, config, *, calibration=None, base_dir=None) -> "EvidenceEngine":
    return cls(
        max_items=int(config.get("evidence.max_items", DEFAULT_MAX_ITEMS)),
        calibration=calibration,
    )

calibration is explicit rather than auto-loaded: whether a fitted artifact exists is an operational fact the caller may know better than the config file does, and silently loading one would make the engine's behaviour depend on filesystem state.

The base_dir parameter is accepted for signature symmetry with evidence.confidence.load_calibration but is not used by this classmethod β€” the engine does not resolve the artifact itself.

Module-level shorthand (evidence/engine.py:668-677):

def aggregate_evidence(results, *, max_items=DEFAULT_MAX_ITEMS, calibration=None) -> EvidenceCollection:
    return EvidenceEngine(max_items=max_items, calibration=calibration).aggregate(results)

Part C β€” The confidence system

7. The honesty rule

evidence/confidence.py:12-30 is the reason the module exists at all. It is worth quoting in full because it is the single most important design statement in this subsystem:

THE HONESTY RULE (the reason this module exists at all)

A calibration that claims to be fitted when it is not is a FALSE CLAIM OF RELIABILITY β€” strictly worse than reporting nothing, because a downstream consumer will trust it. So:

no fitted artifact  ->  pass the raw score through unchanged,
                        set `method="uncalibrated"`,
                        leave `calibrated=None`

calibrated=None is what makes it honest: ConfidenceBreakdown.value then returns raw, and any consumer that wants to distinguish "we calibrated this" from "we did not" can read method or calibrated is None. We never fill in a plausible-looking number to make a schema field look complete.

7.1 The three states

docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1 tabulates the decision:

State calibrated method value
no fitted artifact None "uncalibrated" falls back to raw
fitted artifact, T β‰  1 mapped value "temperature_scaling" calibrated
fitted artifact, T = 1 None "uncalibrated" falls back to raw

The method strings are module constants (evidence/confidence.py:73-78):

METHOD_TEMPERATURE = "temperature_scaling"
METHOD_UNCALIBRATED = "uncalibrated"

Method string reported whenever no mapping was applied. Matches the literal the existing specialists already emit, so trace consumers see one vocabulary.

That last clause is a real constraint: VLMSpecialist._confidence_for and GroundingSpecialist._confidence_for already emit method="uncalibrated", calibrated=None independently. The engine generalises that behaviour rather than introducing a second vocabulary (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1).

7.2 ConfidenceBreakdown β€” the schema

class ConfidenceBreakdown(BaseModel):
    """Never an LLM utterance. Always measurable signals (plan section 26)."""
    model_config = ConfigDict(extra="forbid")

    raw: float = Field(ge=0.0, le=1.0)
    calibrated: float | None = Field(default=None, ge=0.0, le=1.0)
    method: str = "uncalibrated"
    components: dict[str, float] = Field(default_factory=dict)
    degraded: bool = False
    degradation_reason: str | None = None

    @property
    def value(self) -> float:
        return self.calibrated if self.calibrated is not None else self.raw

core/schemas.py:259-273.

Field Type Note
raw float, [0,1] the specialist's hand-weighted score β€” not a probability
calibrated float | None, [0,1] None means no mapping was applied β€” the honesty signal
method str, default "uncalibrated" one of the two constants
components dict[str, float] diagnostic signals; docs/API_CONTRACT.md:682 β€” "Diagnostic only β€” do not compute a confidence from it"
degraded bool "Whether this confidence should be trusted"
degradation_reason str | None why, when degraded

value is a Python @property and is NOT serialised. This is the trap docs/API_CONTRACT.md:571 warns about:

result.confidence.value | NOT a JSON field. It is a Python @property on ConfidenceBreakdown and is not serialised (verified: model_dump() yields only calibrated, components, degradation_reason, degraded, method, raw). To get the number the user should see, read calibrated if it is non-null, otherwise raw.

docs/STEP7_BACKEND_CHAIN_REPORT.md Β§D verifies this against the real model: "confidence.value is not serialised β€” the six-key set matches the contract's parsed claim exactly."

7.3 The four rules the frontend must follow

docs/API_CONTRACT.md:686-694:

  1. Display calibrated when it is not null; otherwise display raw.
  2. Display method next to the value. temperature_scaling means a fitted correction was applied; uncalibrated means it was not.
  3. Never present a confidence as a percentage without its method. A raw 0.99 and a calibrated 0.99 do not mean the same thing.
  4. When degraded is true, show degradation_reason. Confidence that is degraded is not a quality signal.

8. Temperature scaling β€” the derivation

8.1 Why the transform is needed at all

evidence/confidence.py:5-9:

Specialists produce a raw reliability score from signals they can actually point at (max_objectness, input-quality autocorrelation, used_head, …). That raw score is not a probability: it is a hand-weighted sum, and a hand-weighted sum is not calibrated by construction. This module is the one place that maps raw -> calibrated.

The signals named in docs/PHASE13_EVIDENCE_ENGINE.md Β§3.1 are max_objectness, top_mean_objectness, score_contrast, used_head, and input-quality autocorrelation. The plan (Β§26) names per-task signals too: for grounding, *"box confidence + text/image similarity

  • augmentation consistency"*; for change, "mean pixel probability + component stability + registration quality"; for optical-SAR, "fusion classifier margin + optical confidence + SAR confidence + cross-modal agreement".

8.2 The formula

logit(z)   = log(z / (1 - z))
calibrated = sigmoid(logit(z) / T)

evidence/confidence.py:39-47:

The standard formulation maps a logit through sigmoid(logit / T). Our raw scores are already in (0, 1), so they are read as probabilities and converted to log-odds first.

def apply(self, raw: float) -> float:
    """Map a raw score in [0, 1] through the fitted temperature."""
    return _clamp01(_sigmoid(_logit(float(raw)) / self.temperature))

evidence/confidence.py:165-167. Three operations in order: _logit β†’ divide by T β†’ _sigmoid β†’ _clamp01.

T Effect Verified
T > 1 softens β€” pulls scores toward the middle docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3
T < 1 sharpens β€” pushes scores away from the middle docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3; measured: calibrate(0.87, T=0.9772732) β†’ 0.8749186077809417
T = 1 identity handled as uncalibrated, Β§8.3

The transform is monotonic, so calibration respaces scores without reordering them β€” pinned by test_temperature_scaling_is_monotonic (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3).

8.3 _is_effective β€” why a "fitted" T of exactly 1.0 is reported as uncalibrated

#: How close to 1.0 a fitted temperature must land before it is treated as
#: "does nothing". Bit-for-bit 1.0 is the honest test; this tolerance absorbs
#: float round-trip through JSON without accepting a real transform.
_IDENTITY_TOLERANCE = 1e-6

def _is_effective(temperature: float) -> bool:
    """Does this temperature actually transform anything?

    A temperature within `_IDENTITY_TOLERANCE` of 1.0 is the identity map. It is
    treated as "not fitted" so the honest report wins: an artifact that carries
    T=1.0 has learned nothing, and reporting `method="temperature_scaling"` for
    it would assert a correction that was never made.
    """
    return abs(temperature - 1.0) > _IDENTITY_TOLERANCE

evidence/confidence.py:90-93, :118-126.

Verified empirically (read-only check against the shipped module):

Input _is_effective Result
0.9772731820958189 (the fitted T) True a real transform is applied
1.0 False reported as uncalibrated

Why the tolerance and not != 1.0. The comment gives the reason: "Bit-for-bit 1.0 is the honest test; this tolerance absorbs float round-trip through JSON without accepting a real transform." A JSON round-trip of 1.0 can land at 1.0000000000000002; an exact != 1.0 test would then accept a transform that does nothing. 1e-6 is wide enough to absorb the round-trip and far too narrow to swallow any real fitted temperature.

The identity branch records why (evidence/confidence.py:299-303):

if calibration is not None:
    # A real artifact was supplied but it is the identity map. Say so
    # rather than dropping the information on the floor.
    components["calibration_identity"] = 1.0
    components.update(calibration.components)

Measured components for T = 1.0: {'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}.

docs/PHASE13_EVIDENCE_ENGINE.md Β§3.2 frames this as the same failure mode as Β§3.1: "This is a small case the brief did not name explicitly; it is the same false-claim failure mode as Β§3.1 and is handled by the same principle."

8.4 _EPS β€” endpoint clamping so a hard 0.0 / 1.0 cannot become NaN

#: Scores are clamped to this closed interval before the log-odds transform.
#: `logit(0)` and `logit(1)` are infinite; clamping at the boundary keeps the
#: mapping finite and keeps a hard 0.0 / 1.0 from becoming a NaN.
_EPS = 1e-6

def _logit(p: float) -> float:
    """Log-odds of a probability, with the endpoints clamped to stay finite."""
    clamped = min(1.0 - _EPS, max(_EPS, p))
    return math.log(clamped / (1.0 - clamped))

evidence/confidence.py:80-83, :108-111.

Input z Unclamped logit(z) With _EPS clamp Result
0.0 log(0) = -inf logit(1e-6) β‰ˆ -13.8155 finite
1.0 log(1/0) = +inf logit(1 - 1e-6) β‰ˆ +13.8155 finite
0.5 0.0 0.0 unchanged

Pinned by test_endpoint_scores_do_not_produce_nan (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.3).

8.5 _sigmoid β€” numerically stable, and why the branch exists

def _sigmoid(x: float) -> float:
    """Numerically stable logistic function.

    The naive form overflows for large negative `x`; the branch below keeps
    `exp`'s argument non-positive so the result is finite for every input.
    """
    if x >= 0.0:
        return 1.0 / (1.0 + math.exp(-x))
    z = math.exp(x)
    return z / (1.0 + z)

evidence/confidence.py:96-105. For x >= 0 the exp argument is -x <= 0; for x < 0 the exp argument is x < 0. In both branches math.exp receives a non-positive argument, so it can never overflow. The two forms are algebraically identical (1/(1+e^-x) ≑ e^x/(1+e^x)) but only one of them is safe on each side of zero.

def _clamp01(x: float) -> float:
    return min(1.0, max(0.0, x))

evidence/confidence.py:114-115. The final guard: sigmoid returns (0, 1) mathematically, but float rounding can produce exactly 0.0 or 1.0, and _clamp01 keeps the value inside the schema's ge=0.0, le=1.0 bound so Pydantic never rejects an output of this module.

8.6 TemperatureCalibration β€” the artifact object

@dataclass(frozen=True)
class TemperatureCalibration:
    temperature: float
    fitted_on: str | None = None
    artifact: str | None = None
    n_samples: int | None = None

evidence/confidence.py:129-150.

Field Type Meaning (docstring)
temperature float "the fitted scalar. T > 1 softens …, T < 1 sharpens. Validated on construction."
fitted_on str | None "free-text provenance, e.g. "valid" or a split hash. Carried into the breakdown components so a result can be traced back to the artifact that shaped it."
artifact str | None "path/identifier of the source file, for the same reason."
n_samples int | None "how many validation samples the fit used, when known."

Range validation (evidence/confidence.py:85-88, :152-159):

_MIN_TEMPERATURE = 1e-3
_MAX_TEMPERATURE = 1e3

def __post_init__(self) -> None:
    t = float(self.temperature)
    if not math.isfinite(t) or not _MIN_TEMPERATURE <= t <= _MAX_TEMPERATURE:
        raise ValueError(
            f"temperature must be finite and in "
            f"[{_MIN_TEMPERATURE}, {_MAX_TEMPERATURE}], got {self.temperature!r}"
        )
    object.__setattr__(self, "temperature", t)

A temperature of zero or below is not a calibration; it is a division error. (evidence/confidence.py:85-88)

The check is math.isfinite(t) and the range β€” so NaN and Β±inf are both rejected, and object.__setattr__ is used because the dataclass is frozen=True.

Two methods:

Method Returns Purpose
is_effective() bool _is_effective(self.temperature)
apply(raw) float _clamp01(_sigmoid(_logit(raw) / T))
components (property) dict[str, float] {"temperature": T} plus {"calibration_samples": n} when known
describe() dict[str, Any] {method, temperature, effective, fitted_on, artifact, n_samples}

from_dict() β€” accepts two shapes (evidence/confidence.py:187-219):

body: Any = payload
if isinstance(payload.get("temperature_scaling"), dict):
    body = payload["temperature_scaling"]
elif isinstance(payload.get("calibration"), dict):
    body = payload["calibration"]

if not isinstance(body, dict) or "temperature" not in body:
    raise ValueError(
        "calibration artifact has no 'temperature' field "
        f"(source: {source or '<dict>'})"
    )

Accepts either the flat form ({"temperature": 1.4, ...}) or a nested {"temperature_scaling": {...}} form, because the artifact format is owned by training/calibration/ and this loader should not be the thing that decides it.

fitted_on reads body.get("fitted_on") or body.get("split"); artifact prefers the source path over body.get("artifact"); n_samples is coerced to int when present. The shipped artifact uses the nested form β€” artifacts/calibration_v001.json has a top-level "temperature_scaling": {"fitted_on": "Val", "n_samples": 16441, "temperature": 0.9772731820958189} block. This is exactly why from_dict accepts it.

from_json() β€” required semantics (evidence/confidence.py:221-252):

Condition required=False (default) required=True
file absent returns None raises FileNotFoundError
file present, valid JSON returns the object returns the object
file present, malformed JSON raises ValueError raises ValueError

ValueError: the file exists and is malformed (never swallowed: a corrupt artifact silently treated as "no artifact" would hide an operational defect).

This is the same asymmetry as the evidence cap and the config validation: absent is a legitimate state; corrupt is a defect and must be loud.

8.7 calibrate() β€” the single entry point

def calibrate(
    raw: float,
    calibration: TemperatureCalibration | None,
    *,
    extra_components: dict[str, float] | None = None,
    degraded: bool = False,
    degradation_reason: str | None = None,
) -> ConfidenceBreakdown:

evidence/confidence.py:255-323. This is the single entry point used by EvidenceEngine.confidence_for.

Non-finite input RAISES (evidence/confidence.py:289-292):

value = float(raw)
if not math.isfinite(value):
    raise ValueError(f"raw confidence must be finite, got {raw!r}")
clamped = _clamp01(value)

Refusing is deliberate: a non-finite score is a bug upstream, and clamping it to 0.5 would invent a measurement. (evidence/confidence.py:284-287)

Note the asymmetry with _logit's endpoint clamping: an infinite input is refused, but a finite 0.0 or 1.0 is accepted and clamped inside the transform. The distinction is that 0.0 is a real measurement (a specialist is certain it found nothing) while NaN is not a measurement at all.

The honest path (evidence/confidence.py:296-311):

if calibration is None or not calibration.is_effective():
    components["calibrated_applied"] = 0.0
    if calibration is not None:
        components["calibration_identity"] = 1.0
        components.update(calibration.components)
    return ConfidenceBreakdown(
        raw=clamped,
        calibrated=None,
        method=METHOD_UNCALIBRATED,
        components=components,
        degraded=degraded,
        degradation_reason=degradation_reason,
    )

The fitted path (evidence/confidence.py:313-323):

components["calibrated_applied"] = 1.0
components.update(calibration.components)

return ConfidenceBreakdown(
    raw=clamped,
    calibrated=calibration.apply(clamped),
    method=METHOD_TEMPERATURE,
    components=components,
    degraded=degraded,
    degradation_reason=degradation_reason,
)

Measured component blocks (read-only verification against the shipped module):

Case components
calibrate(0.87, T=0.9772731820958189) {'calibrated_applied': 1.0, 'temperature': 0.9772731820958189, 'calibration_samples': 16441.0}
calibrate(0.87, T=1.0) {'calibrated_applied': 0.0, 'calibration_identity': 1.0, 'temperature': 1.0}
calibrate(0.87, None) {'calibrated_applied': 0.0}

The calibrated_applied flag is a float, not a bool, because components is typed dict[str, float]. 0.0 / 1.0 is the encoding. raw is retained in every case β€” the breakdown always carries both the input and the output, so nothing is destroyed by calibrating.

8.8 calibrate_result() β€” preserve the specialist's judgement

def calibrate_result(breakdown, calibration) -> ConfidenceBreakdown:
    return calibrate(
        breakdown.raw,
        calibration,
        extra_components=dict(breakdown.components),
        degraded=breakdown.degraded,
        degradation_reason=breakdown.degradation_reason,
    )

evidence/confidence.py:326-346.

The specialist's own components and degradation reason are carried through untouched: this function adds the calibration, it does not re-judge the specialist's measurement.

Pinned by test_specialist_degradation_survives_calibration (docs/PHASE13_EVIDENCE_ENGINE.md Β§3.5): "A degraded=True fallback result stays degraded after calibration, with its degradation_reason intact."

9. load_calibration() β€” the ordered candidate search

9.1 The frozen keys it reads

confidence:
  temperature_scaling: true
  calibration_file: calibration_v001.json

configs/base.yaml:231-233.

if not bool(config.get("confidence.temperature_scaling", False)):
    return None

filename = config.get("confidence.calibration_file")
if not filename:
    return None

evidence/confidence.py:396-401. It returns None β€” never a fabricated default β€” when:

Condition Returns
the master switch is off None
no filename is configured None
the artifact is absent from every candidate None
the artifact is present but malformed raises ValueError

Every one of those is a legitimate deployment state (an HF Space may ship without the fitted artifact), and each degrades to the pass-through. (evidence/confidence.py:361-364)

9.2 The ordered search

name = str(filename)
candidates: list[Path] = []

if base_dir is not None:
    # An explicit base_dir is honoured first, but a REPO-anchored candidate
    # is still tried: several call sites in this repo pass "." meaning "the
    # repo", which only works when CWD happens to be the repo root.
    candidates.append(Path(base_dir) / name)

from core.config import REPO_ROOT  # local import: keeps this module cheap

if base_dir is not None:
    candidates.append(Path(REPO_ROOT) / name)
candidates.append(Path(REPO_ROOT) / "artifacts" / name)
candidates.append(Path(REPO_ROOT) / "configs" / name)

for candidate in candidates:
    if candidate.exists():
        return TemperatureCalibration.from_json(candidate, required=True)

# Nothing found: the master switch is on but no artifact ships. Degrade.
return None

evidence/confidence.py:403-424.

flowchart TB
  S["load_calibration(config, base_dir=?)"] --> SW{"confidence.temperature_scaling?"}
  SW -->|false| N1["return None"]
  SW -->|true| FN{"calibration_file set?"}
  FN -->|no| N2["return None"]
  FN -->|yes| C1{"base_dir given?"}
  C1 -->|yes| A["1. base_dir / name"]
  A --> A2{"exists?"}
  C1 -->|no| B
  A2 -->|yes| LOAD["from_json(required=True)"]
  A2 -->|no| B["2. REPO_ROOT / name<br/><i>only when base_dir was given</i>"]
  B --> B2{"exists?"}
  B2 -->|yes| LOAD
  B2 -->|no| C["3. REPO_ROOT / artifacts / name<br/><b>the real home</b>"]
  C --> C2{"exists?"}
  C2 -->|yes| LOAD
  C2 -->|no| D["4. REPO_ROOT / configs / name"]
  D --> D2{"exists?"}
  D2 -->|yes| LOAD
  D2 -->|no| N3["return None β€” honest degradation"]
  style C fill:#3fb95022,stroke:#3fb950
  style LOAD fill:#1f6feb22,stroke:#1f6feb

The exact candidate list, by input:

base_dir candidate 1 candidate 2 candidate 3 candidate 4
None β€” β€” REPO_ROOT/artifacts/name REPO_ROOT/configs/name
"." ./name REPO_ROOT/name REPO_ROOT/artifacts/name REPO_ROOT/configs/name
"configs" configs/name REPO_ROOT/name REPO_ROOT/artifacts/name REPO_ROOT/configs/name
"artifacts/calibration/confidence_calibration_v001" that dir / name REPO_ROOT/name REPO_ROOT/artifacts/name REPO_ROOT/configs/name

The search is ordered and deterministic: the first existing candidate wins, and no candidate existing means None (honest degradation, not an error). (evidence/confidence.py:388-389)

9.3 The silent-failure shape this fixes

evidence/confidence.py:366-386 documents the defect precisely:

base_dir is resolved in this order (added 2026-09-22):

  1. an explicit base_dir argument, when the caller passes one;
  2. configs/ under the repository root, but only if the artifact is actually there β€” it never is by default, so this branch exists only to keep an existing caller that parks the artifact beside the config working;
  3. <repo root>/artifacts/ β€” the artifact's real home, and the location scripts/fit_calibration.py writes to.

Step 3 is the important one and it is why this function changed. The frozen config names a bare filename (calibration_v001.json), while the docstring examples in this repository's own docs pass base_dir="." AND base_dir="configs" β€” two different directories, neither of which is artifacts/. A deployment following either example would silently resolve to a nonexistent path, degrade to method="uncalibrated", and report a fitted artifact as never having been fitted. That is precisely the silent-failure shape this project's provenance discipline exists to prevent, so the search now anchors to the repository root using the same REPO_ROOT that app/serving.py uses for the change checkpoint.

The failure shape, stated as a table:

Step Pre-fix behaviour Post-fix behaviour
caller passes base_dir="." from a non-repo CWD resolves to $CWD/calibration_v001.json β†’ absent β†’ None β†’ uncalibrated candidate 1 fails, candidate 3 succeeds β†’ the artifact loads
caller passes base_dir="configs" resolves to configs/calibration_v001.json β†’ absent β†’ None β†’ uncalibrated candidate 1 fails, candidate 3 succeeds β†’ the artifact loads
caller passes nothing resolved to $CWD/... (base_dir defaulted to Path.cwd() pre-fix) candidate 3 succeeds β†’ the artifact loads
the artifact is genuinely absent None None (unchanged)
the artifact is present but corrupt depends on the path resolved raises β€” never swallowed

The critical property is the symptom: the pre-fix failure was silent and indistinguishable from "no artifact was ever fitted". A fitted model reported itself as unfitted. That is the same class of defect as a fabricated confidence number β€” it makes the system lie about its own state.

Verified against the shipped repo (read-only execution):

Call Resolved artifact
load_calibration(config) …/satquery-ai/artifacts/calibration_v001.json
load_calibration(config, base_dir="configs") …/satquery-ai/artifacts/calibration_v001.json
load_calibration(config, base_dir=".") …/satquery-ai/artifacts/calibration_v001.json

All three resolve to the same file via candidate 3. The artifact's own consumer_contract.resolution field recommends load_calibration(config, base_dir='configs') β€” which now works, and works for the reason the fix was made.

The older record, and its status. docs/PHASE13_EVIDENCE_ENGINE.md Β§4.2 was written before the fix and describes the pre-fix behaviour (base_dir defaults to Path.cwd(); "the controller must pass base_dir explicitly"; "Current behaviour on the real config, for the record: load_calibration returns None, because no artifact has been fitted yet"). Two things changed since: the search was repo-anchored (2026-09-22), and artifacts/calibration_v001.json now exists. The current measured behaviour is the table above. The PHASE13 Β§4.1 expectation of artifacts/calibration/confidence_calibration_v001/calibration_v001.json did not hold β€” the artifact lives at artifacts/calibration_v001.json β€” which is fine because the bare filename plus the artifacts/ candidate resolves it.

10. The MEASURED calibration result β€” a negative result, reported as one

10.1 The artifact

artifacts/calibration_v001.json is the fitted artifact. Its top-level keys: consumer_contract, created_utc, fit_diagnostics, metrics, provenance, reliability_diagram, schema, scope, temperature_scaling, type_mask_applied.

fit_diagnostics:

Field Value
temperature 0.9772731820958189
log_temperature -0.022989052824434128
iterations 200
n_samples 16441
nll_before 0.6897411093435998
nll_after 0.689630751845387
nll_improvement 0.0001103574982127542
effective true
hit_bound false

Note log_temperature = log(0.9772731820958189) = -0.022989… β€” the fitter optimised in log-space, and hit_bound: false means the optimum was interior, not at a clamp.

metrics:

Field Value
ece_before 0.013755
ece_after 0.014929
ece_improvement -0.001174 (WORSE)
nll_before 0.689741
nll_after 0.689631
nll_improvement 0.00011
n_bins 15
n_classes 19
n_samples 16441

provenance:

Field Value
checkpoint_path artifacts/change_vqa/run/head.pt
checkpoint_sha256 cfae5e43b97ca930f568dc5b8ae4f36b24e9ff717af226159802206ffd63a82a
config_hash 78f1e3700da15aa1
dataset_id cdvqa
feature_spec change_feat_v1
fitted_on Val
held_out_splits_excluded ["Test", "Test2"]
method temperature_scaling
objective mean_negative_log_likelihood
optimizer golden_section_on_log_temperature
space multiclass_logits
iterations 200
n_classes 19
n_samples 16441

scope β€” what this artifact calibrates, and what it does not:

This temperature calibrates the R-02 change-VQA head's answer confidence. Other specialists emit their own raw scores and are unaffected.

type_mask_applied: false.

10.2 The result, stated exactly

The raw softmax was already near-calibrated (ECE 0.013755) and temperature scaling made ECE very slightly worse (0.014929) while improving NLL marginally. This is a measurement, not a quality judgment, and it must not be described as scaling being "more accurate". β€” docs/STEP7_BACKEND_CHAIN_REPORT.md Β§7

docs/FINAL_DELIVERY_TODO.md:104 records it in the measured-metrics table as 0.013755 β†’ 0.014929 (worse) with status "measured (not an improvement)", sourced to artifacts/calibration_v001.json:26-35. docs/PHASE19_FINAL_HARDENING.md:377 records the same: "Not claimed as an improvement. … Nothing is calibrated in a deployed path".

Metric Before After Ξ” Reading
ECE (15 bins) 0.013755 0.014929 βˆ’0.001174 WORSE
NLL 0.689741 0.689631 +0.000110 marginally better

Why both numbers matter, and neither alone. The fitter's objective was mean_negative_log_likelihood, and NLL did improve β€” that is why T = 0.9773… was selected at all. But the metric a consumer reads as "is this probability trustworthy" is ECE, and ECE got worse. Reporting only NLL would present a fitted artifact as a success; reporting only ECE would hide why the value was chosen. Both are reported, and the headline is the ECE regression.

Why the reliability diagram is not a single score. The artifact's own note:

Equal-width bins over predicted-class confidence. ECE is bin-count sensitive and is not an aggregate score. β€” artifacts/calibration_v001.json, reliability_diagram.note

The diagram is a 15-bin list. The first two bins ([0.0, 0.0667) and [0.0667, 0.1333)) have count: 0 and null accuracy/confidence/gap β€” the model never predicts that low. The remaining thirteen bins are populated; the largest is [0.9333, 1.0] with count: 2814, accuracy: 0.969794, confidence: 0.966895, gap: 0.002899. The largest |gap| is -0.027817 at [0.6667, 0.7333) (count: 1494).

Bin [lo, hi) count accuracy confidence gap
0.0000–0.0667 0 β€” β€” β€”
0.0667–0.1333 0 β€” β€” β€”
0.1333–0.2000 8 0.25 0.188409 +0.061591
0.2000–0.2667 437 0.283753 0.241931 +0.041822
0.2667–0.3333 736 0.290761 0.299846 βˆ’0.009085
0.3333–0.4000 626 0.386581 0.367128 +0.019454
0.4000–0.4667 690 0.450725 0.434665 +0.016060
0.4667–0.5333 1380 0.534783 0.504472 +0.030310
0.5333–0.6000 1521 0.558185 0.565780 βˆ’0.007595
0.6000–0.6667 1395 0.624373 0.633333 βˆ’0.008961
0.6667–0.7333 1494 0.672691 0.700507 βˆ’0.027817
0.7333–0.8000 1611 0.742396 0.766718 βˆ’0.024322
0.8000–0.8667 1701 0.833039 0.834257 βˆ’0.001218
0.8667–0.9333 2028 0.892998 0.903143 βˆ’0.010145
0.9333–1.0000 2814 0.969794 0.966895 +0.002899

(values verbatim from artifacts/calibration_v001.json; the gap column is the artifact's own accuracy βˆ’ confidence)

10.3 A structural note on the transform, recorded rather than glossed

The artifact's consumer_contract states:

"applied_as": "sigmoid(logit(z) / T) for a scalar z; softmax(logits / T) for a distribution",
"class": "TemperatureCalibration",
"module": "evidence.confidence",
"read_keys": ["temperature", "fitted_on|split", "artifact", "n_samples"],
"resolution": "load_calibration(config, base_dir='configs')"

The artifact's provenance records space: "multiclass_logits" and n_classes: 19 β€” the fit happened over a 19-class softmax. The consumer class (evidence.confidence.TemperatureCalibration.apply) implements the scalar form: _clamp01(_sigmoid(_logit(raw) / T)). The class reads exactly the four keys the artifact declares (temperature, fitted_on|split, artifact, n_samples) and reads nothing else β€” the multiclass branch is a documented capability of the artifact format, not a code path in evidence/confidence.py. Both facts are stated here so the boundary is visible; no claim is made about which form a deployed path exercises, because docs/PHASE19_FINAL_HARDENING.md:377 records that "Nothing is calibrated in a deployed path".

10.4 The deployment state

Question Answer Source
Does a fitted artifact exist? yes β€” artifacts/calibration_v001.json filesystem
Does load_calibration(config) find it? yes β€” via candidate 3 verified read-only
Is T effective (non-identity)? yes β€” _is_effective(0.9772731820958189) == True verified read-only
Does it improve ECE? no β€” 0.013755 β†’ 0.014929 artifact metrics
Is it wired into a served path? OPEN / not claimed β€” "Nothing is calibrated in a deployed path" docs/PHASE19_FINAL_HARDENING.md:377
Which specialist does it cover? change_vqa only artifact scope

The artifact's config_hash is 78f1e3700da15aa1 β€” the same frozen hash as configs/base.yaml (Β§07). The artifact is keyed to the config that produced it, which is exactly the provenance link the config hash exists to provide.

11. What is deliberately absent from the confidence system

11.1 No LLM path

NO LLM PATH EXISTS HERE. There is deliberately no function that turns text into a number. Confidence is a measurement (freeze section 5: "No LLM-generated confidence"). The only inputs accepted are a float the specialist computed and a calibration artifact fitted on validation data. β€” evidence/confidence.py:32-37

docs/ARCHITECTURE_FREEZE.md Β§5 lists it as a non-negotiable: "No LLM-generated coordinates. No LLM-generated confidence." And the plan Β§26 opens with it: "No LLM-generated confidence."

11.2 No temperature fitter

A fitter is NOT included: Phase 13 fits T on validation data, and inventing one here would be exactly the kind of unfounded number this module refuses to emit. TemperatureCalibration is the consumer of a fitted T; the artifact loader reads whatever training/calibration/ produces. β€” evidence/confidence.py:57-61

The fitter lives in training/calibration/ (__init__.py, artifact.py, evaluate.py, fitter.py) and produced the artifact in Β§10. evidence/confidence.py deliberately does not contain one β€” the separation is what makes "the number came from a fit" a checkable claim.

11.3 No retrieval path

There is no artifact-serving endpoint in v1 (F-16, Β§2.3). This is why artifact_ref is permanently null and why the confidence module's artifact field carries a path for provenance, not a client-facing reference. TemperatureCalibration.artifact is used in describe() for operator diagnostics; it is not serialised into a client response.


Part D β€” The execution trace and the eight events

12. ExecutionTrace β€” observable facts only

class ExecutionTrace(BaseModel):
    model_config = ConfigDict(extra="forbid")

    run_id: str = Field(default_factory=lambda: _new_id("run"))
    schema_version: str = SCHEMA_VERSION
    task: Task | None = None
    query: str | None = None
    inputs: list[str] = Field(default_factory=list)
    modalities: list[Modality] = Field(default_factory=list)
    intent: Intent | None = None
    validation: dict[str, Any] = Field(default_factory=dict)
    workflow: list[str] = Field(default_factory=list)
    steps: list[TraceStep] = Field(default_factory=list)
    selected_models: list[ModelRef] = Field(default_factory=list)
    parameters: dict[str, Any] = Field(default_factory=dict)
    outputs: list[str] = Field(default_factory=list)
    confidence: ConfidenceBreakdown | None = None
    timings: dict[str, float] = Field(default_factory=dict)
    fallbacks: list[str] = Field(default_factory=list)
    errors: list[dict[str, Any]] = Field(default_factory=list)
    contradiction: bool = False
    config_hash: str | None = None
    started_at: str = Field(default_factory=_utcnow)
    finished_at: str | None = None

core/schemas.py:296-319. The section header above it reads: # Execution trace (observable facts only β€” never chain-of-thought) (core/schemas.py:277).

Field Type Carries
run_id str run_<12 hex>, generated per run
schema_version str SCHEMA_VERSION = "1.0"
task Task | None the routed task
query str | None the user's question, verbatim
inputs list[str] asset references
modalities list[Modality] inferred modalities
intent Intent | None the router's advisory output
validation dict[str, Any] input-validation facts (input_count, format, modality in the plan Β§27 shape)
workflow list[str] the planned step names
steps list[TraceStep] the state machine's visits
selected_models list[ModelRef] which models ran
parameters dict[str, Any] parameters used
outputs list[str] output kinds (bbox, answer, …)
confidence ConfidenceBreakdown | None the confidence breakdown
timings dict[str, float] measured durations
fallbacks list[str] which fallbacks fired
errors list[dict[str, Any]] error records
contradiction bool whether a contradiction was detected
config_hash str | None the frozen config identity β€” 78f1e3700da15aa1 for this revision
started_at / finished_at str ISO-8601 UTC

TraceStep (core/schemas.py:279-285):

class TraceStep(BaseModel):
    model_config = ConfigDict(extra="forbid")

    state: ControllerState
    started_at: str = Field(default_factory=_utcnow)
    duration_ms: float | None = None
    detail: dict[str, Any] = Field(default_factory=dict)

ModelRef (core/schemas.py:288-293):

class ModelRef(BaseModel):
    model_config = ConfigDict(extra="forbid")

    name: str
    revision: str | None = None
    role: str | None = None

ControllerState β€” the nine states (core/schemas.py:79-88):

# State Meaning
1 RECEIVE the request arrived
2 PARSE the query was parsed
3 VALIDATE inputs validated
4 PLAN the workflow was chosen
5 PREPROCESS tiling / normalisation
6 EXECUTE specialists ran
7 AGGREGATE evidence was aggregated
8 VERIFY confidence was computed
9 RESPOND the envelope was assembled

configs/base.yaml Β§agent.states declares exactly this list, in this order β€” so the config and the enum cannot drift.

12.1 There is no field for reasoning

ExecutionTrace has no reasoning, thought, rationale, or chain_of_thought field, and model_config = ConfigDict(extra="forbid") means one cannot be smuggled in. This is a structural guarantee, not a convention. docs/ARCHITECTURE_FREEZE.md Β§5:

Every result carries an observable execution trace. No chain-of-thought.

The plan Β§27 says the same at the end of its trace example: "No chain-of-thought. Only observable execution facts." The summary() methods in the evidence layer obey the same rule β€” EvidenceCollection.summary()'s docstring: "Observable trace facts. No chain-of-thought, no interpretation."

13. The eight execution events

The frontend's event protocol is the integration seam between the backend and the instrument UI. It is declared in frontend/assets/js/core.js:616-620:

SQ.EVENT_NAMES = [
  'QUERY_RECEIVED', 'QUERY_UNDERSTOOD', 'ROUTE_SELECTED',
  'SPECIALIST_STARTED', 'SPECIALIST_COMPLETED',
  'EVIDENCE_GENERATED', 'CONFIDENCE_COMPUTED', 'RESULT_ASSEMBLED'
];
# Event What it reports Drives the stage
1 QUERY_RECEIVED the query text, the scene seed, the AOI, the GSD, the date pair QUERY
2 QUERY_UNDERSTOOD the routed task, the parsed slots, the policy id UNDERSTAND
3 ROUTE_SELECTED the task, the specialist list, the rule trace, the policy version ROUTE
4 SPECIALIST_STARTED which component, its index/ordinal, its stage, its model, its device ANALYZE / GROUND
5 SPECIALIST_COMPLETED which component, its duration, its output kind (advances the active marker)
6 EVIDENCE_GENERATED the evidence regions, the count, the threshold, the registration RMSE EVIDENCE
7 CONFIDENCE_COMPUTED the reported/calibrated values, the method, the reliability bins CONFIDENCE
8 RESULT_ASSEMBLED the answer text, the task, the specialists, the evidence ids, the confidence, the provenance ANSWER

The stage rail is a separate, coarser list (frontend/assets/js/core.js:605-614):

SQ.STAGES = [
  { id: 'QUERY', label: 'QUERY' },
  { id: 'UNDERSTAND', label: 'UNDERSTAND' },
  { id: 'ROUTE', label: 'ROUTE' },
  { id: 'ANALYZE', label: 'ANALYZE' },
  { id: 'GROUND', label: 'GROUND' },
  { id: 'EVIDENCE', label: 'EVIDENCE' },
  { id: 'CONFIDENCE', label: 'CONFIDENCE' },
  { id: 'ANSWER', label: 'ANSWER' }
];

Eight stages for eight events β€” but not a 1:1 mapping: events 4 and 5 share the ANALYZE/GROUND stages, because a specialist's start and completion are two events about one stage visit.

13.1 The run engine is deliberately dumb

frontend/assets/js/core.js:8-11:

The run engine is deliberately dumb: it renders whatever events it receives. Nothing about the visuals depends on the events being synthetic. Swapping the mock driver for a websocket / SSE feed of the same event names is the entire integration surface.

/**
 * THE INTEGRATION SEAM.
 * Feed real execution events here β€” same names, same payload shapes β€” and
 * every visual state in the prototype updates identically.
 */
ingest: function (type, payload) { ... }

frontend/assets/js/core.js:737-742. ingest() is a switch over the eight names, each case updating state and calling emit(). emit() records {type, t, seq, payload} where t = performance.now() - t0 and seq is a monotonic counter (frontend/assets/js/core.js:709-718).

Two drivers exist, and they are not interchangeable:

Driver When Payload honesty
startMock(query, scene) preview only β€” when no file is selected synthetic; marked is-mock
runLive(query) in frontend/assets/js/mission.js the production path payloads come from the real network response

docs/FINAL_DELIVERY_TODO.md:73 records the distinction: "runMock only when no file is selected; emits empty payloads, marked is-mock; not in production path." The live driver emits all eight events around real network calls (docs/FINAL_DELIVERY_TODO.md:71): "runLive emits 8 events around real network calls; markState uses real values."

13.2 markState β€” the trace is driven by events, never by a timer

frontend/assets/js/mission.js:404-425:

function markState(id, note, isLive) {
  var n = traceNodes[id];
  if (!n) return;
  n.node.classList.toggle('is-mock', !isLive);
  n.tm.textContent = note !== undefined ? note : (STATE_NOTE[id] || '');

  /* Advance the trace to the furthest state reached. This is driven by the
     same events that carry the real result, so the bar and the node states
     move only when the run actually reaches a stage β€” never on a timer. The
     fill spans from the left edge to the centre of the current node. */
  var idx = STATES.indexOf(id);
  if (idx > traceProgress) traceProgress = idx;
  ...
  if (traceFill) {
    traceFill.style.width = (((traceProgress + 0.5) / STATES.length) * 100) + '%';
  }
}

The design property in the comment is the important one: the progress bar is a function of reached states, not of elapsed time. A run that stalls shows a stalled bar.

13.3 The measured 94.4444 % trace fill

docs/FINAL_DELIVERY_TODO.md:72:

Area Status Evidence
Execution trace (progress bar) VERIFIED .trace__fill width is set from real event count (measured 94.4444 % live, 2026-09-25)

The number is derivable from the formula, which is why it is a real measurement rather than a coincidence. The live path reaches RESPOND, the ninth of nine ControllerState values:

width = ((traceProgress + 0.5) / STATES.length) * 100
      = ((8 + 0.5) / 9) * 100
      = (8.5 / 9) * 100
      = 94.4444… %

STATES is the nine-state ControllerState list (Β§12); traceProgress = 8 is the zero-based index of RESPOND. The + 0.5 is what makes the bar span to the centre of the current node rather than to its left edge β€” so a fully-completed run stops at 94.4444 %, not 100 %, because the last node's centre is half a node-width short of the right edge. A bar that read 100 % on a nine-state rail would be reporting something the run never did.

14. Worked example β€” one item's full journey

The following traces a single grounding claim from a specialist's Box to a client-facing evidence item. Every step is a real code path.

sequenceDiagram
  participant G as GroundingSpecialist
  participant E as EvidenceEngine
  participant C as confidence.calibrate
  participant T as ExecutionTrace
  participant F as Frontend

  G->>G: compute Box(x1,y1,x2,y2, score, label, coordinate_system)
  G->>E: SpecialistResult.evidence = [Evidence(type=bounding_box, ...)]
  Note over E: aggregate([grounding_result])
  E->>E: collect β†’ [ev_a1b2c3d4e5f6]
  E->>E: _identity_key β†’ (bounding_box, grounding, normalized_0_1, (...), 0.87)
  E->>E: _deduplicate β†’ no collision
  E->>E: sorted(_sort_key)
  E->>E: _record_agreement β†’ single specialist, no annotation
  E->>E: _renumber β†’ evidence_001
  E-->>T: EvidenceCollection(items=[evidence_001], sources=['grounding'], ...)
  E->>C: confidence_for(grounding_result)
  C->>C: _is_effective(0.9772731820958189) == True
  C->>C: _logit(0.87) / 0.9772732 β†’ _sigmoid β†’ _clamp01
  C-->>T: ConfidenceBreakdown(raw=0.87, calibrated=0.8749186…, method="temperature_scaling")
  T->>F: emit EVIDENCE_GENERATED {regions, count, threshold, registrationRMSE}
  T->>F: emit CONFIDENCE_COMPUTED {calibrated, method, bins, observed}
  T->>F: emit RESULT_ASSEMBLED {text, evidenceIds, confidence, provenance}

The state of the item at each step:

Step evidence_id type source_specialist coordinate_system coordinates score artifact_ref payload
specialist output ev_a1b2c3d4e5f6 bounding_box grounding normalized_0_1 [0.21,0.33,0.47,0.61] 0.87 null {"label": "water"}
after dedup unchanged unchanged unchanged unchanged unchanged unchanged unchanged unchanged
after sort unchanged unchanged unchanged unchanged unchanged unchanged unchanged unchanged
after agreement unchanged unchanged unchanged unchanged unchanged unchanged unchanged {"label": "water"} (no corroboration β€” single specialist)
after renumber evidence_001 … … … … … … …

And the source object is untouched. After aggregate returns, the original SpecialistResult.evidence[0] still has evidence_id == "ev_a1b2c3d4e5f6". This is the purity contract (Β§5.4) made concrete: _renumber calls model_copy(update={"evidence_id": ...}), which produces a new object and leaves the source alone.

Verified confidence values (read-only execution against the shipped module):

Input raw calibrated method
calibrate(0.87, T=0.9772731820958189) 0.87 0.8749186077809417 temperature_scaling
calibrate(0.87, T=1.0) 0.87 None uncalibrated
calibrate(0.0, None) 0.0 None uncalibrated

Part E β€” Honest boundaries

15. What is NOT RUN, OPEN, BLOCKED or REJECTED for this topic

Item Status Note
Calibration improves ECE REJECTED measured 0.013755 β†’ 0.014929 (worse). The artifact is retained because it is in the frozen config, not because it helps.
Calibration applied in a served path OPEN docs/PHASE19_FINAL_HARDENING.md:377 β€” "Nothing is calibrated in a deployed path"
Calibration covers specialists other than change_vqa NOT RUN the artifact's scope covers change_vqa only; every other specialist's confidence is uncalibrated
A second fitted artifact for the other specialists NOT RUN no artifact exists
EvidenceEngine wired into core/controller.py IMPLEMENTED (Phase 14) docs/PHASE13_EVIDENCE_ENGINE.md Β§8 recorded it as not-yet-wired at Phase 13; the live run path emits the eight events
Artifact rendering (crops, masks, change maps served to a client) OPEN "artifact_ref and the evidence_from_* methods provide the hooks; rendering crops, masks and change maps is controller/GUI territory." (docs/PHASE13_EVIDENCE_ENGINE.md Β§8)
Artifact retrieval endpoint REJECTED for v1 F-16: no artifact-serving endpoint exists, so artifact_ref is permanently null
contributing_specialists field on Evidence REJECTED proposed and not adopted; core/schemas.py unchanged (docs/PHASE13_EVIDENCE_ENGINE.md Β§5)
Change-degradation clause in SpecialistResult RESOLVED (removed, F-16c) the narrower replacement rule is REJECTED (core/schemas.py:362-395)
payload["corroborated_by"] exercised by a live multi-specialist run UNKNOWN β€” not established from the available evidence the annotation is pinned by unit tests; no live run was found that produced two specialists making an identical claim
The multiclass (softmax(logits/T)) form of the calibration applied in code UNKNOWN β€” not established from the available evidence the artifact documents the capability; evidence/confidence.py implements the scalar form (Β§10.3)
Whether _sort_key's coordinate tie-break is exercised with non-equal coordinates in production UNKNOWN β€” not established from the available evidence pinned by unit test test_equal_scores_order_deterministically_by_coordinates
The exact count of live runs that populated dropped_over_limit > 0 UNKNOWN β€” not established from the available evidence no run record was found
evidence.max_items = 32 ever being the binding constraint in a live run UNKNOWN β€” not established from the available evidence the cap is config-declared; no live run's total_before_limit was found

16. Where the evidence lives

Claim Source
Evidence fields and the spatial validator core/schemas.py:216-253
artifact_ref is permanently null (F-16) core/schemas.py:225-235; docs/API_CONTRACT.md Β§2.4; docs/DEPLOYMENT_ARCHITECTURE.md Β§5.5
F-16c and the removed change clause core/schemas.py:362-395; docs/DEPLOYMENT_ARCHITECTURE.md:855-866
EvidenceType β€” all 11 members core/schemas.py:65-76
CoordinateSystem β€” 3 values core/schemas.py:57-62
Purity contract evidence/engine.py:38-50
_identity_key and its exclusions evidence/engine.py:125-142; :52-68
_claim_key evidence/engine.py:145-162
_sort_key and the four-level rationale evidence/engine.py:165-190
_deduplicate and the payload merge evidence/engine.py:352-389
_record_agreement and AGREEMENT_KEY evidence/engine.py:391-425; :234-238
_renumber and ID_PREFIX evidence/engine.py:427-430; :104-107
The cap and the five accounting fields evidence/engine.py:193-217; :337-350
evidence_digest evidence/engine.py:680-704
evidence_type_for evidence/engine.py:434-458
evidence_from_box/region/geospatial evidence/engine.py:462-572
confidence_for and the degradation-wins rule evidence/engine.py:576-637
The honesty rule evidence/confidence.py:12-30
logit/sigmoid/_EPS/_clamp01 evidence/confidence.py:80-115
_is_effective and _IDENTITY_TOLERANCE evidence/confidence.py:90-93, :118-126
TemperatureCalibration and its range guard evidence/confidence.py:129-252
calibrate / calibrate_result evidence/confidence.py:255-346
load_calibration and the ordered search evidence/confidence.py:349-424
The 77-test phase record docs/PHASE13_EVIDENCE_ENGINE.md
The measured calibration result artifacts/calibration_v001.json; docs/STEP7_BACKEND_CHAIN_REPORT.md Β§7; docs/FINAL_DELIVERY_TODO.md:104
The eight events frontend/assets/js/core.js:616-620
The stage rail frontend/assets/js/core.js:605-614
markState and the fill formula frontend/assets/js/mission.js:404-425
The 94.4444 % measurement docs/FINAL_DELIVERY_TODO.md:72
ExecutionTrace / TraceStep / ModelRef core/schemas.py:279-319
ControllerState β€” nine states core/schemas.py:79-88; configs/base.yaml Β§agent.states
evidence.max_items: 32 configs/base.yaml Β§evidence
confidence.* frozen keys configs/base.yaml Β§confidence
Layer verbs docs/ARCHITECTURE_FREEZE.md Β§5

Next: 07 Configuration freeze β€” the registry, the enforced invariants, and the hash 78f1e3700da15aa1 that every artifact in this chapter is keyed to.