SatQuery / docs /EVALUATION.md
thundercode's picture
release: add docs/EVALUATION.md
c2283ee verified
|
Raw History Blame
7.19 kB

Evaluation

This document describes how each number in BENCHMARKS.md was produced, and enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was not measured is tagged NOT RUN, and a metric that got worse is shown getting worse.

Status tags: VERIFIED · MEASURED · NOT RUN · OPEN · REJECTED.


1. Evaluation-honesty rules (enforced by convention and by tooling)

  1. Evidence before claims. Every reported number has an artifact path. A number with no artifact does not appear in the docs.
  2. Two protocols are never collapsed. Grounding is reported under canonical and matched6.
  3. Two test sets are never collapsed. Change-VQA is reported on test and test2.
  4. accuracy never travels without macro-F1 for imbalanced multi-class heads.
  5. Validation is not test. The router figure is labelled validation, ungated, n = 86.
  6. A negative result stays negative. Calibration ECE worsened; it is shown worsening.
  7. USABLE ≠ ACCEPTED. The VLM adapter's metrics are real; its status is acceptance-rejected.
  8. No composite/vanity score. There is no single headline accuracy for the system, and none is invented by averaging the per-task numbers.

2. Per-task evaluation protocols

2.1 Change detection — VERIFIED

  • Split: LEVIR-CD-256 test, n = 2,048 (immutable public split).
  • Threshold: 0.50, frozen in config.
  • Metrics: pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can be recomputed.
  • Why both pooled and macro: the test split is only ≈ 5 % changed pixels (mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different questions about that imbalance.
  • Artifact: artifacts/change/eval_test/eval_result.json.

2.2 Grounding — MEASURED, two protocols × two decode variants

  • Split: VRSBench, n = 16,159.
  • Box convention: VRSBench 0–100 → project 0–1, via declared benchmark_box_scale: 100.0.
  • Protocols: canonical and matched6 — both reported.
  • Decode variants: head threshold (the shipped decode), head argmax, and zero-shot matched (baseline).
  • Resolution: frozen at 224; 448 rejected by a pre-registered paired test.
Protocol decode mean best IoU recall@0.5
canonical head threshold 0.2838 0.2198
canonical head argmax 0.1215 —
canonical zero-shot matched 0.0972 —
matched6 head threshold 0.2566 0.1938

Artifacts: …/eval_result_canonical.json, …/eval_result_matched6.json.

2.3 Optical-SAR fusion — MEASURED, ruling OPEN

  • Split: held-out test, n = 4,000.
  • Label space: 19 CLC classes; 14 present, 5 absent in the scored split.
  • macro-F1 denominator: all 19 classes (absent classes contribute 0.0) — recorded explicitly.
  • Pre-registration: the metric is a pre-registered protocol (pre_registered_11.5); is_deciding_statistic: False.
  • Artifact: artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json.

2.4 Change-VQA — MEASURED, ruling OPEN

  • Splits: test (n = 39,686) and test2.
  • Selection: epoch 8, chosen on val answer accuracy 0.700018.
  • Artifact: artifacts/change_vqa/run/PROMOTION.json.

2.5 VLM adapter — MEASURED, ACCEPTANCE-REJECTED

  • Split: frozen 1,000-question subset.
  • Metrics: exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline).
  • Decision: the artifact is not accepted for promotion; the record of why is stored in artifacts/vlm/phase6_closure.json under why_acceptance_rejected.
  • Artifact: artifacts/vlm/phase6_closure.json (status: CLOSED).

2.6 Router — MEASURED, TEST NOT RUN

  • Split scored: validation, n = 86, corpus-limited.
  • Metric: overall ungated accuracy 0.965116.
  • Not done: the test split was never scored. The val number is indicative only.

2.7 Calibration — MEASURED, not an improvement

  • Fit split: val, n = 16,441. Temperature T = 0.9772732.
  • Result: ECE 0.013755 → 0.014929 (ece_improvement = −0.001174).
  • Retention rationale: part of the frozen config, not because it helped.
  • Artifact: artifacts/calibration_v001.json.

3. Behavioural evaluation (live validation)

Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser, one upload per case, with per-case screenshots and recorded run ids:

Property Result
Independent full passes 3
Cases per pass 8 (6 regression + 2 router-defect)
Passes at 8/8 3 of 3
Live runs 24
Correct dispatches 24
Mock-node contamination 0
Trace fill 94.4444 %
Frontend regression suite 106 passed

3.1 Harness integrity — a bug that was caught

An earlier harness revision typed queries with synthetic CDP key events, which Chrome silently drops when the window lacks OS focus. The harness therefore dispatched the page's default query and still recorded a "result" — a false pass.

The current harness asserts form state before dispatch: that the query box really holds the intended query, that #obsTail reads ready, and that both frames are attached for pair tasks. The earlier 8/8 run was independently checked and confirmed not to have been infected (its answers were query-specific and the query text was embedded in the answers). This failure mode is recorded here because it is exactly the kind of silent false-positive an evaluation harness must not have.

4. Test suite

Suite Result Notes
tests/unit/test_frontend_live_wiring.py 106 passed re-run this session
Doc/frontend suite (5 files) 183 passed
Full tests/unit 5–6 failures, all environmental/ordering 4× test_safe_delete_shim (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped)
Affected files re-run together 137 passed confirms the failures are not real regressions

The environmental failures are not hidden. They are attributable to the sandbox delete guard and test ordering, not to the code under test.

5. What is NOT evaluated

Item State
System-level end-to-end accuracy NOT RUN — none exists
Router test split NOT RUN
Benchmark adapters NOT RUN
End-to-end latency benchmark NOT RUN
Cross-dataset generalisation NOT RUN
Human evaluation NOT RUN
Adversarial / robustness evaluation NOT RUN

6. Reproducing the metric check

python release/tools/verify_readme_metrics.py

This walks every claim in ../README.md and BENCHMARKS.md to its source artifact and compares values at the printed precision. It prints ALL CLAIMS VERIFIED (exit 0) only when all 20 numeric claims match and the status assertions hold. Committed output: release/tools/readme_metrics_report.txt.