# Evaluation This document describes **how** each number in [`BENCHMARKS.md`](BENCHMARKS.md) was produced, and enforces the project's evaluation-honesty rules. It is deliberately conservative: a metric that was not measured is tagged `NOT RUN`, and a metric that got worse is shown getting worse. **Status tags:** `VERIFIED` · `MEASURED` · `NOT RUN` · `OPEN` · `REJECTED`. --- ## 1. Evaluation-honesty rules (enforced by convention and by tooling) 1. **Evidence before claims.** Every reported number has an artifact path. A number with no artifact does not appear in the docs. 2. **Two protocols are never collapsed.** Grounding is reported under canonical **and** matched6. 3. **Two test sets are never collapsed.** Change-VQA is reported on `test` **and** `test2`. 4. **accuracy never travels without macro-F1** for imbalanced multi-class heads. 5. **Validation is not test.** The router figure is labelled validation, ungated, n = 86. 6. **A negative result stays negative.** Calibration ECE worsened; it is shown worsening. 7. **USABLE ≠ ACCEPTED.** The VLM adapter's metrics are real; its status is acceptance-rejected. 8. **No composite/vanity score.** There is no single headline accuracy for the system, and none is invented by averaging the per-task numbers. ## 2. Per-task evaluation protocols ### 2.1 Change detection — `VERIFIED` - **Split:** LEVIR-CD-256 test, n = 2,048 (immutable public split). - **Threshold:** 0.50, frozen in config. - **Metrics:** pooled IoU / macro IoU / pooled F1, plus the full confusion counts so any metric can be recomputed. - **Why both pooled and macro:** the test split is only ≈ 5 % changed pixels (mean change fraction 0.0509). Pooled IoU (0.8122) and macro IoU (0.8457) answer different questions about that imbalance. - **Artifact:** `artifacts/change/eval_test/eval_result.json`. ### 2.2 Grounding — `MEASURED`, two protocols × two decode variants - **Split:** VRSBench, n = 16,159. - **Box convention:** VRSBench 0–100 → project 0–1, via declared `benchmark_box_scale: 100.0`. - **Protocols:** *canonical* and *matched6* — both reported. - **Decode variants:** *head threshold* (the shipped decode), *head argmax*, and *zero-shot matched* (baseline). - **Resolution:** frozen at 224; 448 rejected by a pre-registered paired test. | Protocol | decode | mean best IoU | recall@0.5 | |---|---|---|---| | canonical | head threshold | 0.2838 | 0.2198 | | canonical | head argmax | 0.1215 | — | | canonical | zero-shot matched | 0.0972 | — | | matched6 | head threshold | 0.2566 | 0.1938 | **Artifacts:** `…/eval_result_canonical.json`, `…/eval_result_matched6.json`. ### 2.3 Optical-SAR fusion — `MEASURED`, ruling `OPEN` - **Split:** held-out test, n = 4,000. - **Label space:** 19 CLC classes; 14 present, 5 absent in the scored split. - **macro-F1 denominator:** all 19 classes (absent classes contribute 0.0) — recorded explicitly. - **Pre-registration:** the metric is a pre-registered protocol (`pre_registered_11.5`); `is_deciding_statistic: False`. - **Artifact:** `artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`. ### 2.4 Change-VQA — `MEASURED`, ruling `OPEN` - **Splits:** `test` (n = 39,686) and `test2`. - **Selection:** epoch 8, chosen on val answer accuracy 0.700018. - **Artifact:** `artifacts/change_vqa/run/PROMOTION.json`. ### 2.5 VLM adapter — `MEASURED`, `ACCEPTANCE-REJECTED` - **Split:** frozen 1,000-question subset. - **Metrics:** exact_match 0.963, F1 0.96432 (+49.5 pp over unadapted baseline). - **Decision:** the artifact is **not accepted** for promotion; the record of *why* is stored in `artifacts/vlm/phase6_closure.json` under `why_acceptance_rejected`. - **Artifact:** `artifacts/vlm/phase6_closure.json` (`status: CLOSED`). ### 2.6 Router — `MEASURED`, `TEST NOT RUN` - **Split scored:** validation, n = 86, corpus-limited. - **Metric:** overall **ungated** accuracy 0.965116. - **Not done:** the test split was never scored. The val number is indicative only. ### 2.7 Calibration — `MEASURED`, not an improvement - **Fit split:** val, n = 16,441. Temperature **T = 0.9772732**. - **Result:** ECE 0.013755 → **0.014929** (`ece_improvement = −0.001174`). - **Retention rationale:** part of the frozen config, **not** because it helped. - **Artifact:** `artifacts/calibration_v001.json`. ## 3. Behavioural evaluation (live validation) Accuracy and behaviour are evaluated separately. The deployed stack was driven in a headed browser, one upload per case, with per-case screenshots and recorded run ids: | Property | Result | |---|---| | Independent full passes | **3** | | Cases per pass | 8 (6 regression + 2 router-defect) | | Passes at 8/8 | **3 of 3** | | Live runs | **24** | | Correct dispatches | **24** | | Mock-node contamination | **0** | | Trace fill | **94.4444 %** | | Frontend regression suite | **106 passed** | ### 3.1 Harness integrity — a bug that was caught An earlier harness revision typed queries with **synthetic CDP key events**, which Chrome **silently drops when the window lacks OS focus**. The harness therefore dispatched the page's *default* query and still recorded a "result" — a false pass. The current harness **asserts form state before dispatch**: that the query box really holds the intended query, that `#obsTail` reads `ready`, and that both frames are attached for pair tasks. The earlier 8/8 run was independently checked and confirmed **not** to have been infected (its answers were query-specific and the query text was embedded in the answers). This failure mode is recorded here because it is exactly the kind of silent false-positive an evaluation harness must not have. ## 4. Test suite | Suite | Result | Notes | |---|---|---| | `tests/unit/test_frontend_live_wiring.py` | **106 passed** | re-run this session | | Doc/frontend suite (5 files) | **183 passed** | | | Full `tests/unit` | 5–6 failures, **all environmental/ordering** | 4× `test_safe_delete_shim` (sandbox delete guard), 1 ordering flake (passes in isolation), 1 stale adapter test (CROMA now shipped) | | Affected files re-run together | **137 passed** | confirms the failures are not real regressions | **The environmental failures are not hidden.** They are attributable to the sandbox delete guard and test ordering, not to the code under test. ## 5. What is NOT evaluated | Item | State | |---|---| | System-level end-to-end accuracy | **NOT RUN — none exists** | | Router test split | **NOT RUN** | | Benchmark adapters | **NOT RUN** | | End-to-end latency benchmark | **NOT RUN** | | Cross-dataset generalisation | **NOT RUN** | | Human evaluation | **NOT RUN** | | Adversarial / robustness evaluation | **NOT RUN** | ## 6. Reproducing the metric check ```bash python release/tools/verify_readme_metrics.py ``` This walks every claim in [`../README.md`](../README.md) and [`BENCHMARKS.md`](BENCHMARKS.md) to its source artifact and compares values at the printed precision. It prints `ALL CLAIMS VERIFIED` (exit 0) only when **all 20** numeric claims match and the status assertions hold. Committed output: `release/tools/readme_metrics_report.txt`.