# Datasets Every figure below was **measured from the data on disk**, not copied from a paper or a dataset schema. Where a corpus was only partially acquired, or where the local subset differs from the official release, that is stated — not smoothed over. **Status tags:** `ACQUIRED` · `MEASURED` · `PARTIAL` · `NOT DOWNLOADED` · `REJECTED`. --- ## 1. Summary | Task | Dataset | Role | Local state | |---|---|---|---| | `change` | **LEVIR-CD-256** | train / val / test | ACQUIRED — split 7120 / 1024 / 2048 | | `change_vqa` | **CDVQA** (+ **SECOND** imagery) | train / val / test | annotations + imagery ACQUIRED; name-verified 2,968/2,968 | | `grounding` | **VRSBench** | eval (16,159 records) | ACQUIRED — boxes normalised 0–100 | | `optical_sar` | **BigEarthNet** (CLC-19) | train / held-out test | PARTIAL — 28,000-patch local subset | | `vlm` | **BigEarthNet** (instruction pairs) | LoRA adaptation | PARTIAL — same 28k subset | | — | SECOND (SCD) | CDVQA imagery source | ACQUIRED, extracted to `data/cdvqa/{im1,im2,label1,label2}/` | ## 2. LEVIR-CD-256 — change detection The standard change-detection benchmark, 256×256 tiles. Split used (from `configs/base.yaml`): | Split | Count | |---|---| | train | 7,120 | | val | 1,024 | | test | **2,048** | **Measured test-split detail** (`artifacts/change/eval_test/eval_result.json`): | Property | Value | |---|---| | n | 2,048 | | images containing change | **935** | | mean change fraction | **0.0509** (≈ 5 % of pixels) | | threshold | 0.50 | | tile size | 256 | | device (eval) | cuda | | total pixels scored | 134,217,728 | The 5 % change fraction is why **pooled** and **macro** metrics are both reported: with a strong class imbalance, pooled IoU (0.8122) and macro IoU (0.8457) answer different questions. The per-pixel confusion counts (tp 5,978,997 / fp 523,658 / fn 858,407 / tn 126,856,666) are stored so any metric can be recomputed rather than trusted. ## 3. CDVQA (+ SECOND) — change-VQA **The CDVQA repository publishes annotations only — no imagery.** The imagery is publicly available as **SECOND (SCD)**. This project acquired SECOND, **name-verified it (2,968 / 2,968 MATCH)**, extracted it to `data/cdvqa/{im1,im2,label1,label2}/`, and verified a CDVQA example loads against it end-to-end. **Measured corpus figures:** | Property | Value | |---|---| | Val images | 400 | | Val questions | **16,441** | | Test questions | **39,686** | | Test2 questions | second held-out set (accuracy 0.651469) | **Temporal and label semantics — established from evidence, with honest uncertainty:** | Mapping | State | |---|---| | `label1` = pre, `label2` = post | **established** | | `im1` = pre, `im2` = post | **supported, not proven** | **Corpus root:** `data/cdvqa` — the directory that *contains* `annotations/`. Passing `data/cdvqa/annotations` is **rejected** by `_resolve_root` (verified by execution: it raises `no 'annotations' directory under data\cdvqa\annotations`). The correct call is `load_cdvqa('data/cdvqa')`. > **Not established: trainability of the raw loader.** The adapter is pair-aware and > `require_images=True` succeeds, but the raw corpus has no training loop of its own — the shipped > `change_vqa` head trains on **cached features**, not on the raw loader. See > [`TRAINING.md`](TRAINING.md). ## 4. VRSBench — grounding evaluation The grounding head is evaluated on VRSBench. **Measured: 16,159 eval records.** VRSBench stores boxes **normalised to 0–100**; this project stores boxes **normalised to 0–1**. The conversion is declared explicitly as `grounding.benchmark_box_scale: 100.0` so it cannot be applied twice or forgotten. Because the box convention is a common source of silent error, grounding is reported under **two protocols** (canonical and matched6) and **two decode variants** — see [`BENCHMARKS.md`](BENCHMARKS.md) §1. ## 5. BigEarthNet — optical-SAR fusion and VLM adaptation BigEarthNet is used in two places: as the **label space and evaluation benchmark** for the optical-SAR fusion head, and as the **instruction-pair source** for the SmolVLM LoRA adaptation. **Measured local subset:** | Property | Value | |---|---| | S2 patches (local subset) | **28,000** | | tiles | 98 | | bands per patch | 12 | | full official corpus | 480,038 patches | | label space | **19 CLC classes** | ### 5.1 Label-semantics caveat (important) > **The local 28k subset is 100 % single-label, against the official 1–11 multi-label scheme.** > Metrics computed on this subset are therefore **not comparable** to published multi-label > BigEarthNet numbers. Any statement of the form "BigEarthNet mAP = X" is **false** for this subset. ### 5.2 Fusion evaluation detail **Measured** (`artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`): | Property | Value | |---|---| | split | test | | n scored | **4,000** | | classes in label space | 19 | | classes present in the scored split | 14 | | classes absent | 5 | | macro-F1 denominator | **all 19 classes (absent classes contribute 0.0)** | | accuracy | 0.931 | | macro F1 | 0.434161 | | deciding statistic | **False** | The macro-F1 denominator is recorded explicitly so a reader can see that absent classes drag the macro score down by construction. This is why accuracy (0.931) and macro-F1 (0.434161) must be read together. ### 5.3 The BigEarthNet data-format contradiction > **PARTIALLY — reported, not silently resolved.** The BigEarthNet data format contradicts the > original plan. This was reported rather than quietly patched, because silently changing the > preprocessing would move the frozen config hash. See > [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) for the full record. The BigEarthNet documentation — its uses, mentions, or endorsements — does **not** specify a percentile stretch. This project nevertheless applies percentile normalisation for optical inputs (2/98) to match the CROMA contract. That is a deliberate, documented choice, not an upstream fact. ## 6. Data hygiene and leakage controls - **Splits are by group, never by example**, for the router: template / hard-negative families are kept whole, and hard-negative families are placed in the **test** split so their accuracy measures generalisation rather than memorisation. - **Leakage split key: `scene_id`** (declared in `configs/base.yaml`). - **Immutable public test: true.** `hidden_data_access: false`. The evaluation config forbids touching hidden data. ## 7. What is NOT available | Corpus | State | |---|---| | BigEarthNet S1+S2 full corpus | **NOT DOWNLOADED** (only the 28k S2 subset is local) | | BigEarthNet multi-label (reBEN) results | **not produced** — the local subset is single-label | | Cross-dataset generalisation sets | **not used** | | Any private / hidden evaluation data | **not accessed** (`hidden_data_access: false`) |