SatQuery / docs /DATASETS.md
thundercode's picture
release: add docs/DATASETS.md
0adf671 verified
|
Raw History Blame
6.89 kB
# Datasets
Every figure below was **measured from the data on disk**, not copied from a paper or a dataset
schema. Where a corpus was only partially acquired, or where the local subset differs from the
official release, that is stated β€” not smoothed over.
**Status tags:** `ACQUIRED` Β· `MEASURED` Β· `PARTIAL` Β· `NOT DOWNLOADED` Β· `REJECTED`.
---
## 1. Summary
| Task | Dataset | Role | Local state |
|---|---|---|---|
| `change` | **LEVIR-CD-256** | train / val / test | ACQUIRED β€” split 7120 / 1024 / 2048 |
| `change_vqa` | **CDVQA** (+ **SECOND** imagery) | train / val / test | annotations + imagery ACQUIRED; name-verified 2,968/2,968 |
| `grounding` | **VRSBench** | eval (16,159 records) | ACQUIRED β€” boxes normalised 0–100 |
| `optical_sar` | **BigEarthNet** (CLC-19) | train / held-out test | PARTIAL β€” 28,000-patch local subset |
| `vlm` | **BigEarthNet** (instruction pairs) | LoRA adaptation | PARTIAL β€” same 28k subset |
| β€” | SECOND (SCD) | CDVQA imagery source | ACQUIRED, extracted to `data/cdvqa/{im1,im2,label1,label2}/` |
## 2. LEVIR-CD-256 β€” change detection
The standard change-detection benchmark, 256Γ—256 tiles. Split used (from `configs/base.yaml`):
| Split | Count |
|---|---|
| train | 7,120 |
| val | 1,024 |
| test | **2,048** |
**Measured test-split detail** (`artifacts/change/eval_test/eval_result.json`):
| Property | Value |
|---|---|
| n | 2,048 |
| images containing change | **935** |
| mean change fraction | **0.0509** (β‰ˆ 5 % of pixels) |
| threshold | 0.50 |
| tile size | 256 |
| device (eval) | cuda |
| total pixels scored | 134,217,728 |
The 5 % change fraction is why **pooled** and **macro** metrics are both reported: with a strong
class imbalance, pooled IoU (0.8122) and macro IoU (0.8457) answer different questions. The
per-pixel confusion counts (tp 5,978,997 / fp 523,658 / fn 858,407 / tn 126,856,666) are stored so
any metric can be recomputed rather than trusted.
## 3. CDVQA (+ SECOND) β€” change-VQA
**The CDVQA repository publishes annotations only β€” no imagery.** The imagery is publicly available
as **SECOND (SCD)**. This project acquired SECOND, **name-verified it (2,968 / 2,968 MATCH)**,
extracted it to `data/cdvqa/{im1,im2,label1,label2}/`, and verified a CDVQA example loads against it
end-to-end.
**Measured corpus figures:**
| Property | Value |
|---|---|
| Val images | 400 |
| Val questions | **16,441** |
| Test questions | **39,686** |
| Test2 questions | second held-out set (accuracy 0.651469) |
**Temporal and label semantics β€” established from evidence, with honest uncertainty:**
| Mapping | State |
|---|---|
| `label1` = pre, `label2` = post | **established** |
| `im1` = pre, `im2` = post | **supported, not proven** |
**Corpus root:** `data/cdvqa` β€” the directory that *contains* `annotations/`. Passing
`data/cdvqa/annotations` is **rejected** by `_resolve_root` (verified by execution: it raises
`no 'annotations' directory under data\cdvqa\annotations`). The correct call is
`load_cdvqa('data/cdvqa')`.
> **Not established: trainability of the raw loader.** The adapter is pair-aware and
> `require_images=True` succeeds, but the raw corpus has no training loop of its own β€” the shipped
> `change_vqa` head trains on **cached features**, not on the raw loader. See
> [`TRAINING.md`](TRAINING.md).
## 4. VRSBench β€” grounding evaluation
The grounding head is evaluated on VRSBench. **Measured: 16,159 eval records.**
VRSBench stores boxes **normalised to 0–100**; this project stores boxes **normalised to 0–1**. The
conversion is declared explicitly as `grounding.benchmark_box_scale: 100.0` so it cannot be applied
twice or forgotten.
Because the box convention is a common source of silent error, grounding is reported under **two
protocols** (canonical and matched6) and **two decode variants** β€” see
[`BENCHMARKS.md`](BENCHMARKS.md) Β§1.
## 5. BigEarthNet β€” optical-SAR fusion and VLM adaptation
BigEarthNet is used in two places: as the **label space and evaluation benchmark** for the
optical-SAR fusion head, and as the **instruction-pair source** for the SmolVLM LoRA adaptation.
**Measured local subset:**
| Property | Value |
|---|---|
| S2 patches (local subset) | **28,000** |
| tiles | 98 |
| bands per patch | 12 |
| full official corpus | 480,038 patches |
| label space | **19 CLC classes** |
### 5.1 Label-semantics caveat (important)
> **The local 28k subset is 100 % single-label, against the official 1–11 multi-label scheme.**
> Metrics computed on this subset are therefore **not comparable** to published multi-label
> BigEarthNet numbers. Any statement of the form "BigEarthNet mAP = X" is **false** for this subset.
### 5.2 Fusion evaluation detail
**Measured** (`artifacts/optical_sar/fusion_head_production_v001/pre_registered_115_metric.json`):
| Property | Value |
|---|---|
| split | test |
| n scored | **4,000** |
| classes in label space | 19 |
| classes present in the scored split | 14 |
| classes absent | 5 |
| macro-F1 denominator | **all 19 classes (absent classes contribute 0.0)** |
| accuracy | 0.931 |
| macro F1 | 0.434161 |
| deciding statistic | **False** |
The macro-F1 denominator is recorded explicitly so a reader can see that absent classes drag the
macro score down by construction. This is why accuracy (0.931) and macro-F1 (0.434161) must be read
together.
### 5.3 The BigEarthNet data-format contradiction
> **PARTIALLY β€” reported, not silently resolved.** The BigEarthNet data format contradicts the
> original plan. This was reported rather than quietly patched, because silently changing the
> preprocessing would move the frozen config hash. See
> [`RESEARCH_NOTES.md`](RESEARCH_NOTES.md) for the full record.
The BigEarthNet documentation β€” its uses, mentions, or endorsements β€” does **not** specify a
percentile stretch. This project nevertheless applies percentile normalisation for optical inputs
(2/98) to match the CROMA contract. That is a deliberate, documented choice, not an upstream fact.
## 6. Data hygiene and leakage controls
- **Splits are by group, never by example**, for the router: template / hard-negative families are
kept whole, and hard-negative families are placed in the **test** split so their accuracy measures
generalisation rather than memorisation.
- **Leakage split key: `scene_id`** (declared in `configs/base.yaml`).
- **Immutable public test: true.** `hidden_data_access: false`. The evaluation config forbids
touching hidden data.
## 7. What is NOT available
| Corpus | State |
|---|---|
| BigEarthNet S1+S2 full corpus | **NOT DOWNLOADED** (only the 28k S2 subset is local) |
| BigEarthNet multi-label (reBEN) results | **not produced** β€” the local subset is single-label |
| Cross-dataset generalisation sets | **not used** |
| Any private / hidden evaluation data | **not accessed** (`hidden_data_access: false`) |