File size: 4,905 Bytes
32dd8b9
 
2227bb5
 
 
32dd8b9
2227bb5
 
 
 
 
32dd8b9
 
 
2227bb5
 
 
 
32dd8b9
2227bb5
32dd8b9
2227bb5
32dd8b9
2227bb5
 
 
 
32dd8b9
2227bb5
32dd8b9
2227bb5
 
 
 
 
 
 
 
 
32dd8b9
2227bb5
 
32dd8b9
2227bb5
 
 
 
 
 
 
 
 
32dd8b9
2227bb5
32dd8b9
2227bb5
 
32dd8b9
2227bb5
 
32dd8b9
2227bb5
 
 
 
 
 
 
 
32dd8b9
2227bb5
 
32dd8b9
2227bb5
32dd8b9
2227bb5
 
32dd8b9
2227bb5
 
8ea64b0
 
 
 
 
 
 
 
 
 
 
 
 
2227bb5
 
 
 
823a45d
 
8ea64b0
 
2227bb5
 
 
 
 
 
32dd8b9
2227bb5
32dd8b9
2227bb5
32dd8b9
2227bb5
 
 
 
 
 
32dd8b9
2227bb5
 
823a45d
39cc2c9
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
# Typical Marine Ecological Environment Feature Recognition

This project is being refactored from a single green-tide detector into a
registry-driven framework for typical marine ecological environment feature
recognition.

The goal is not a closed multiclass model. Each element, satellite, sensor,
resolution, and fused-product policy can be expressed as a small Markdown
capability card, then composed into a task profile for training or inference.
This keeps the system extensible for green tide, red tide, golden tide,
aquaculture, ships, oil spill, sea ice, and future targets.

## Core Constraint

Do not assume coastline vectors, land-mask vectors, or external GIS masks are
available. Land, black borders, no-data areas, water background, cloud, and other
hard negatives must be handled by labels, validity/context heads, hard-negative
samples, or image-derived validity rules.

## Markdown Capability Registry

Cards live under [docs/registry](docs/registry):

- `elements/`: target feature or object cards.
- `satellites/`: satellite/platform cards.
- `sensors/`: sensor, product, and fusion cards.
- `resolutions/`: spatial-resolution policy cards.

Example for an already fused GF6 product:

```powershell
python scripts/compose_task_profile.py `
  --element green_tide `
  --satellite GF6 `
  --sensor PMS `
  --fusion FUSED_OPTICAL `
  --resolution 2m `
  --output configs/profiles/gf6_green_tide_fused_2m.json
```

Example for GF2/GF1-style PAN+MSS imagery where fusion should happen tile by
tile during inference:

```powershell
python scripts/compose_task_profile.py `
  --element green_tide `
  --satellite GF2 `
  --sensor PMS `
  --fusion STREAM_FUSION `
  --resolution 2m `
  --output configs/profiles/gf2_green_tide_stream_fusion_2m.json
```

## Fused Image Expression

Fused imagery is represented as a derived observation, not as a plain
multispectral raster and not as `is_fused=true`.

Every sample that uses a fused raster or runtime fusion should keep a `fusion`
object in the manifest:

- `state`: `none`, `fused_product`, `runtime_fusion`, or `unknown`.
- `method`: vendor method, implemented method, or `unknown_vendor_product`.
- `sources`: PAN/MSS/source product roles, paths, and native resolutions.
- `target_resolution_m`: output grid resolution.
- `native_multispectral_resolution_m`: original multispectral resolution.
- `persisted`: whether the fused raster exists on disk.
- `reproducible`: whether source data and method can reproduce it.
- `spectral_preservation`: known or estimated spectral preservation risk.

This is important because a GF6 fused image, a GF2 tile-wise fusion stream, and a
native multispectral image should not be treated as identical observations.

## Dataset Standard

The normalized dataset is manifest-driven and may retain different patch sizes,
satellites, sensors, resolutions, and fusion states.

See [docs/dataset_standard.md](docs/dataset_standard.md).

For the SAMPoly-style polygon head, the dataset gate is stricter: only real
mask or polygon annotations enter the main polygon training set. Bbox-only
labels are rejected because rectangles cannot supervise true boundaries,
vertices, or polygon ordering. The polygon-ready layout is:

```text
images/{train,val,test}
masks/{train,val,test}
manifests/accepted_polygon_samples.jsonl
manifests/rejected_polygon_samples.jsonl
dataset_card.json
```

## Key Scripts

- `scripts/build_marine_feature_dataset.py`: scan local or server-side data roots
  and write normalized manifests.
- `scripts/search_hf_marine_datasets.py`: search Hugging Face datasets and write
  normalized `hf://` manifest references without downloading full repositories.
- `scripts/prepare_polygon_dataset.py`: build the strict mask/polygon dataset
  for SAMPoly-style training and reject bbox-only samples.
- `scripts/compose_task_profile.py`: compose Markdown capability cards into a
  JSON task profile.
- `scripts/infer_whole_scene.py`: sliding-window whole-scene inference with
  overlap weighting and optional PAN+MSS runtime fusion.
- `scripts/postprocess_mask.py`: temporary cleanup utility for obvious invalid
  areas; it is not a substitute for the final context/validity model heads.

## Model Direction

The target architecture is a unified, explainable, multi-task model:

- shared DINOv3-style visual encoder;
- clear FPN/PAN multi-scale neck;
- validity head for invalid/no-data areas;
- scene-context head for water, land, cloud/shadow, and other;
- element-specific semantic, instance, detection, polygon, or change heads
  selected by Markdown task profiles.

Current legacy green-tide weights can still be loaded for candidate generation,
but final products should come from the registry-driven multi-task framework.

No bbox-only accepted dataset is retained. New training samples should be
imported only when they provide mask or polygon annotations suitable for the
SAMPoly-style head.