File size: 6,338 Bytes
f212e9a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
---
license: mit
language:
- en
tags:
- OneScience
- Earth Science
- Remote Sensing Representation Learning
- Multimodal Remote Sensing
frameworks: PyTorch
datasets:
- Sentinel-1
- Sentinel-2
- NAIP
- EnMAP
- Gaofen
---

<p align="center">
  <strong>
    <span style="font-size: 30px;">DOFA</span>
  </strong>
</p>

# Model Introduction

DOFA is a dynamic single-backbone foundation model for multisensor remote-sensing imagery. It generates band-adaptive weights through a wavelength-conditioned hypernetwork, enabling one model to process observations with different channel counts and spectral responses.

Paper: Neural Plasticity-Inspired Multimodal Foundation Model for Earth Observation  
https://arxiv.org/abs/2403.15356

# Model Description

DOFA was proposed by a research team from institutions including Wuhan University. The model uses five modalities, Sentinel-1, Sentinel-2, NAIP, EnMAP, and Gaofen, for multisensor masked pretraining, with 4-channel Gaofen inputs. It is suitable for multimodal remote-sensing representation learning, image reconstruction, and cross-sensor downstream task adaptation.

# Use Cases

| Scenario | Description |
| :---: | :--- |
| Multisensor pretraining | Use central wavelengths to drive dynamic patch-embedding and decoding weights. |
| Cross-modal representation learning | Use one checkpoint to process five remote-sensing modalities with different channel counts. |
| Land-cover and scene classification | Transfer and fine-tune shared representations for land-cover and remote-sensing scene classification across different sensors. |
| Semantic segmentation | Transfer wavelength-aware features to pixel-level remote-sensing interpretation tasks such as flood and land-cover segmentation and fine-tune them. |
| Local engineering validation | Use a small amount of synthetic data to check training, inference, and evaluation workflows. |
| Multi-accelerator training | Launch distributed training with `torchrun`. |

# Usage Guide

## 1.OneCode

Experience intelligent one-click AI4S programming through the OneCode online environment:

[Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)

## 2.Download and Installation

```bash
hf download OneScience-Group/DOFA --local-dir ./DOFA
cd DOFA
```

### Environment Dependencies

**Hardware Requirements**

- A GPU or DCU is recommended.
- CPU can be used for connectivity validation with a small configuration; full training and inference are slower.
- DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience-recommended version matching the current cluster, is recommended.

**DCU Environment**

```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```

**GPU Environment**

```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```

### Training Data Introduction

By default, one `224x224` synthetic sample per data split is used for each modality to validate the engineering workflow. The five modalities are Sentinel-1 with 2 channels, Sentinel-2 with 9 channels, NAIP with 3 channels, EnMAP with 202 channels, and Gaofen with 4 channels. The synthetic EnMAP wavelengths are evenly spaced approximations used for protocol validation.

The synthetic data retain the channel counts, `224x224` spatial size, and per-channel wavelength input specifications of the authors' official pretraining configuration for the five sensors.

Real data must be preprocessed and converted to the following NPZ training protocol. This protocol is consistent with the model input specification but is not equivalent to the original datasets' download format.

```text
images: float32 [N,C,224,224]
wavelengths: float32 [C]
modality: string scalar
data_range: float scalar
```

`fake_data.py` automatically writes the `protocol`, `data_source`, and `wavelength_mode` protocol metadata. These fields must be retained when using real data.

```bash
python scripts/fake_data.py
```

### Training

```bash
python scripts/train.py
```

Multi-accelerator training can use:

```bash
torchrun --nproc_per_node=8 scripts/train.py
```

Training shares the dynamic backbone across five modalities and saves a checkpoint and overall training metrics. The default configuration is intended for quick workflow validation. Formal experiments should use the multimodal data scale, model configuration, and training schedule corresponding to the paper.

```text
result/checkpoints/dofa.pt
result/training/metrics.json
```

### Training Weights

This repository will provide DOFA training weights in the `weight/` folder. The weight files will be uploaded soon and are expected to be available in the near future.

### Inference

```bash
python scripts/inference.py
```

Inference loads the training checkpoint, performs masked reconstruction for each modality in the configuration, and saves the results to:

```text
result/output/
```

### Evaluation and Visualization

```bash
python scripts/result.py
```

Evaluation reports overall MSE, MAE, PSNR, and mask ratio on masked regions, and generates reconstruction comparison figures for each modality. Synthetic-data results are only for validating the engineering workflow and do not represent the full performance reported in the paper.

```text
result/evaluation/metrics.json
result/evaluation/
```

# Official OneScience Resources

| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |

# Citation and License

This repository is a reproduction of the original DOFA paper.

The use of the code and data in this repository remains subject to the licenses and terms of use of their respective projects.