File size: 5,163 Bytes
1ca0208
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
---
frameworks: PyTorch
language:
- en
license: apache-2.0
tags:
- OneScience
- Earth Science
- Weather Forecast
- Ensemble Weather Forecasting
- Diffusion Model
- GEFS
- ERA5
tasks: []
datasets:
- ERA5
---
<p align="center">
  <strong>
    <span style="font-size: 30px;">SEEDS</span>
  </strong>
</p>

# Model Introduction

SEEDS is a generative weather model released by Google in March 2024. The name stands for Scalable Ensemble Envelope Diffusion Sampler; the associated paper was published in *Science Advances*.

Paper: SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models

https://arxiv.org/abs/2306.14066

# Model Description

SEEDS is a conditional diffusion model that uses a small number of numerical weather prediction seed members to efficiently generate large forecast ensembles.

# Use Cases

| Scenario | Description |
| :---: | :--- |
| Ensemble weather forecast research | Train a conditional diffusion model on data following the project's cubed-sphere NPZ protocol and generate forecast ensembles. |
| Local quick validation | Use synthetic data to check training, inference, ensemble evaluation, and visualization. |
| ModelScope / OneCode execution | Download the standalone model package, install dependencies, and run the scripts directly. |
| Multi-GPU training | Launch PyTorch DistributedDataParallel with `torchrun`. |

# Usage Guide

## 1. OneCode Usage

Experience intelligent one-click AI4S programming through the OneCode online environment:

[Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)

## 2. Manual Installation and Usage

**Hardware Requirements**

- Training and inference require a GPU or DCU recognized by PyTorch. CPU can be used to generate synthetic data and inspect configuration, but cannot run the current training and inference scripts.
- Multi-GPU training uses the NCCL backend. Ensure that the device driver, communication libraries, and PyTorch version are compatible.
- DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended.

### Download the Model Package

```bash
hf download OneScience-Group/SEEDS --local-dir ./SEEDS
cd SEEDS
```

### Install the Runtime Environment

**DCU Environment**

```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
```

**GPU Environment**

```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/  --trusted-host mirrors.onescience.ai
```

### Training Data Introduction

The SEEDS paper uses GEFS reforecast data for training, operational GEFS members as conditioning inputs, and ERA5 as the evaluation reference. The official data and preprocessing are not bundled with this repository; prepare the NPZ files specified by `conf/config.yaml` before training.

### Generate Synthetic Data

When real data is unavailable, generate a default `6x48x48` cubed-sphere fixture. Synthetic data only validates the program flow and does not represent GEFS, ERA5, or the paper's forecast quality:

```bash
python scripts/fake_data.py
```

### Training

Single GPU:

```bash
python scripts/train.py
```

Multi-GPU:

```bash
torchrun --nproc_per_node=8 scripts/train.py
```

The number of epochs is controlled by `conf/config.yaml`, and the default checkpoint is saved to `data/checkpoint/model_bak.pth`.

### Fine-tuning

To continue from an existing checkpoint, use the explicit fine-tuning flag:

```bash
python scripts/train.py --finetune
```

### Training Weights

This repository provides a `weight/` directory for checkpoints trained on the official data. The weight files will be uploaded soon and are expected to be available in the near future.

### Inference

Inference reads `data/checkpoint/model_bak.pth` by default and generates ensemble members in batches controlled by `sampling.member_batch_size`:

```bash
python scripts/inference.py
```

Predictions, targets, and seed members are saved under `result/output/`.

### Evaluation and Visualization

```bash
python scripts/result.py
```

The script computes ensemble-mean RMSE, ACC, and empirical CRPS, saves the corresponding NPY metrics, and generates `result/forecast.png`. If training loss files are available, it also writes `result/loss.png`.

# Official OneScience Resources

| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |

# Citation and License

- This repository is an independent adaptation of the SEEDS paper and is not an official Google product.