SEEDS / README.md
yzt15806542928's picture
Upload folder using huggingface_hub
1ca0208 verified
|
Raw
History Blame Contribute Delete
5.16 kB
---
frameworks: PyTorch
language:
- en
license: apache-2.0
tags:
- OneScience
- Earth Science
- Weather Forecast
- Ensemble Weather Forecasting
- Diffusion Model
- GEFS
- ERA5
tasks: []
datasets:
- ERA5
---
<p align="center">
<strong>
<span style="font-size: 30px;">SEEDS</span>
</strong>
</p>
# Model Introduction
SEEDS is a generative weather model released by Google in March 2024. The name stands for Scalable Ensemble Envelope Diffusion Sampler; the associated paper was published in *Science Advances*.
Paper: SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models
https://arxiv.org/abs/2306.14066
# Model Description
SEEDS is a conditional diffusion model that uses a small number of numerical weather prediction seed members to efficiently generate large forecast ensembles.
# Use Cases
| Scenario | Description |
| :---: | :--- |
| Ensemble weather forecast research | Train a conditional diffusion model on data following the project's cubed-sphere NPZ protocol and generate forecast ensembles. |
| Local quick validation | Use synthetic data to check training, inference, ensemble evaluation, and visualization. |
| ModelScope / OneCode execution | Download the standalone model package, install dependencies, and run the scripts directly. |
| Multi-GPU training | Launch PyTorch DistributedDataParallel with `torchrun`. |
# Usage Guide
## 1. OneCode Usage
Experience intelligent one-click AI4S programming through the OneCode online environment:
[Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home)
## 2. Manual Installation and Usage
**Hardware Requirements**
- Training and inference require a GPU or DCU recognized by PyTorch. CPU can be used to generate synthetic data and inspect configuration, but cannot run the current training and inference scripts.
- Multi-GPU training uses the NCCL backend. Ensure that the device driver, communication libraries, and PyTorch version are compatible.
- DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended.
### Download the Model Package
```bash
hf download OneScience-Group/SEEDS --local-dir ./SEEDS
cd SEEDS
```
### Install the Runtime Environment
**DCU Environment**
```bash
# Please activate DTK and CONDA first
conda create -n onescience311 python=3.11 -y
conda activate onescience311
# uv installation is supported
pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
**GPU Environment**
```bash
# Please activate CONDA first
conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12
conda activate onescience311
# uv installation is supported
pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai
```
### Training Data Introduction
The SEEDS paper uses GEFS reforecast data for training, operational GEFS members as conditioning inputs, and ERA5 as the evaluation reference. The official data and preprocessing are not bundled with this repository; prepare the NPZ files specified by `conf/config.yaml` before training.
### Generate Synthetic Data
When real data is unavailable, generate a default `6x48x48` cubed-sphere fixture. Synthetic data only validates the program flow and does not represent GEFS, ERA5, or the paper's forecast quality:
```bash
python scripts/fake_data.py
```
### Training
Single GPU:
```bash
python scripts/train.py
```
Multi-GPU:
```bash
torchrun --nproc_per_node=8 scripts/train.py
```
The number of epochs is controlled by `conf/config.yaml`, and the default checkpoint is saved to `data/checkpoint/model_bak.pth`.
### Fine-tuning
To continue from an existing checkpoint, use the explicit fine-tuning flag:
```bash
python scripts/train.py --finetune
```
### Training Weights
This repository provides a `weight/` directory for checkpoints trained on the official data. The weight files will be uploaded soon and are expected to be available in the near future.
### Inference
Inference reads `data/checkpoint/model_bak.pth` by default and generates ensemble members in batches controlled by `sampling.member_batch_size`:
```bash
python scripts/inference.py
```
Predictions, targets, and seed members are saved under `result/output/`.
### Evaluation and Visualization
```bash
python scripts/result.py
```
The script computes ensemble-mean RMSE, ACC, and empirical CRPS, saves the corresponding NPY metrics, and generates `result/forecast.png`. If training loss files are available, it also writes `result/loss.png`.
# Official OneScience Resources
| Platform | OneScience Main Repository | Skills Repository |
| --- | --- | --- |
| Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills |
| GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills |
# Citation and License
- This repository is an independent adaptation of the SEEDS paper and is not an official Google product.