--- frameworks: PyTorch language: - en license: apache-2.0 tags: - OneScience - Earth Science - Weather Forecast - Ensemble Weather Forecasting - Diffusion Model - GEFS - ERA5 tasks: [] datasets: - ERA5 ---
SEEDS
# Model Introduction SEEDS is a generative weather model released by Google in March 2024. The name stands for Scalable Ensemble Envelope Diffusion Sampler; the associated paper was published in *Science Advances*. Paper: SEEDS: Emulation of Weather Forecast Ensembles with Diffusion Models https://arxiv.org/abs/2306.14066 # Model Description SEEDS is a conditional diffusion model that uses a small number of numerical weather prediction seed members to efficiently generate large forecast ensembles. # Use Cases | Scenario | Description | | :---: | :--- | | Ensemble weather forecast research | Train a conditional diffusion model on data following the project's cubed-sphere NPZ protocol and generate forecast ensembles. | | Local quick validation | Use synthetic data to check training, inference, ensemble evaluation, and visualization. | | ModelScope / OneCode execution | Download the standalone model package, install dependencies, and run the scripts directly. | | Multi-GPU training | Launch PyTorch DistributedDataParallel with `torchrun`. | # Usage Guide ## 1. OneCode Usage Experience intelligent one-click AI4S programming through the OneCode online environment: [Click to Experience Intelligent One-Click AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home) ## 2. Manual Installation and Usage **Hardware Requirements** - Training and inference require a GPU or DCU recognized by PyTorch. CPU can be used to generate synthetic data and inspect configuration, but cannot run the current training and inference scripts. - Multi-GPU training uses the NCCL backend. Ensure that the device driver, communication libraries, and PyTorch version are compatible. - DCU users must install DTK in advance. DTK 25.04.2 or above, or the OneScience recommended version matching your cluster, is recommended. ### Download the Model Package ```bash hf download OneScience-Group/SEEDS --local-dir ./SEEDS cd SEEDS ``` ### Install the Runtime Environment **DCU Environment** ```bash # Please activate DTK and CONDA first conda create -n onescience311 python=3.11 -y conda activate onescience311 # uv installation is supported pip install onescience[earth-dcu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai ``` **GPU Environment** ```bash # Please activate CONDA first conda create -n onescience311 python=3.11 -y libstdcxx-ng=12 libgcc-ng=12 gcc_linux-64=12 gxx_linux-64=12 conda activate onescience311 # uv installation is supported pip install onescience[earth-gpu] -i http://mirrors.onescience.ai:3141/pypi/simple/ --trusted-host mirrors.onescience.ai ``` ### Training Data Introduction The SEEDS paper uses GEFS reforecast data for training, operational GEFS members as conditioning inputs, and ERA5 as the evaluation reference. The official data and preprocessing are not bundled with this repository; prepare the NPZ files specified by `conf/config.yaml` before training. ### Generate Synthetic Data When real data is unavailable, generate a default `6x48x48` cubed-sphere fixture. Synthetic data only validates the program flow and does not represent GEFS, ERA5, or the paper's forecast quality: ```bash python scripts/fake_data.py ``` ### Training Single GPU: ```bash python scripts/train.py ``` Multi-GPU: ```bash torchrun --nproc_per_node=8 scripts/train.py ``` The number of epochs is controlled by `conf/config.yaml`, and the default checkpoint is saved to `data/checkpoint/model_bak.pth`. ### Fine-tuning To continue from an existing checkpoint, use the explicit fine-tuning flag: ```bash python scripts/train.py --finetune ``` ### Training Weights This repository provides a `weight/` directory for checkpoints trained on the official data. The weight files will be uploaded soon and are expected to be available in the near future. ### Inference Inference reads `data/checkpoint/model_bak.pth` by default and generates ensemble members in batches controlled by `sampling.member_batch_size`: ```bash python scripts/inference.py ``` Predictions, targets, and seed members are saved under `result/output/`. ### Evaluation and Visualization ```bash python scripts/result.py ``` The script computes ensemble-mean RMSE, ACC, and empirical CRPS, saves the corresponding NPY metrics, and generates `result/forecast.png`. If training loss files are available, it also writes `result/loss.png`. # Official OneScience Resources | Platform | OneScience Main Repository | Skills Repository | | --- | --- | --- | | Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills | | GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills | # Citation and License - This repository is an independent adaptation of the SEEDS paper and is not an official Google product.