Model Card for ESFM/ESFM_s_wm_pre
Released intermediate checkpoint in the default masked ERA5 lineage, initialized from the knowledge-distilled ESFM encoder. It learns six-hour forecasting while variables, pressure levels, or contiguous regions are withheld from inputs.
Checkpoint selection: Use for reproducing the released masked training lineage or inspecting the checkpoint before its final continuation. Prefer
ESFM_s_wmfor the final default masked model.
Model Details
- Developed by: The ESFM research team, with the full contributor and author lists linked below.
- Shared by: ESFM on Hugging Face
- Model type: Intermediate deterministic masked ERA5 checkpoint; modified 3D Swin-UNet encoder-decoder
- Model size: Approximately 115 million parameters
- Masking protocol: Variable, pressure-level, and spatial masking
- Forecast lead time: 6 hours
- License: MIT
- Repository: https://huggingface.co/ESFM/ESFM_s_wm_pre
Model Sources
- Code: https://github.com/swiss-ai/ESFM
- Paper: https://arxiv.org/abs/2605.00850
- Project page: https://swiss-ai.github.io/ESFM/
The paper is currently available as an arXiv preprint.
Uses
Direct Use
Use for reproducing the released masked training lineage or inspecting the checkpoint before its final continuation. Prefer ESFM_s_wm for the final default masked model.
Downstream Use
Base checkpoint for the released ESFM_s_wm continuation and for the MODIS, ECMWF-11K, and Weather-5K specializations documented in the repository experiment guide.
Out-of-Scope Use
Not a substitute for the final checkpoint when reporting manuscript-level performance. It should not be treated as an operational forecast service.
Bias, Risks, and Limitations
Missing-input skill depends on the training distribution and masking protocol; entire withheld pressure levels are particularly difficult. The model inherits ERA5 biases and long-rollout limitations.
All ESFM checkpoints are research artifacts. Users should validate forecasts for their variables, regions, seasons, lead times, missingness pattern, and decision context. Do not use the model as the sole basis for safety-critical decisions.
How to Get Started
The checkpoint is not packaged as a Hugging Face Transformers from_pretrained model. Construct the ESFM architecture with the matching repository config, then load the state dictionary. The released notebook contains the complete download, model-construction, normalization, and inference workflow.
git clone https://github.com/swiss-ai/ESFM.git
cd ESFM
# Open notebooks/inference_ESFMs_on_ERA5.ipynb
In the notebook, set:
EXPERIMENT_NAME = "ESFM_s_wm_pre"
To download the weights directly:
from huggingface_hub import hf_hub_download
model_name = "ESFM_s_wm_pre"
weights_path = hf_hub_download(
repo_id=f"ESFM/{model_name}",
filename=f"{model_name}.safetensors",
)
print(weights_path)
Set EXPERIMENT_NAME = "ESFM_s_wm_pre" in notebooks/inference_ESFMs_on_ERA5.ipynb, or use configs/config_ESFM_s_wm_pre.yaml with the released inference code.
Training Details
Training Data
WeatherBench2 ERA5 at 0.25-degree resolution, trained on 1979 through 2020 with the manuscript's variable/level/spatial masking protocol.
Dataset preprocessing and the exact variable registry are documented in the ESFM repository and preprint.
Training Procedure
Initialized from ESFM_s_enc_KD_nm and trained for 100,000 steps on ERA5 with the missing-data masking protocol. This experiment was run on 16 GPUs. The detailed masking probabilities and conditional ratios are documented in the linked preprint and released masking configuration.
- Training objective: Six-hour forecast learning, as specified above
- Nominal architecture: ESFM small, approximately 115M parameters
- Software environment: PyTorch/Lightning in the released NVIDIA PhysicsNeMo 25.03 container; lightning==2.5.1 is pinned in the Dockerfile
- Training regime: Lightning
precision="32-true"with FP32 parameters and optimizer state; selected model forward operations use CUDA BF16 autocasting throughtorch.autocast(dtype=torch.bfloat16).
Evaluation
The manuscript evaluates the completed masked lineage under dense and deliberately withheld-input settings. This intermediate checkpoint is not separately tabulated.
The manuscript uses held-out temporal data and reports task-appropriate metrics: latitude-weighted MAE and Pearson correlation for gridded deterministic forecasts, relative MAE for MODIS comparisons, station metrics for station models, and CRPS for ensembles. Detailed values are intentionally not copied into this card.
Technical Specifications
ESFM retains Aurora's 3D Swin-UNet backbone and adds variable-specific tokenization, axial attention across variables, perceiver aggregation across variables and pressure levels, learnable NaN tokens for missing patches, resolution-specific tokenizers where configured, and a decoder queried at target pressure levels. The small configuration uses a 256-dimensional embedding and approximately 115M parameters.
Environmental Impact
- Hardware type: NVIDIA GH200 systems with four GPUs per node. This experiment was run on four nodes, totaling 16 GPUs.
- Total training time: 200 hours
- Compute location: Training used CSCS Alps infrastructure.
Citation
@misc{ozdemir2026esfm,
title={Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting},
author={Firat Ozdemir and Yun Cheng and Salman Mohebi and Fanny Lehmann and Simon Adamov and Zhenyi Zhang and Leonardo Trentini and Dana Grund and Oliver Fuhrer and Torsten Hoefler and Siddhartha Mishra and Sebastian Schemm and Benedikt Soja and Mathieu Salzmann},
year={2026},
eprint={2605.00850},
archivePrefix={arXiv},
primaryClass={physics.ao-ph},
url={https://arxiv.org/abs/2605.00850}
}
More Information
Model Card Contact
Firat Ozdemir: firat.ozdemir@sdsc.ethz.ch