Model Card for ESFM/ESFM_s_nm_e11k_lt6h
ESFM small checkpoint directly finetuned for six-hour forecasting from raw, non-imputed ECMWF-11K surface-station observations.
Checkpoint selection: Six-hour forecasts for the documented ECMWF-11K station variables and preprocessing layout, including the manuscript's trained-station and holdout evaluation settings.
Model Details
- Developed by: The ESFM research team, with the full contributor and author lists linked below.
- Shared by: ESFM on Hugging Face
- Model type: Deterministic ECMWF-11K station checkpoint; modified 3D Swin-UNet encoder-decoder
- Model size: Approximately 115 million parameters
- Masking protocol: No additional masking during station finetuning
- Ensemble members: 1
- Forecast lead time: 6 hours
- License: MIT
- Repository: https://huggingface.co/ESFM/ESFM_s_nm_e11k_lt6h
Model Sources
- Code: https://github.com/swiss-ai/ESFM
- Paper: https://arxiv.org/abs/2605.00850
- Project page: https://swiss-ai.github.io/ESFM/
The paper is currently available as an arXiv preprint.
Uses
Direct Use
Six-hour forecasts for the documented ECMWF-11K station variables and preprocessing layout, including the manuscript's trained-station and holdout evaluation settings.
Downstream Use
Research on raw irregular station forecasting and adaptation to other station networks after rebuilding the station map and finetuning.
Out-of-Scope Use
Not valid for arbitrary station orderings or variable sets without reproducing preprocessing similar to the manuscript; not an operational station forecast service.
Bias, Risks, and Limitations
Station coverage and missingness are geographically and temporally uneven. Generalization degrades at held-out locations, especially when no input is available at target coordinates; surface pressure is the weakest reported variable under extrapolation.
All ESFM checkpoints are research artifacts. Validate forecasts for the target variables, stations or regions, seasons, lead times, missingness pattern, and decision context. Do not use the model as the sole basis for safety-critical decisions.
How to Get Started
The checkpoint is not packaged as a Hugging Face Transformers from_pretrained model. Construct ESFM with the matching config and load the state dictionary. The released notebook demonstrates checkpoint download and architecture construction.
from huggingface_hub import hf_hub_download
model_name = "ESFM_s_nm_e11k_lt6h"
weights_path = hf_hub_download(
repo_id=f"ESFM/{model_name}",
filename=f"{model_name}.safetensors",
)
print(weights_path)
This checkpoint requires locally preprocessed station data. Use configs/config_ESFM_s_nm_e11k_lt6h.yaml; holdout evaluation uses configs/config_ESFM_s_nm_e11k_lt6h_ho.yaml and the matching scripts under scripts/inference/.
Clone the implementation first:
git clone https://github.com/swiss-ai/ESFM.git
cd ESFM
Training Details
Training Data
The ECMWF-11K dataset assembled from Copernicus surface-land observations v2.0.0. The manuscript reports 11,863 stations spanning 2000-2024, with 1,000 stations held out from training, and no imputation.
Preprocessing and the exact variable registry are documented in the ESFM repository, preprocessing repository, and preprint.
Training Procedure
Initialized from the masked ERA5 lineage and finetuned for 30,000 steps on 4 GPUs without an added masking protocol because the raw station data are already highly sparse. Stations are mapped without value interpolation to a compact 90x180 grid while retaining their coordinates for positional encoding.
- Nominal architecture: ESFM small, approximately 115M parameters
- Software environment: PyTorch/Lightning in the released NVIDIA PhysicsNeMo 25.03 container; lightning==2.5.1 is pinned in the Dockerfile
- Training regime: Lightning
precision="32-true"with FP32 parameters and optimizer state; selected model forward operations use CUDA BF16 autocasting throughtorch.autocast(dtype=torch.bfloat16).
Evaluation
The manuscript evaluates trained stations, held-out stations supplied at inference, and extrapolation to held-out locations without target-position input. It reports useful skill through 24 hours for the corresponding lead-time models; detailed station metrics remain in the preprint.
Detailed numerical results are intentionally not copied into this card.
Technical Specifications
ESFM uses variable-specific tokenization, axial attention across variables, perceiver aggregation, a 3D Swin-UNet backbone, and a decoder queried at target pressure levels. Missing patches are represented by learnable NaN tokens. Resolution-specific tokenizers and station mapping are enabled for the relevant sparse-data configs. The ensemble checkpoint additionally applies member-conditioned AdaLN-Zero after the backbone.
Environmental Impact
- Hardware type: NVIDIA GH200 systems with four GPUs per node. This experiment was run on four nodes, totaling 16 GPUs.
- Total training time: 24 hours
- Compute location: Training used CSCS Alps infrastructure.
Citation
@misc{ozdemir2026esfm,
title={Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting},
author={Firat Ozdemir and Yun Cheng and Salman Mohebi and Fanny Lehmann and Simon Adamov and Zhenyi Zhang and Leonardo Trentini and Dana Grund and Oliver Fuhrer and Torsten Hoefler and Siddhartha Mishra and Sebastian Schemm and Benedikt Soja and Mathieu Salzmann},
year={2026},
eprint={2605.00850},
archivePrefix={arXiv},
primaryClass={physics.ao-ph},
url={https://arxiv.org/abs/2605.00850}
}
More Information
- ESFM code and configurations
- Current ESFM preprint
- Data preprocessing scripts
- ECMWF-11K dataset
- Project page
Model Card Contact
Firat Ozdemir: firat.ozdemir@sdsc.ethz.ch