Model Card for ESFM/ESFM_s_nm_e11k_lt24h

ESFM small checkpoint directly finetuned for 24-hour forecasting from raw, non-imputed ECMWF-11K surface-station observations.

Checkpoint selection: Direct 24-hour forecasts for the documented ECMWF-11K station variables and station mapping.

Model Details

  • Developed by: The ESFM research team, with the full contributor and author lists linked below.
  • Shared by: ESFM on Hugging Face
  • Model type: Deterministic ECMWF-11K station checkpoint; modified 3D Swin-UNet encoder-decoder
  • Model size: Approximately 115 million parameters
  • Masking protocol: No additional masking during station finetuning
  • Ensemble members: 1
  • Forecast lead time: 24 hours
  • License: MIT
  • Repository: https://huggingface.co/ESFM/ESFM_s_nm_e11k_lt24h

Model Sources

The paper is currently available as an arXiv preprint.

Uses

Direct Use

Direct 24-hour forecasts for the documented ECMWF-11K station variables and station mapping.

Downstream Use

Research on day-ahead raw station forecasting and adaptation to other station networks after explicit remapping and finetuning.

Out-of-Scope Use

Do not treat this as four applications of the 6-hour station checkpoint; it is a separately direct-trained 24-hour model. Not valid for arbitrary station orderings or variable sets without reproducing preprocessing similar to the manuscript; not an operational station forecast service.

Bias, Risks, and Limitations

Errors grow with forecast horizon and under extrapolation to unseen locations. Surface pressure is the most challenging reported variable, and raw station missingness and geographic imbalance remain embedded in the data.

All ESFM checkpoints are research artifacts. Validate forecasts for the target variables, stations or regions, seasons, lead times, missingness pattern, and decision context. Do not use the model as the sole basis for safety-critical decisions.

How to Get Started

The checkpoint is not packaged as a Hugging Face Transformers from_pretrained model. Construct ESFM with the matching config and load the state dictionary. The released notebook demonstrates checkpoint download and architecture construction.

from huggingface_hub import hf_hub_download

model_name = "ESFM_s_nm_e11k_lt24h"
weights_path = hf_hub_download(
    repo_id=f"ESFM/{model_name}",
    filename=f"{model_name}.safetensors",
)
print(weights_path)

Use configs/config_ESFM_s_nm_e11k_lt24h.yaml; holdout evaluation uses configs/config_ESFM_s_nm_e11k_lt24h_ho.yaml and the matching scripts under scripts/inference/.

Clone the implementation first:

git clone https://github.com/swiss-ai/ESFM.git
cd ESFM

Training Details

Training Data

The ECMWF-11K dataset assembled from Copernicus surface-land observations v2.0.0. The manuscript reports 11,863 stations spanning 2000-2024, with 1,000 held out and no imputation.

Preprocessing and the exact variable registry are documented in the ESFM repository, preprocessing repository, and preprint.

Training Procedure

Initialized from the 30,000-step ESFM_s_nm_e11k_lt6h checkpoint and finetuned directly for a 24-hour target for 15,000 steps on 16 GPUs. Station values are not spatially interpolated during compact-grid mapping.

  • Nominal architecture: ESFM small, approximately 115M parameters
  • Software environment: PyTorch/Lightning in the released NVIDIA PhysicsNeMo 25.03 container; lightning==2.5.1 is pinned in the Dockerfile
  • Training regime: Lightning precision="32-true" with FP32 parameters and optimizer state; selected model forward operations use CUDA BF16 autocasting through torch.autocast(dtype=torch.bfloat16).

Evaluation

The manuscript reports that forecasts remain skillful at 24 hours across trained and held-out settings, including extrapolation without input at target locations. Detailed values remain in the linked preprint.

Detailed numerical results are intentionally not copied into this card.

Technical Specifications

ESFM uses variable-specific tokenization, axial attention across variables, perceiver aggregation, a 3D Swin-UNet backbone, and a decoder queried at target pressure levels. Missing patches are represented by learnable NaN tokens. Resolution-specific tokenizers and station mapping are enabled for the relevant sparse-data configs. The ensemble checkpoint additionally applies member-conditioned AdaLN-Zero after the backbone.

Environmental Impact

  • Hardware type: NVIDIA GH200 systems with four GPUs per node. This experiment was run on four nodes, totaling 16 GPUs.
  • Total training time: 12 hours
  • Compute location: Training used CSCS Alps infrastructure.

Citation

@misc{ozdemir2026esfm,
  title={Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting},
  author={Firat Ozdemir and Yun Cheng and Salman Mohebi and Fanny Lehmann and Simon Adamov and Zhenyi Zhang and Leonardo Trentini and Dana Grund and Oliver Fuhrer and Torsten Hoefler and Siddhartha Mishra and Sebastian Schemm and Benedikt Soja and Mathieu Salzmann},
  year={2026},
  eprint={2605.00850},
  archivePrefix={arXiv},
  primaryClass={physics.ao-ph},
  url={https://arxiv.org/abs/2605.00850}
}

More Information

Model Card Contact

Firat Ozdemir: firat.ozdemir@sdsc.ethz.ch

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ESFM/ESFM_s_nm_e11k_lt24h

Base model

microsoft/aurora
Finetuned
(2)
this model

Paper for ESFM/ESFM_s_nm_e11k_lt24h