Model Card for ESFM/ESFM_s_nm_e11k_lt12h

ESFM small checkpoint directly finetuned for 12-hour forecasting from raw, non-imputed ECMWF-11K surface-station observations.

Checkpoint selection: Direct 12-hour forecasts for the documented ECMWF-11K station variables and station mapping.

Model Details

  • Developed by: The ESFM research team, with the full contributor and author lists linked below.
  • Shared by: ESFM on Hugging Face
  • Model type: Deterministic ECMWF-11K station checkpoint; modified 3D Swin-UNet encoder-decoder
  • Model size: Approximately 115 million parameters
  • Masking protocol: No additional masking during station finetuning
  • Ensemble members: 1
  • Forecast lead time: 12 hours
  • License: MIT
  • Repository: https://huggingface.co/ESFM/ESFM_s_nm_e11k_lt12h

Model Sources

The paper is currently available as an arXiv preprint.

Uses

Direct Use

Direct 12-hour forecasts for the documented ECMWF-11K station variables and station mapping.

Downstream Use

Research on raw station forecasting at a 12-hour horizon or adaptation to new networks after explicit remapping and finetuning.

Out-of-Scope Use

Do not treat this as two applications of the 6-hour checkpoint; it is a separately direct-trained 12-hour model. Not valid for arbitrary station orderings or variable sets without reproducing preprocessing similar to the manuscript; not an operational station forecast service.

Bias, Risks, and Limitations

Station coverage and missingness are uneven, held-out-location extrapolation is harder than forecasting trained stations, and surface pressure is the weakest reported variable under extrapolation.

All ESFM checkpoints are research artifacts. Validate forecasts for the target variables, stations or regions, seasons, lead times, missingness pattern, and decision context. Do not use the model as the sole basis for safety-critical decisions.

How to Get Started

The checkpoint is not packaged as a Hugging Face Transformers from_pretrained model. Construct ESFM with the matching config and load the state dictionary. The released notebook demonstrates checkpoint download and architecture construction.

from huggingface_hub import hf_hub_download

model_name = "ESFM_s_nm_e11k_lt12h"
weights_path = hf_hub_download(
    repo_id=f"ESFM/{model_name}",
    filename=f"{model_name}.safetensors",
)
print(weights_path)

Use configs/config_ESFM_s_nm_e11k_lt12h.yaml; holdout evaluation uses configs/config_ESFM_s_nm_e11k_lt12h_ho.yaml and the matching scripts under scripts/inference/.

Clone the implementation first:

git clone https://github.com/swiss-ai/ESFM.git
cd ESFM

Training Details

Training Data

The ECMWF-11K dataset assembled from Copernicus surface-land observations v2.0.0. The manuscript reports 11,863 stations spanning 2000-2024, with 1,000 held out and no imputation.

Preprocessing and the exact variable registry are documented in the ESFM repository, preprocessing repository, and preprint.

Training Procedure

Initialized from the 30,000-step ESFM_s_nm_e11k_lt6h checkpoint and finetuned directly for a 12-hour target for 15,000 steps on 16 GPUs. Station values are not spatially interpolated during the compact-grid mapping.

  • Nominal architecture: ESFM small, approximately 115M parameters
  • Software environment: PyTorch/Lightning in the released NVIDIA PhysicsNeMo 25.03 container; lightning==2.5.1 is pinned in the Dockerfile
  • Training regime: Lightning precision="32-true" with FP32 parameters and optimizer state; selected model forward operations use CUDA BF16 autocasting through torch.autocast(dtype=torch.bfloat16).

Evaluation

The manuscript evaluates trained, held-out-input, and no-target-input settings for the ECMWF-11K lead-time family. Results are summarized qualitatively here and linked to the current preprint.

Detailed numerical results are intentionally not copied into this card.

Technical Specifications

ESFM uses variable-specific tokenization, axial attention across variables, perceiver aggregation, a 3D Swin-UNet backbone, and a decoder queried at target pressure levels. Missing patches are represented by learnable NaN tokens. Resolution-specific tokenizers and station mapping are enabled for the relevant sparse-data configs. The ensemble checkpoint additionally applies member-conditioned AdaLN-Zero after the backbone.

Environmental Impact

  • Hardware type: NVIDIA GH200 systems with four GPUs per node. This experiment was run on four nodes, totaling 16 GPUs.
  • Total training time: 12 hours
  • Compute location: Training used CSCS Alps infrastructure.

Citation

@misc{ozdemir2026esfm,
  title={Earth System Foundation Model (ESFM): A unified framework for heterogeneous data integration and forecasting},
  author={Firat Ozdemir and Yun Cheng and Salman Mohebi and Fanny Lehmann and Simon Adamov and Zhenyi Zhang and Leonardo Trentini and Dana Grund and Oliver Fuhrer and Torsten Hoefler and Siddhartha Mishra and Sebastian Schemm and Benedikt Soja and Mathieu Salzmann},
  year={2026},
  eprint={2605.00850},
  archivePrefix={arXiv},
  primaryClass={physics.ao-ph},
  url={https://arxiv.org/abs/2605.00850}
}

More Information

Model Card Contact

Firat Ozdemir: firat.ozdemir@sdsc.ethz.ch

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ESFM/ESFM_s_nm_e11k_lt12h

Base model

microsoft/aurora
Finetuned
(2)
this model

Paper for ESFM/ESFM_s_nm_e11k_lt12h