LDSR-S2 — Latent Diffusion Super-Resolution for Sentinel-2

LDSR-S2 is a latent diffusion model for ×4 spatial super-resolution of the 10 m RGB-NIR Sentinel-2 bands, producing imagery at a nominal 2.5 m spatial sampling.

This repository hosts the pretrained model weights used by ESAOpenSR/opensr-model, the reference implementation accompanying the paper:

Trustworthy Super-Resolution of Multispectral Sentinel-2 Imagery With Latent Diffusion Simon Donike, Cesar Aybar, Luis Gómez-Chova, Freddie Kalaitzis IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 18, 6940–6952, 2025. DOI: 10.1109/JSTARS.2025.3542220

LDSR-S2 architecture

Model summary

LDSR-S2 adapts latent diffusion to multispectral Earth-observation super-resolution. Instead of performing diffusion directly in image space, the model operates in a learned latent representation to make inference practical for large remote-sensing imagery.

The low-resolution Sentinel-2 observation is encoded and used to condition the diffusion process. This conditioning is intended to constrain the generated high-frequency spatial information using the observed multispectral image and improve consistency with the original measurement.

The model processes the four native 10 m Sentinel-2 bands:

Channel Sentinel-2 band Description
1 B04 Red
2 B03 Green
3 B02 Blue
4 B08 Near infrared

The expected channel order is therefore R, G, B, NIR (B04, B03, B02, B08).

Input and output

  • Input: Sentinel-2 L2A reflectance tensor
  • Bands: B04, B03, B02, B08
  • Channel order: R, G, B, NIR
  • Expected reflectance range: approximately [0, 1]
  • Input tensor: C × H × W or B × C × H × W
  • Number of channels: 4
  • Scale factor: ×4
  • Native input resolution: 10 m
  • Nominal output sampling: 2.5 m
  • Output tensor: C × 4H × 4W or B × C × 4H × 4W

The reference inference configuration uses 128 × 128 LR patches, corresponding to 512 × 512 SR patches.

Important: a nominal 2.5 m output pixel size does not imply that all reconstructed spatial information is equivalent to an independent 2.5 m physical measurement. LDSR-S2 is a learned generative reconstruction model.


Checkpoints

Recommended checkpoint

opensr-ldsrs2_v1_0_0.ckpt

This is the current LDSR-S2 v1.0.0 checkpoint and the checkpoint referenced by the current opensr-model configuration.

For normal use, this is the checkpoint you should use.

Legacy checkpoint

opensr_10m_v4_v6.ckpt

This checkpoint is retained for backwards compatibility and reproducibility of earlier experiments. New applications should use opensr-ldsrs2_v1_0_0.ckpt.

Example tensors

The repository additionally contains:

  • example_rural.pt
  • example_urban.pt

These are small example tensors intended for testing and demonstration rather than pretrained model parameters.


Installation

The recommended way to use these weights is through the opensr-model Python package:

pip install opensr-model

The package contains the LDSR-S2 architecture, inference code, configuration, preprocessing/postprocessing logic, and automatic checkpoint loading.


Usage

Recommended: load through opensr-model

from io import StringIO

import requests
import torch
from omegaconf import OmegaConf

import opensr_model

# Load the official LDSR-S2 configuration
config_url = (
    "https://raw.githubusercontent.com/ESAOpenSR/"
    "opensr-model/main/opensr_model/configs/config_10m.yaml"
)

response = requests.get(config_url)
response.raise_for_status()

config = OmegaConf.load(StringIO(response.text))

# Create model
device = "cuda" if torch.cuda.is_available() else "cpu"

model = opensr_model.SRLatentDiffusion(
    config,
    device=device,
)

# Downloads the corresponding pretrained checkpoint
model.load_pretrained(config.ckpt_version)

# Example:
# Sentinel-2 L2A reflectance in R,G,B,NIR order
# [B04, B03, B02, B08]
x = torch.rand(1, 4, 128, 128, device=device)

with torch.no_grad():
    sr = model(
        x,
        sampling_steps=100,
    )

print(sr.shape)
# torch.Size([1, 4, 512, 512])

The default configuration currently points to:

opensr-ldsrs2_v1_0_0.ckpt

Direct checkpoint download

If you need the checkpoint independently of the package, it can be downloaded from this Hugging Face repository:

from huggingface_hub import hf_hub_download

checkpoint = hf_hub_download(
    repo_id="simon-donike/RS-SR-LTDF",
    filename="opensr-ldsrs2_v1_0_0.ckpt",
)

print(checkpoint)

Loading the raw checkpoint directly requires the architecture defined in ESAOpenSR/opensr-model. For most users, model.load_pretrained(...) is therefore preferable.


Example results

LDSR-S2 example

Additional example output:

LDSR-S2 super-resolution example


Uncertainty estimation

LDSR-S2 is a probabilistic generative model. Repeated diffusion sampling can therefore be used to characterize variability between plausible reconstructions and derive pixel-level uncertainty estimates.

This is particularly relevant for Earth-observation applications because spatial detail generated by a super-resolution model is not necessarily directly observed in the Sentinel-2 input.

An example uncertainty product is shown below:

LDSR-S2 uncertainty map

See the opensr-model repository and the accompanying paper for the uncertainty methodology and example workflows.


Full Sentinel-2 scenes

The raw LDSR-S2 model performs tensor-level inference. It does not itself provide a complete geospatial processing pipeline for large .SAFE products, GeoTIFF tiling, reprojection, stitching, or metadata preservation.

For operational inference on full Sentinel-2 products or large raster files, use the OpenSR tooling:

The opensr-model repository also contains interactive Colab notebooks for end-to-end examples.


Training data

LDSR-S2 was developed as part of the ESA OpenSR project in conjunction with the SEN2NAIP Sentinel-2 super-resolution dataset.

SEN2NAIP provides paired and synthetic data designed for Sentinel-2 super-resolution research. For the exact training-data construction, preprocessing, experimental configuration, and geographic composition used for LDSR-S2, refer to the model paper and the SEN2NAIP publication.

Dataset:

isp-uv-es/SEN2NAIP

SEN2NAIP publication:

Cesar Aybar, David Montero, Julio Contreras, Simon Donike, Freddie Kalaitzis, Luis Gómez-Chova. SEN2NAIP: A large-scale dataset for Sentinel-2 Image Super-Resolution. Scientific Data, 11, 1389, 2024. DOI: 10.1038/s41597-024-04214-y


Intended uses

LDSR-S2 is intended for research and Earth-observation applications involving spatial enhancement of Sentinel-2 RGB-NIR imagery, including:

  • visualization and mapping;
  • development and evaluation of remote-sensing super-resolution methods;
  • analysis of the effect of spatial enhancement on downstream EO tasks;
  • uncertainty-aware super-resolution research;
  • generation of spatially enhanced Sentinel-2 RGB-NIR products.

The model is particularly intended for applications where preserving the information contained in the original multispectral observation is important.


Limitations and responsible use

LDSR-S2 is a generative super-resolution model. Its output should not be interpreted as a direct high-resolution measurement of the Earth's surface.

In particular:

  • Fine spatial structures may be reconstructed incorrectly.
  • Plausible high-frequency detail may be generated even when it is not uniquely determined by the LR observation.
  • Small objects may be omitted, displaced, altered, or synthesized.
  • Performance may degrade for geographic regions, land-cover classes, atmospheric conditions, acquisition conditions, or image statistics that differ substantially from the training distribution.
  • The 2.5 m output grid represents the model's ×4 reconstruction target and should not be interpreted as equivalent to a native 2.5 m sensor measurement.
  • Quantitative or high-stakes applications should validate the SR product for the specific downstream task and, where possible, use independent reference data.
  • Spectral correction and consistency with the LR input reduce some forms of radiometric inconsistency but do not guarantee that generated high-frequency spatial information is physically correct.

For applications requiring confidence information, multiple stochastic reconstructions and the uncertainty methodology described in the paper are recommended.

For scientific benchmarking, evaluation should include spatial, spectral/radiometric, and task-specific criteria rather than relying exclusively on perceptual image quality.


Model architecture

LDSR-S2 consists principally of:

  1. a learned autoencoder that maps four-band imagery into a latent representation;
  2. an encoded low-resolution Sentinel-2 observation used as conditioning;
  3. a conditional denoising U-Net operating in latent space;
  4. DDIM-based iterative sampling;
  5. decoding back to four-band image space;
  6. optional channel-wise histogram matching against the original LR observation.

The current v1.0 configuration uses:

  • four input/output channels;
  • latent dimensionality of 4;
  • 1000 diffusion training timesteps;
  • epsilon (eps) prediction;
  • DDIM inference;
  • 100 sampling steps by default;
  • ×4 spatial enlargement.

The implementation is available at:

https://github.com/ESAOpenSR/opensr-model


License

Model weights

The pretrained model weights distributed in this Hugging Face repository are released under the MIT License.

Copyright © OpenSR contributors.

Permission is granted under the terms of the MIT License to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the model weights, subject to the conditions of the license.

See the LICENSE file in this repository for the complete license text.

Source code

The reference ESAOpenSR/opensr-model implementation is also distributed under the MIT License. Portions of the latent-diffusion implementation are adapted from the CompVis latent-diffusion codebase and retain the corresponding MIT notice.


Citation

If you use LDSR-S2 or these pretrained weights in scientific work, please cite:

@article{donike2025trustworthy,
  author  = {Donike, Simon and Aybar, Cesar and G{\'o}mez-Chova, Luis and Kalaitzis, Freddie},
  title   = {Trustworthy Super-Resolution of Multispectral Sentinel-2 Imagery With Latent Diffusion},
  journal = {IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing},
  volume  = {18},
  pages   = {6940--6952},
  year    = {2025},
  doi     = {10.1109/JSTARS.2025.3542220}
}

Resources


Acknowledgements

LDSR-S2 was developed within the OpenSR initiative, supported by the European Space Agency (ESA), with contributions from the Image Processing Laboratory at the Universitat de València and project collaborators.

The latent-diffusion implementation includes adaptations of components originally developed by the CompVis group at LMU Munich.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train simon-donike/RS-SR-LTDF