LDSR-S2 — Latent Diffusion Super-Resolution for Sentinel-2
LDSR-S2 is a latent diffusion model for ×4 spatial super-resolution of the 10 m RGB-NIR Sentinel-2 bands, producing imagery at a nominal 2.5 m spatial sampling.
This repository hosts the pretrained model weights used by ESAOpenSR/opensr-model, the reference implementation accompanying the paper:
Trustworthy Super-Resolution of Multispectral Sentinel-2 Imagery With Latent Diffusion Simon Donike, Cesar Aybar, Luis Gómez-Chova, Freddie Kalaitzis IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 18, 6940–6952, 2025. DOI: 10.1109/JSTARS.2025.3542220
Model summary
LDSR-S2 adapts latent diffusion to multispectral Earth-observation super-resolution. Instead of performing diffusion directly in image space, the model operates in a learned latent representation to make inference practical for large remote-sensing imagery.
The low-resolution Sentinel-2 observation is encoded and used to condition the diffusion process. This conditioning is intended to constrain the generated high-frequency spatial information using the observed multispectral image and improve consistency with the original measurement.
The model processes the four native 10 m Sentinel-2 bands:
| Channel | Sentinel-2 band | Description |
|---|---|---|
| 1 | B04 | Red |
| 2 | B03 | Green |
| 3 | B02 | Blue |
| 4 | B08 | Near infrared |
The expected channel order is therefore R, G, B, NIR (B04, B03, B02, B08).
Input and output
- Input: Sentinel-2 L2A reflectance tensor
- Bands: B04, B03, B02, B08
- Channel order: R, G, B, NIR
- Expected reflectance range: approximately
[0, 1] - Input tensor:
C × H × WorB × C × H × W - Number of channels: 4
- Scale factor: ×4
- Native input resolution: 10 m
- Nominal output sampling: 2.5 m
- Output tensor:
C × 4H × 4WorB × C × 4H × 4W
The reference inference configuration uses 128 × 128 LR patches, corresponding to 512 × 512 SR patches.
Important: a nominal 2.5 m output pixel size does not imply that all reconstructed spatial information is equivalent to an independent 2.5 m physical measurement. LDSR-S2 is a learned generative reconstruction model.
Checkpoints
Recommended checkpoint
opensr-ldsrs2_v1_0_0.ckpt
This is the current LDSR-S2 v1.0.0 checkpoint and the checkpoint referenced by the current opensr-model configuration.
For normal use, this is the checkpoint you should use.
Legacy checkpoint
opensr_10m_v4_v6.ckpt
This checkpoint is retained for backwards compatibility and reproducibility of earlier experiments. New applications should use opensr-ldsrs2_v1_0_0.ckpt.
Example tensors
The repository additionally contains:
example_rural.ptexample_urban.pt
These are small example tensors intended for testing and demonstration rather than pretrained model parameters.
Installation
The recommended way to use these weights is through the opensr-model Python package:
pip install opensr-model
The package contains the LDSR-S2 architecture, inference code, configuration, preprocessing/postprocessing logic, and automatic checkpoint loading.
Usage
Recommended: load through opensr-model
from io import StringIO
import requests
import torch
from omegaconf import OmegaConf
import opensr_model
# Load the official LDSR-S2 configuration
config_url = (
"https://raw.githubusercontent.com/ESAOpenSR/"
"opensr-model/main/opensr_model/configs/config_10m.yaml"
)
response = requests.get(config_url)
response.raise_for_status()
config = OmegaConf.load(StringIO(response.text))
# Create model
device = "cuda" if torch.cuda.is_available() else "cpu"
model = opensr_model.SRLatentDiffusion(
config,
device=device,
)
# Downloads the corresponding pretrained checkpoint
model.load_pretrained(config.ckpt_version)
# Example:
# Sentinel-2 L2A reflectance in R,G,B,NIR order
# [B04, B03, B02, B08]
x = torch.rand(1, 4, 128, 128, device=device)
with torch.no_grad():
sr = model(
x,
sampling_steps=100,
)
print(sr.shape)
# torch.Size([1, 4, 512, 512])
The default configuration currently points to:
opensr-ldsrs2_v1_0_0.ckpt
Direct checkpoint download
If you need the checkpoint independently of the package, it can be downloaded from this Hugging Face repository:
from huggingface_hub import hf_hub_download
checkpoint = hf_hub_download(
repo_id="simon-donike/RS-SR-LTDF",
filename="opensr-ldsrs2_v1_0_0.ckpt",
)
print(checkpoint)
Loading the raw checkpoint directly requires the architecture defined in ESAOpenSR/opensr-model. For most users, model.load_pretrained(...) is therefore preferable.
Example results
Additional example output:
Uncertainty estimation
LDSR-S2 is a probabilistic generative model. Repeated diffusion sampling can therefore be used to characterize variability between plausible reconstructions and derive pixel-level uncertainty estimates.
This is particularly relevant for Earth-observation applications because spatial detail generated by a super-resolution model is not necessarily directly observed in the Sentinel-2 input.
An example uncertainty product is shown below:
See the opensr-model repository and the accompanying paper for the uncertainty methodology and example workflows.
Full Sentinel-2 scenes
The raw LDSR-S2 model performs tensor-level inference. It does not itself provide a complete geospatial processing pipeline for large .SAFE products, GeoTIFF tiling, reprojection, stitching, or metadata preservation.
For operational inference on full Sentinel-2 products or large raster files, use the OpenSR tooling:
ESAOpenSR/opensr-model— LDSR-S2 architecture and inferenceESAOpenSR/opensr-utils— tiled large-image inference and geospatial I/OESAOpenSR/SEN2SR— complementary processing including Sentinel-2 20 m bandsESAOpenSR/opensr-test— remote-sensing SR evaluation
The opensr-model repository also contains interactive Colab notebooks for end-to-end examples.
Training data
LDSR-S2 was developed as part of the ESA OpenSR project in conjunction with the SEN2NAIP Sentinel-2 super-resolution dataset.
SEN2NAIP provides paired and synthetic data designed for Sentinel-2 super-resolution research. For the exact training-data construction, preprocessing, experimental configuration, and geographic composition used for LDSR-S2, refer to the model paper and the SEN2NAIP publication.
Dataset:
SEN2NAIP publication:
Cesar Aybar, David Montero, Julio Contreras, Simon Donike, Freddie Kalaitzis, Luis Gómez-Chova. SEN2NAIP: A large-scale dataset for Sentinel-2 Image Super-Resolution. Scientific Data, 11, 1389, 2024. DOI: 10.1038/s41597-024-04214-y
Intended uses
LDSR-S2 is intended for research and Earth-observation applications involving spatial enhancement of Sentinel-2 RGB-NIR imagery, including:
- visualization and mapping;
- development and evaluation of remote-sensing super-resolution methods;
- analysis of the effect of spatial enhancement on downstream EO tasks;
- uncertainty-aware super-resolution research;
- generation of spatially enhanced Sentinel-2 RGB-NIR products.
The model is particularly intended for applications where preserving the information contained in the original multispectral observation is important.
Limitations and responsible use
LDSR-S2 is a generative super-resolution model. Its output should not be interpreted as a direct high-resolution measurement of the Earth's surface.
In particular:
- Fine spatial structures may be reconstructed incorrectly.
- Plausible high-frequency detail may be generated even when it is not uniquely determined by the LR observation.
- Small objects may be omitted, displaced, altered, or synthesized.
- Performance may degrade for geographic regions, land-cover classes, atmospheric conditions, acquisition conditions, or image statistics that differ substantially from the training distribution.
- The 2.5 m output grid represents the model's ×4 reconstruction target and should not be interpreted as equivalent to a native 2.5 m sensor measurement.
- Quantitative or high-stakes applications should validate the SR product for the specific downstream task and, where possible, use independent reference data.
- Spectral correction and consistency with the LR input reduce some forms of radiometric inconsistency but do not guarantee that generated high-frequency spatial information is physically correct.
For applications requiring confidence information, multiple stochastic reconstructions and the uncertainty methodology described in the paper are recommended.
For scientific benchmarking, evaluation should include spatial, spectral/radiometric, and task-specific criteria rather than relying exclusively on perceptual image quality.
Model architecture
LDSR-S2 consists principally of:
- a learned autoencoder that maps four-band imagery into a latent representation;
- an encoded low-resolution Sentinel-2 observation used as conditioning;
- a conditional denoising U-Net operating in latent space;
- DDIM-based iterative sampling;
- decoding back to four-band image space;
- optional channel-wise histogram matching against the original LR observation.
The current v1.0 configuration uses:
- four input/output channels;
- latent dimensionality of 4;
- 1000 diffusion training timesteps;
- epsilon (
eps) prediction; - DDIM inference;
- 100 sampling steps by default;
- ×4 spatial enlargement.
The implementation is available at:
https://github.com/ESAOpenSR/opensr-model
License
Model weights
The pretrained model weights distributed in this Hugging Face repository are released under the MIT License.
Copyright © OpenSR contributors.
Permission is granted under the terms of the MIT License to use, copy, modify, merge, publish, distribute, sublicense, and/or sell copies of the model weights, subject to the conditions of the license.
See the LICENSE file in this repository for the complete license text.
Source code
The reference ESAOpenSR/opensr-model implementation is also distributed under the MIT License. Portions of the latent-diffusion implementation are adapted from the CompVis latent-diffusion codebase and retain the corresponding MIT notice.
Citation
If you use LDSR-S2 or these pretrained weights in scientific work, please cite:
@article{donike2025trustworthy,
author = {Donike, Simon and Aybar, Cesar and G{\'o}mez-Chova, Luis and Kalaitzis, Freddie},
title = {Trustworthy Super-Resolution of Multispectral Sentinel-2 Imagery With Latent Diffusion},
journal = {IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing},
volume = {18},
pages = {6940--6952},
year = {2025},
doi = {10.1109/JSTARS.2025.3542220}
}
Resources
- Source code: https://github.com/ESAOpenSR/opensr-model
- OpenSR: https://opensr.eu/
- Paper: https://doi.org/10.1109/JSTARS.2025.3542220
- SEN2NAIP: https://huggingface.co/datasets/isp-uv-es/SEN2NAIP
- OpenSR Utils: https://github.com/ESAOpenSR/opensr-utils
- OpenSR Test: https://github.com/ESAOpenSR/opensr-test
- SEN2SR: https://github.com/ESAOpenSR/SEN2SR
Acknowledgements
LDSR-S2 was developed within the OpenSR initiative, supported by the European Space Agency (ESA), with contributions from the Image Processing Laboratory at the Universitat de València and project collaborators.
The latent-diffusion implementation includes adaptations of components originally developed by the CompVis group at LMU Munich.