SAR-StableDiffusion / README.md
sylviaHoch's picture
Update README.md
fe1bb67 verified
|
Raw
History Blame Contribute Delete
4.68 kB
metadata
license: cc-by-nc-sa-4.0
language:
  - en
library_name: diffusers
pipeline_tag: text-to-image
tags:
  - sar
  - synthetic-aperture-radar
  - remote-sensing
  - stable-diffusion
  - image-generation
  - synthetic-data
  - ship-detection
  - earth-observation
  - sentinel-1

SAR Stable Diffusion

Overview

Fine-tuned Stable Diffusion 1.5 backbone for generating synthetic Sentinel-1 SAR amplitude images from text prompts.
Published alongside the upcoming paper "Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations" (Hochstuhl et al., 2026; accepted for GCPR conference 2026).

The model generates single-channel SAR amplitude images for specific land cover types similar to Sentinel-1 GRD image products (in VH polarization).

For spatially controlled generation via ship/land masks, use this model together with
sylviaHoch/SAR-ControlNet.


Database

The model was fine-tuned on Sentinel-1 acquisitions from the SEN12MS dataset dataset. The corresponding text prompts were derived from the dataset's land cover maps, using the dominant land cover class within each image patch. Possible land cover classes are:

  • Evergreen Needleleaf Forest
  • Evergreen Broadleaf Forest
  • Deciduous Needleleaf Forest
  • Deciduous Broadleaf Forest
  • Mixed Forest
  • Closed Shrublands
  • Open Shrublands
  • Woody Savannas
  • Savannas
  • Grasslands
  • Permanent Wetlands
  • Croplands
  • Urban and Built-up
  • Cropland/Natural Vegetation Mosaic
  • Snow and Ice
  • Barren or Sparsely Vegetated
  • Water Bodies

Architecture

Component Details
Base model Stable Diffusion 1.5
VAE decoder Adapted to single-channel output
Text encoder Fine-tuned with a LoRA adapter (adapter_text_encoder/)
Output Single-channel float32 SAR amplitude image

Repository Structure

SAR-StableDiffusion/
β”œβ”€β”€ model_index.json
β”œβ”€β”€ unet/
β”œβ”€β”€ vae/                        # Modified: single-channel conv_out
β”œβ”€β”€ text_encoder/
β”œβ”€β”€ tokenizer/
β”œβ”€β”€ scheduler/
β”œβ”€β”€ feature_extractor/
└── adapter_text_encoder/       # LoRA adapter for text encoder
    β”œβ”€β”€ adapter_config.json
    └── adapter_model.safetensors

Usage

This model can be used directly with the πŸ€— diffusers pipeline. The LoRA adapter for the text encoder is included in this repository and must be loaded manually after initialising the pipeline.

Requirements:

  • torch (CUDA recommended)
  • diffusers
  • peft

Text-to-Image Generation

import torch
from diffusers import StableDiffusionPipeline
from peft import PeftModel

MODEL_ID = "sylviaHoch/SAR-StableDiffusion"

# Load pipeline
pipeline = StableDiffusionPipeline.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.float16
).to("cuda")

# Load LoRA text-encoder adapter
pipeline.text_encoder = PeftModel.from_pretrained(
    pipeline.text_encoder,
    MODEL_ID,
    subfolder="adapter_text_encoder"
)

# Generate
image = pipeline(
    "Woody Savannas.",
    num_inference_steps=50,
    guidance_scale=3.0,
    output_type="pt"
).images

image is a torch.Tensor of shape (N, C, H, W) with values in the range [0, 1] and dtype float32, where N is the batch size (here 1), C the number of channels (here 1, grayscale), and H and W the native model resolution (512 x 512).


Intended Use

  • Synthetic SAR data generation for given land cover types.
  • Research on generative models for remote sensing.

Limitations

  • Tuned specifically to Sentinel-1 sensor characteristics based on GRD products in VH polarization; transferability to other sensors and imaging properties is very limited.
  • Generated images are synthetic and may not fully capture all real SAR image properties.

License

This model is released under CC BY-NC-SA 4.0.
Commercial use is not permitted. Derivatives must be shared under the same license.
See LICENSE for details.