--- license: cc-by-nc-sa-4.0 language: - en library_name: diffusers pipeline_tag: text-to-image tags: - sar - synthetic-aperture-radar - remote-sensing - stable-diffusion - image-generation - synthetic-data - ship-detection - earth-observation - sentinel-1 --- # SAR Stable Diffusion ## Overview Fine-tuned **Stable Diffusion 1.5** backbone for generating synthetic **Sentinel-1 SAR** amplitude images from text prompts. Published alongside the upcoming paper **"Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations"** (Hochstuhl et al., 2026; accepted for GCPR conference 2026). The model generates single-channel SAR amplitude images for specific land cover types similar to Sentinel-1 GRD image products (in VH polarization). > For spatially controlled generation via ship/land masks, use this model together with > [`sylviaHoch/SAR-ControlNet`](https://huggingface.co/sylviaHoch/SAR-ControlNet). --- ## Database The model was fine-tuned on Sentinel-1 acquisitions from the [`SEN12MS dataset`](https://github.com/schmitt-muc/SEN12MS) dataset. The corresponding text prompts were derived from the dataset's land cover maps, using the dominant land cover class within each image patch. Possible land cover classes are: - Evergreen Needleleaf Forest - Evergreen Broadleaf Forest - Deciduous Needleleaf Forest - Deciduous Broadleaf Forest - Mixed Forest - Closed Shrublands - Open Shrublands - Woody Savannas - Savannas - Grasslands - Permanent Wetlands - Croplands - Urban and Built-up - Cropland/Natural Vegetation Mosaic - Snow and Ice - Barren or Sparsely Vegetated - Water Bodies --- ## Architecture | Component | Details | |---|---| | Base model | Stable Diffusion 1.5 | | VAE decoder | Adapted to **single-channel** output | | Text encoder | Fine-tuned with a **LoRA adapter** (`adapter_text_encoder/`) | | Output | Single-channel float32 SAR amplitude image | --- ## Repository Structure ```text SAR-StableDiffusion/ ├── model_index.json ├── unet/ ├── vae/ # Modified: single-channel conv_out ├── text_encoder/ ├── tokenizer/ ├── scheduler/ ├── feature_extractor/ └── adapter_text_encoder/ # LoRA adapter for text encoder ├── adapter_config.json └── adapter_model.safetensors ``` --- ## Usage This model can be used directly with the 🤗 [diffusers](https://github.com/huggingface/diffusers) pipeline. The LoRA adapter for the text encoder is included in this repository and must be loaded manually after initialising the pipeline. **Requirements:** - `torch` (CUDA recommended) - `diffusers` - `peft` ### Text-to-Image Generation ```python import torch from diffusers import StableDiffusionPipeline from peft import PeftModel MODEL_ID = "sylviaHoch/SAR-StableDiffusion" # Load pipeline pipeline = StableDiffusionPipeline.from_pretrained( MODEL_ID, torch_dtype=torch.float16 ).to("cuda") # Load LoRA text-encoder adapter pipeline.text_encoder = PeftModel.from_pretrained( pipeline.text_encoder, MODEL_ID, subfolder="adapter_text_encoder" ) # Generate image = pipeline( "Woody Savannas.", num_inference_steps=50, guidance_scale=3.0, output_type="pt" ).images ``` >image is a torch.Tensor of shape (N, C, H, W) with values in the range [0, 1] and dtype float32, where N is the batch size (here 1), C the number of channels (here 1, grayscale), and H and W the native model resolution (512 x 512). --- ## Intended Use - Synthetic SAR data generation for given land cover types. - Research on generative models for remote sensing. ## Limitations - Tuned specifically to **Sentinel-1** sensor characteristics based on GRD products in VH polarization; transferability to other sensors and imaging properties is very limited. - Generated images are synthetic and may not fully capture all real SAR image properties. --- ## License This model is released under **CC BY-NC-SA 4.0**. Commercial use is not permitted. Derivatives must be shared under the same license. See [LICENSE](https://creativecommons.org/licenses/by-nc-sa/4.0/) for details.