SAR-StableDiffusion / README.md
sylviaHoch's picture
Update README.md
fe1bb67 verified
|
Raw
History Blame Contribute Delete
4.68 kB
---
license: cc-by-nc-sa-4.0
language:
- en
library_name: diffusers
pipeline_tag: text-to-image
tags:
- sar
- synthetic-aperture-radar
- remote-sensing
- stable-diffusion
- image-generation
- synthetic-data
- ship-detection
- earth-observation
- sentinel-1
---
# SAR Stable Diffusion
## Overview
Fine-tuned **Stable Diffusion 1.5** backbone for generating synthetic **Sentinel-1 SAR** amplitude images from text prompts.
Published alongside the upcoming paper **"Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations"** (Hochstuhl et al., 2026; accepted for GCPR conference 2026).
The model generates single-channel SAR amplitude images for specific land cover types similar to Sentinel-1 GRD image products (in VH polarization).
> For spatially controlled generation via ship/land masks, use this model together with
> [`sylviaHoch/SAR-ControlNet`](https://huggingface.co/sylviaHoch/SAR-ControlNet).
---
## Database
The model was fine-tuned on Sentinel-1 acquisitions from the [`SEN12MS dataset`](https://github.com/schmitt-muc/SEN12MS) dataset. The corresponding text prompts were derived from the dataset's land cover maps, using the dominant land cover class within each image patch. Possible land cover classes are:
- Evergreen Needleleaf Forest
- Evergreen Broadleaf Forest
- Deciduous Needleleaf Forest
- Deciduous Broadleaf Forest
- Mixed Forest
- Closed Shrublands
- Open Shrublands
- Woody Savannas
- Savannas
- Grasslands
- Permanent Wetlands
- Croplands
- Urban and Built-up
- Cropland/Natural Vegetation Mosaic
- Snow and Ice
- Barren or Sparsely Vegetated
- Water Bodies
---
## Architecture
| Component | Details |
|---|---|
| Base model | Stable Diffusion 1.5 |
| VAE decoder | Adapted to **single-channel** output |
| Text encoder | Fine-tuned with a **LoRA adapter** (`adapter_text_encoder/`) |
| Output | Single-channel float32 SAR amplitude image |
---
## Repository Structure
```text
SAR-StableDiffusion/
β”œβ”€β”€ model_index.json
β”œβ”€β”€ unet/
β”œβ”€β”€ vae/ # Modified: single-channel conv_out
β”œβ”€β”€ text_encoder/
β”œβ”€β”€ tokenizer/
β”œβ”€β”€ scheduler/
β”œβ”€β”€ feature_extractor/
└── adapter_text_encoder/ # LoRA adapter for text encoder
β”œβ”€β”€ adapter_config.json
└── adapter_model.safetensors
```
---
## Usage
This model can be used directly with the πŸ€— [diffusers](https://github.com/huggingface/diffusers) pipeline.
The LoRA adapter for the text encoder is included in this repository and must be loaded manually after initialising the pipeline.
**Requirements:**
- `torch` (CUDA recommended)
- `diffusers`
- `peft`
### Text-to-Image Generation
```python
import torch
from diffusers import StableDiffusionPipeline
from peft import PeftModel
MODEL_ID = "sylviaHoch/SAR-StableDiffusion"
# Load pipeline
pipeline = StableDiffusionPipeline.from_pretrained(
MODEL_ID,
torch_dtype=torch.float16
).to("cuda")
# Load LoRA text-encoder adapter
pipeline.text_encoder = PeftModel.from_pretrained(
pipeline.text_encoder,
MODEL_ID,
subfolder="adapter_text_encoder"
)
# Generate
image = pipeline(
"Woody Savannas.",
num_inference_steps=50,
guidance_scale=3.0,
output_type="pt"
).images
```
>image is a torch.Tensor of shape (N, C, H, W) with values in the range [0, 1] and dtype float32, where N is the batch size (here 1), C the number of channels (here 1, grayscale), and H and W the native model resolution (512 x 512).
---
## Intended Use
- Synthetic SAR data generation for given land cover types.
- Research on generative models for remote sensing.
## Limitations
- Tuned specifically to **Sentinel-1** sensor characteristics based on GRD products in VH polarization; transferability to other sensors and imaging properties is very limited.
- Generated images are synthetic and may not fully capture all real SAR image properties.
---
<!-- ## Citation
If you use this model, please cite:
```bibtex
@inproceedings{sar_diffusion_gcpr2026,
title = {Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations},
booktitle = {German Conference on Pattern Recognition (GCPR)},
year = {2026},
note = {accepted, to be published},
authors = {}
}
```
--- -->
## License
This model is released under **CC BY-NC-SA 4.0**.
Commercial use is not permitted. Derivatives must be shared under the same license.
See [LICENSE](https://creativecommons.org/licenses/by-nc-sa/4.0/) for details.