File size: 4,440 Bytes
88bc5cc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9ff9ba1
6e1e680
88bc5cc
 
 
 
9ff9ba1
88bc5cc
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
---

license: cc-by-nc-sa-4.0
language:
  - en
library_name: diffusers
pipeline_tag: text-to-image
tags:
  - sar
  - synthetic-aperture-radar
  - remote-sensing
  - stable-diffusion
  - image-generation
  - synthetic-data
  - ship-detection
  - earth-observation
  - capella
---


# HR-SAR Stable Diffusion

## Overview

Fine-tuned **Stable Diffusion 1.5** backbone for generating synthetic high-resolution **SAR** amplitude images from text prompts.  
Published alongside the upcoming paper **"Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations"** (Hochstuhl et al., 2026; accepted for GCPR conference 2026).

---

## Database
Fine-tuning was conducted on a SAR–text dataset compiled from high-resolution **Capella (X-band)** imagery originating from the [SpaceNet6 challenge dataset](https://spacenet.ai/sn6-challenge/), which covers the Rotterdam harbor area. The dataset contains geocoded, fully polarimetric Ground Range Detected (GRD) amplitude images in dB-scale with a ground sample distance of 0.5 m, along with co-registered optical WorldView-2 imagery (RGB, 0.5 m GSD). To create SAR–text pairs, the imagery was tiled into **512 Γ— 512** patches (~256 Γ— 256 m), and text captions were generated from the corresponding optical patches using a BLIP model fine-tuned on remote sensing imagery ([BLIP_RSCID](https://huggingface.co/Gurveer05/blip-image-captioning-base-rscid-finetuned)).

---

## Architecture

| Component | Details |
|---|---|
| Base model | Stable Diffusion 1.5 |
| VAE decoder | Adapted to **single-channel** output |
| Text encoder | Fine-tuned with a **LoRA adapter** (`adapter_text_encoder/`) |
| Output | Single-channel float32 SAR amplitude image |
---

## Repository Structure

```text

HR-SAR-StableDiffusion/

β”œβ”€β”€ model_index.json

β”œβ”€β”€ unet/

β”œβ”€β”€ vae/                        # Modified: single-channel conv_out

β”œβ”€β”€ text_encoder/

β”œβ”€β”€ tokenizer/

β”œβ”€β”€ scheduler/

β”œβ”€β”€ feature_extractor/

└── adapter_text_encoder/       # LoRA adapter for text encoder

    β”œβ”€β”€ adapter_config.json

    └── adapter_model.safetensors

```

---

## Usage
This model can be used directly with the πŸ€— [diffusers](https://github.com/huggingface/diffusers) pipeline.
The LoRA adapter for the text encoder is included in this repository and must be loaded manually after initialising the pipeline.

**Requirements:**
- `torch` (CUDA recommended)
- `diffusers`
- `peft`

### Text-to-Image Generation

```python

import torch

from diffusers import StableDiffusionPipeline

from peft import PeftModel



MODEL_ID = "sylviaHoch/HR-SAR-StableDiffusion"



# Load pipeline

pipeline = StableDiffusionPipeline.from_pretrained(

    MODEL_ID,

    torch_dtype=torch.float16

).to("cuda")



# Load LoRA text-encoder adapter

pipeline.text_encoder = PeftModel.from_pretrained(

    pipeline.text_encoder,

    MODEL_ID,

    subfolder="adapter_text_encoder"

)



# Generate

image = pipeline(

    "An aerial view of oil tanks.",

    num_inference_steps=50,

    guidance_scale=3.0,

    output_type="pt"

).images



```

>image is a torch.Tensor of shape (N, C, H, W) with values in the range [0, 1] and dtype float32, where N is the batch size (here 1), C the number of channels (here 1, grayscale), and H and W the native model resolution (512 x 512).
---

## Intended Use

- Synthetic SAR data generation.
- Research on generative models for remote sensing.

## Limitations

- Tuned specifically to **Capella** sensor characteristics based on multilooked, geocoded images in VH polarization; transferability to other sensors and imaging properties is very limited.
- Generated images are synthetic and may not fully capture all real SAR image properties.

---

## Citation

If you use this model, please cite:

```bibtex

@inproceedings{sar_diffusion_gcpr2026,

  title     = {Diffusion-Based SAR Training Data Synthesis Controlled by Spatial Annotations},

  booktitle = {German Conference on Pattern Recognition (GCPR)},

  year      = {2026},

  note      = {accepted, to be published},

  authors   = {}

}

```

---

## License

This model is released under **CC BY-NC-SA 4.0**.  
Commercial use is not permitted. Derivatives must be shared under the same license.  
See [LICENSE](https://creativecommons.org/licenses/by-nc-sa/4.0/) for details.