Text-to-Image
Diffusers
Safetensors
English
satellite-imagery
remote-sensing
diffusion-transformer
controllable-generation
geospatial
Instructions to use BiliSakura/TerraDiT-diffusers with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use BiliSakura/TerraDiT-diffusers with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("BiliSakura/TerraDiT-diffusers", dtype=torch.bfloat16, device_map="cuda") prompt = "The satellite image shows a coastal town where a marina with many small boats meets the dark blue sea, with a sandy beach along the shore and low-rise buildings with red roofs behind it." image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
File size: 7,472 Bytes
19b0ddb | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | ---
license: cc-by-nc-4.0
library_name: diffusers
pipeline_tag: text-to-image
tags:
- satellite-imagery
- remote-sensing
- diffusion-transformer
- controllable-generation
- geospatial
datasets:
- lcybuaa/Git-10M
- MVRL/TerraDiT-data
inference: true
language:
- en
widget:
- text: The satellite image shows a coastal town where a marina with many small boats meets the dark blue sea, with a sandy beach along the shore and low-rise buildings with red roofs behind it.
output:
url: TerraDiT-Alpha-XL/demo/demo.png
- text: The satellite image shows a suburban neighborhood with detached houses along curving residential streets, a school with a large sports field, and a small pond surrounded by trees in the lower right.
output:
url: TerraDiT-Sigma-XL/demo/demo.png
- text: The satellite image shows a small town where a river with dark water runs from the top left to the bottom right, crossed by a road bridge. A park with green lawns and trees lies north of the river, a large warehouse with a flat grey roof sits in the east, and residential houses with red roofs fill the south-west.
output:
url: TerraDiT-Omega-XL/demo/demo.png
- text: The satellite image shows a small town where a river with dark water runs from the top left to the bottom right, crossed by a road bridge. A park with green lawns and trees lies north of the river, a large warehouse with a flat grey roof sits in the east, and residential houses with red roofs fill the south-west.
output:
url: TerraDiT-Omega-B/demo/demo.png
---
# TerraDiT-diffusers
Self-contained [Diffusers](https://github.com/huggingface/diffusers) checkpoints for **TerraDiT-α**, **TerraDiT-Σ**, and **TerraDiT-Ω** (ECCV 2026): satellite image synthesis at 256×256 from text, geolocation, point prompts, and arbitrary geospatial primitives.
Converted from [`MVRL/TerraDiT`](https://huggingface.co/MVRL/TerraDiT) with [`scripts/convert_to_diffusers.py`](https://github.com/mvrl/TerraDiT). Each subfolder is a SiT-diffusers-style repo (`pipeline.py` + family transformer + optional `geolocation_encoder/`) and loads without installing the training code.
[Code](https://github.com/mvrl/TerraDiT) ·
[Collection](https://huggingface.co/collections/MVRL/terradit-6a9a248fef20075fe3b1967c) ·
[Data](https://huggingface.co/datasets/MVRL/TerraDiT-data) ·
[TerraDiT-Ω paper](https://arxiv.org/abs/2606.31029) ·
[TerraDiT paper](https://arxiv.org/abs/2603.02172)
| Subfolder | Model | Conditioning | Backbone | Params |
| --- | --- | --- | --- | ---: |
| `TerraDiT-Alpha-XL` | TerraDiT-α | text | SiT-XL/2 | 828M |
| `TerraDiT-Sigma-XL` | TerraDiT-Σ | text + geolocation + point prompts | SiT-XL/2 | 1.10B |
| `TerraDiT-Omega-XL` | TerraDiT-Ω | text + geolocation + polygons / polylines / boxes / points | SiT-XL/2 | 1.18B |
| `TerraDiT-Omega-B` | TerraDiT-Ω base | same as Ω | SiT-B/2 | 299M |
Shared components in every variant: SDXL VAE (`scaling_factor=0.13025`), LongCLIP text encoder (`zer0int/LongCLIP-KO-LITE-TypoAttack-Attn-ViT-L-14`, 144 tokens), `FlowMatchEulerDiscreteScheduler` (100 Euler steps, no CFG). Σ / Ω also ship RANGE+ under `geolocation_encoder/` (SatCLIP location encoder + `range_db.npz`).
## Demo
256×256, 100 Euler steps, `guidance_scale=0.0`, seed `42`. α is text-only; Σ adds St. Louis lat/lon plus point prompts; Ω / Ω-B add Paris lat/lon plus mixed geospatial primitives.
| α XL | Σ XL | Ω XL | Ω B |
| --- | --- | --- | --- |
|  |  |  |  |
## Load a local subfolder
```python
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"./TerraDiT-Alpha-XL",
trust_remote_code=True,
).to("cuda")
image = pipe(
prompt="The satellite image shows a coastal town with a marina and red-roofed houses.",
height=256,
width=256,
num_inference_steps=100,
guidance_scale=0.0,
generator=torch.Generator(device="cuda").manual_seed(42),
).images[0]
image.save("terradit_alpha.png")
```
### TerraDiT-Σ — text + lat/lon + point prompts
```python
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"./TerraDiT-Sigma-XL",
trust_remote_code=True,
).to("cuda")
image = pipe(
prompt="The satellite image shows a suburban neighborhood with a school and a pond.",
lat=38.65,
lon=-90.31,
points=[
[40, 40, "building house"],
[90, 170, "amenity school"],
[205, 215, "natural water"],
],
height=256,
width=256,
num_inference_steps=100,
generator=torch.Generator(device="cuda").manual_seed(42),
).images[0]
```
Coordinates are tile pixels in `[0, 256)`, x right and y down. Tags follow the OSM `"<key> <value>"` convention.
### TerraDiT-Ω — any geospatial primitive
```python
import torch
from diffusers import DiffusionPipeline
pipe = DiffusionPipeline.from_pretrained(
"./TerraDiT-Omega-XL", # or ./TerraDiT-Omega-B
trust_remote_code=True,
).to("cuda")
image = pipe(
prompt="A small town crossed by a river and a road bridge.",
lat=48.86,
lon=2.35,
instances=[
{"type": "polygon", "coords": [[20, 20], [110, 15], [120, 90], [30, 100]], "tag": "leisure park"},
{"type": "polyline", "coords": [[0, 30], [60, 110], [140, 170], [255, 235]], "tag": "waterway river"},
{"type": "bbox", "coords": [[20, 170], [70, 215]], "tag": "building house"},
{"type": "point", "coords": [200, 200], "tag": "amenity parking"},
],
condition_type="omega",
height=256,
width=256,
num_inference_steps=100,
generator=torch.Generator(device="cuda").manual_seed(42),
).images[0]
```
## Sources
| Component | Source |
| --- | --- |
| Transformers | [`MVRL/TerraDiT`](https://huggingface.co/MVRL/TerraDiT) (`alpha_xl`, `sigma_xl`, `omega_xl`, `omega_base`) |
| VAE | [`stabilityai/sdxl-vae`](https://huggingface.co/stabilityai/sdxl-vae) |
| Text encoder / tokenizer | [`zer0int/LongCLIP-KO-LITE-TypoAttack-Attn-ViT-L-14`](https://huggingface.co/zer0int/LongCLIP-KO-LITE-TypoAttack-Attn-ViT-L-14) |
| Location encoder | [`MVRL/satclip-loc-enc-vit16-l40`](https://huggingface.co/MVRL/satclip-loc-enc-vit16-l40) |
| RANGE+ database | [`mvrl/RANGE-database`](https://huggingface.co/datasets/mvrl/RANGE-database) (`range_db_large.npz`) |
## License
Weights: **CC-BY-NC-4.0** (non-commercial research), because the training imagery
([Git-10M](https://huggingface.co/datasets/lcybuaa/Git-10M)) is CC-BY-NC-ND-4.0.
Code: Apache 2.0. Conditioning data derives from OpenStreetMap (ODbL, © OpenStreetMap contributors).
SDXL VAE and LongCLIP are MIT.
```bibtex
@inproceedings{wei2026terraditomega,
title = {TerraDiT-{\Omega}: Unified Spatial Control for Satellite Image Synthesis with Any Geospatial Primitive},
author = {Wei, Brian and Sastry, Srikumar and Cher, Daniel and Xing, Eric and Jacobs, Nathan},
booktitle = {European Conference on Computer Vision},
year = {2026}
}
@article{sastry2026terradit,
title = {TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis},
author = {Sastry, Srikumar and Cher, Daniel and Wei, Brian and Dhakal, Aayush and Khanal, Subash and Gupta, Dev and Jacobs, Nathan},
journal = {arXiv:2603.02172},
year = {2026}
}
```
|