BiliSakura commited on Mar 7

Commit

2e8a26c

verified ·

1 Parent(s): 6e9e94f

Add files using upload-large-folder tool

Browse files

Files changed (27) hide show

.gitattributes +1 -0
README.md +173 -0
controlnet/GeoSynth-Canny/config.json +56 -0
controlnet/GeoSynth-Canny/diffusion_pytorch_model.safetensors +3 -0
controlnet/GeoSynth-OSM/config.json +56 -0
controlnet/GeoSynth-OSM/diffusion_pytorch_model.safetensors +3 -0
controlnet/GeoSynth-SAM/config.json +56 -0
controlnet/GeoSynth-SAM/diffusion_pytorch_model.safetensors +3 -0
demo_images/GeoSynth-Canny/input.jpeg +3 -0
demo_images/GeoSynth-Canny/output.jpeg +0 -0
demo_images/GeoSynth-OSM/input.jpeg +0 -0
demo_images/GeoSynth-OSM/output.jpeg +0 -0
demo_images/GeoSynth-SAM/input.jpeg +0 -0
demo_images/GeoSynth-SAM/output.jpeg +0 -0
feature_extractor/preprocessor_config.json +27 -0
model_index.json +38 -0
scheduler/scheduler_config.json +29 -0
text_encoder/config.json +25 -0
text_encoder/model.safetensors +3 -0
tokenizer/merges.txt +0 -0
tokenizer/special_tokens_map.json +24 -0
tokenizer/tokenizer_config.json +38 -0
tokenizer/vocab.json +0 -0
unet/config.json +72 -0
unet/diffusion_pytorch_model.safetensors +3 -0
vae/config.json +34 -0
vae/diffusion_pytorch_model.safetensors +3 -0

.gitattributes CHANGED Viewed

@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text

 *.zip filter=lfs diff=lfs merge=lfs -text
 *.zst filter=lfs diff=lfs merge=lfs -text
 *tfevents* filter=lfs diff=lfs merge=lfs -text
+demo_images/GeoSynth-Canny/input.jpeg filter=lfs diff=lfs merge=lfs -text

README.md ADDED Viewed

	@@ -0,0 +1,173 @@

+---
+license: apache-2.0
+library_name: diffusers
+pipeline_tag: image-to-image
+tags:
+- controlnet
+- remote-sensing
+- arxiv:2404.06637
+widget:
+# GeoSynth-OSM: OSM tile -> satellite image
+- src: demo_images/GeoSynth-OSM/input.jpeg
+  prompt: Satellite image features a city neighborhood
+  output:
+    url: demo_images/GeoSynth-OSM/output.jpeg
+# GeoSynth-Canny: Canny edges -> satellite image
+- src: demo_images/GeoSynth-Canny/input.jpeg
+  prompt: Satellite image features a city neighborhood
+  output:
+    url: demo_images/GeoSynth-Canny/output.jpeg
+# GeoSynth-SAM: SAM segmentation -> satellite image
+- src: demo_images/GeoSynth-SAM/input.jpeg
+  prompt: Satellite image features a city neighborhood
+  output:
+    url: demo_images/GeoSynth-SAM/output.jpeg
+---
+# GeoSynth-ControlNets
+We maintain **two repositories**—one per base checkpoint—each with its compatible ControlNets:
+| Repo | Base Model | ControlNets |
+|------|------------|-------------|
+| **This repo** | GeoSynth (text encoder & UNet same as SD 2.1) | GeoSynth-OSM, GeoSynth-Canny, GeoSynth-SAM |
+| **[GeoSynth-ControlNets-Location](https://huggingface.co/BiliSakura/GeoSynth-ControlNets-Location)** | GeoSynth-Location (adds CoordNet branch) | GeoSynth-Location-OSM, GeoSynth-Location-SAM*, GeoSynth-Location-Canny |
+*[GeoSynth-Location-SAM](https://huggingface.co/MVRL/GeoSynth-Location-SAM) controlnet ckpt is missing from source.*
+### This repository
+1. **GeoSynth checkpoint** — A remote sensing visual generative model. The text encoder and UNet are the same as [Stable Diffusion 2.1](https://huggingface.co/sd2-community/stable-diffusion-2-1-base) (not fine-tuned).
+2. **ControlNet models** — OSM, Canny, and SAM conditioning, located under [`controlnet/`](controlnet/).
+### Architecture note: location-conditioned models
+Location-conditioned variants (GeoSynth-Location-*) use a **different base checkpoint** that adds a CoordNet branch. The branch takes `[lon, lat]` as input, passes it through a **SatCLIP** location encoder, then through a **CoordNet** (13 stacked cross-attention blocks, inner dim 256, 4 heads). ControlNet and CoordNet both condition the UNet. See the [GeoSynth paper](https://huggingface.co/papers/2404.06637) Figure 3.
+### ControlNet variants (this repo)
+| Control | Subfolder | Status |
+|---------|-----------|--------|
+| OSM     | `controlnet/GeoSynth-OSM` | ✅ Integrated |
+| Canny   | `controlnet/GeoSynth-Canny` | ✅ Integrated |
+| SAM     | `controlnet/GeoSynth-SAM` | ✅ Integrated |
+Use it with 🧨 [diffusers](#examples) or the [Stable Diffusion](https://github.com/Stability-AI/stablediffusion) repository.
+### Model Sources
+- **Source:** [GeoSynth](https://github.com/mvrl/GeoSynth)
+- **Paper:** [GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis](https://huggingface.co/papers/2404.06637)
+- **Base model:** [Stable Diffusion 2.1](https://huggingface.co/sd2-community/stable-diffusion-2-1-base)
+## Examples
+### Text-to-Image (base GeoSynth)
+```python
+from diffusers import StableDiffusionPipeline
+pipe = StableDiffusionPipeline.from_pretrained("BiliSakura/GeoSynth-ControlNets")
+pipe = pipe.to("cuda")
+image = pipe("Satellite image features a city neighborhood").images[0]
+image.save("generated_city.jpg")
+```
+### ControlNet (diffusers integration)
+Use the 🧨 diffusers `ControlNetModel` wrapper with `StableDiffusionControlNetPipeline`:
+**GeoSynth-OSM** — synthesizes satellite images from OpenStreetMap tiles (RGB):
+```python
+from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
+from PIL import Image
+import torch
+controlnet = ControlNetModel.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    subfolder="controlnet/GeoSynth-OSM",
+)
+pipe = StableDiffusionControlNetPipeline.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    controlnet=controlnet,
+)
+pipe = pipe.to("cuda")
+img = Image.open("osm_tile.jpeg")  # OSM tile (RGB, 512x512)
+generator = torch.manual_seed(42)
+image = pipe("Satellite image features a city neighborhood", image=img, generator=generator, num_inference_steps=20).images[0]
+image.save("generated_city.jpg")
+```
+**GeoSynth-Canny** — synthesizes satellite images from Canny edge maps:
+```python
+from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
+from PIL import Image
+import torch
+controlnet = ControlNetModel.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    subfolder="controlnet/GeoSynth-Canny",
+)
+pipe = StableDiffusionControlNetPipeline.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    controlnet=controlnet,
+)
+pipe = pipe.to("cuda")
+img = Image.open("canny_edges.jpeg")  # Canny edge image (RGB, 512x512)
+generator = torch.manual_seed(42)
+image = pipe("Satellite image features a city neighborhood", image=img, generator=generator, num_inference_steps=20).images[0]
+image.save("generated_city.jpg")
+```
+**GeoSynth-SAM** — synthesizes satellite images from SAM (Segment Anything Model) segmentation masks:
+```python
+from diffusers import StableDiffusionControlNetPipeline, ControlNetModel
+from PIL import Image
+import torch
+controlnet = ControlNetModel.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    subfolder="controlnet/GeoSynth-SAM",
+)
+pipe = StableDiffusionControlNetPipeline.from_pretrained(
+    "BiliSakura/GeoSynth-ControlNets",
+    controlnet=controlnet,
+)
+pipe = pipe.to("cuda")
+img = Image.open("sam_segmentation.jpeg")  # SAM mask (RGB, 512x512)
+generator = torch.manual_seed(42)
+image = pipe("Satellite image features a city neighborhood", image=img, generator=generator, num_inference_steps=20).images[0]
+image.save("generated_city.jpg")
+```
+*For location-conditioned variants (GeoSynth-Location-OSM, GeoSynth-Location-SAM, GeoSynth-Location-Canny), see the separate [GeoSynth-ControlNets-Location](https://huggingface.co/BiliSakura/GeoSynth-ControlNets-Location) repo.*
+## Citation
+If you use this model, please cite the GeoSynth paper. For location-conditioned variants, also cite SatCLIP.
+```bibtex
+@inproceedings{sastry2024geosynth,
+  title={GeoSynth: Contextually-Aware High-Resolution Satellite Image Synthesis},
+  author={Sastry, Srikumar and Khanal, Subash and Dhakal, Aayush and Jacobs, Nathan},
+  booktitle={IEEE/ISPRS Workshop: Large Scale Computer Vision for Remote Sensing (EARTHVISION)},
+  year={2024}
+}
+@article{klemmer2025satclip,
+  title={{SatCLIP}: {Global}, General-Purpose Location Embeddings with Satellite Imagery},
+  author={Klemmer, Konstantin and Rolf, Esther and Robinson, Caleb and Mackey, Lester and Ru{\ss}wurm, Marc},
+  journal={Proceedings of the AAAI Conference on Artificial Intelligence},
+  volume={39},
+  number={4},
+  pages={4347--4355},
+  year={2025},
+  doi={10.1609/aaai.v39i4.32457}
+}
+```

controlnet/GeoSynth-Canny/config.json ADDED Viewed

	@@ -0,0 +1,56 @@

+{
+  "_class_name": "ControlNetModel",
+  "_diffusers_version": "0.27.0",
+  "act_fn": "silu",
+  "addition_embed_type": null,
+  "addition_embed_type_num_heads": 64,
+  "addition_time_embed_dim": null,
+  "attention_head_dim": [
+    5,
+    10,
+    20,
+    20
+  ],
+  "block_out_channels": [
+    320,
+    640,
+    1280,
+    1280
+  ],
+  "class_embed_type": null,
+  "conditioning_channels": 3,
+  "conditioning_embedding_out_channels": [
+    16,
+    32,
+    96,
+    256
+  ],
+  "controlnet_conditioning_channel_order": "rgb",
+  "cross_attention_dim": 1024,
+  "down_block_types": [
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "DownBlock2D"
+  ],
+  "downsample_padding": 1,
+  "encoder_hid_dim": null,
+  "encoder_hid_dim_type": null,
+  "flip_sin_to_cos": true,
+  "freq_shift": 0,
+  "global_pool_conditions": false,
+  "in_channels": 4,
+  "layers_per_block": 2,
+  "mid_block_scale_factor": 1,
+  "mid_block_type": "UNetMidBlock2DCrossAttn",
+  "norm_eps": 1e-05,
+  "norm_num_groups": 32,
+  "num_attention_heads": null,
+  "num_class_embeds": null,
+  "only_cross_attention": false,
+  "projection_class_embeddings_input_dim": null,
+  "resnet_time_scale_shift": "default",
+  "transformer_layers_per_block": 1,
+  "upcast_attention": false,
+  "use_linear_projection": true
+}

controlnet/GeoSynth-Canny/diffusion_pytorch_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:580077209d4c9a88d0d1c40bd9506e46d5a22b02e09d2ca192010da7846bb1e9
+size 1456953560

controlnet/GeoSynth-OSM/config.json ADDED Viewed

	@@ -0,0 +1,56 @@

+{
+  "_class_name": "ControlNetModel",
+  "_diffusers_version": "0.27.0",
+  "act_fn": "silu",
+  "addition_embed_type": null,
+  "addition_embed_type_num_heads": 64,
+  "addition_time_embed_dim": null,
+  "attention_head_dim": [
+    5,
+    10,
+    20,
+    20
+  ],
+  "block_out_channels": [
+    320,
+    640,
+    1280,
+    1280
+  ],
+  "class_embed_type": null,
+  "conditioning_channels": 3,
+  "conditioning_embedding_out_channels": [
+    16,
+    32,
+    96,
+    256
+  ],
+  "controlnet_conditioning_channel_order": "rgb",
+  "cross_attention_dim": 1024,
+  "down_block_types": [
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "DownBlock2D"
+  ],
+  "downsample_padding": 1,
+  "encoder_hid_dim": null,
+  "encoder_hid_dim_type": null,
+  "flip_sin_to_cos": true,
+  "freq_shift": 0,
+  "global_pool_conditions": false,
+  "in_channels": 4,
+  "layers_per_block": 2,
+  "mid_block_scale_factor": 1,
+  "mid_block_type": "UNetMidBlock2DCrossAttn",
+  "norm_eps": 1e-05,
+  "norm_num_groups": 32,
+  "num_attention_heads": null,
+  "num_class_embeds": null,
+  "only_cross_attention": false,
+  "projection_class_embeddings_input_dim": null,
+  "resnet_time_scale_shift": "default",
+  "transformer_layers_per_block": 1,
+  "upcast_attention": false,
+  "use_linear_projection": true
+}

controlnet/GeoSynth-OSM/diffusion_pytorch_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fcdc1311bbc19ac2e30252a298537fa2d0b3f7c8a876010329a22d5935718b16
+size 1456953560

controlnet/GeoSynth-SAM/config.json ADDED Viewed

	@@ -0,0 +1,56 @@

+{
+  "_class_name": "ControlNetModel",
+  "_diffusers_version": "0.27.0",
+  "act_fn": "silu",
+  "addition_embed_type": null,
+  "addition_embed_type_num_heads": 64,
+  "addition_time_embed_dim": null,
+  "attention_head_dim": [
+    5,
+    10,
+    20,
+    20
+  ],
+  "block_out_channels": [
+    320,
+    640,
+    1280,
+    1280
+  ],
+  "class_embed_type": null,
+  "conditioning_channels": 3,
+  "conditioning_embedding_out_channels": [
+    16,
+    32,
+    96,
+    256
+  ],
+  "controlnet_conditioning_channel_order": "rgb",
+  "cross_attention_dim": 1024,
+  "down_block_types": [
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "DownBlock2D"
+  ],
+  "downsample_padding": 1,
+  "encoder_hid_dim": null,
+  "encoder_hid_dim_type": null,
+  "flip_sin_to_cos": true,
+  "freq_shift": 0,
+  "global_pool_conditions": false,
+  "in_channels": 4,
+  "layers_per_block": 2,
+  "mid_block_scale_factor": 1,
+  "mid_block_type": "UNetMidBlock2DCrossAttn",
+  "norm_eps": 1e-05,
+  "norm_num_groups": 32,
+  "num_attention_heads": null,
+  "num_class_embeds": null,
+  "only_cross_attention": false,
+  "projection_class_embeddings_input_dim": null,
+  "resnet_time_scale_shift": "default",
+  "transformer_layers_per_block": 1,
+  "upcast_attention": false,
+  "use_linear_projection": true
+}

controlnet/GeoSynth-SAM/diffusion_pytorch_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:54ab8c768ce9fc17660595b34215f7615bda4931226069f40836972daea7e0d3
+size 1456953560

demo_images/GeoSynth-Canny/input.jpeg ADDED Viewed

Git LFS Details

SHA256: dba59aca702f3a6a6878e74a69a8fc70298e92933d676067852af168cd098294
Pointer size: 131 Bytes
Size of remote file: 102 kB

demo_images/GeoSynth-Canny/output.jpeg ADDED Viewed

demo_images/GeoSynth-OSM/input.jpeg ADDED Viewed

demo_images/GeoSynth-OSM/output.jpeg ADDED Viewed

demo_images/GeoSynth-SAM/input.jpeg ADDED Viewed

demo_images/GeoSynth-SAM/output.jpeg ADDED Viewed

feature_extractor/preprocessor_config.json ADDED Viewed

	@@ -0,0 +1,27 @@

+{
+  "crop_size": {
+    "height": 224,
+    "width": 224
+  },
+  "do_center_crop": true,
+  "do_convert_rgb": true,
+  "do_normalize": true,
+  "do_rescale": true,
+  "do_resize": true,
+  "image_mean": [
+    0.48145466,
+    0.4578275,
+    0.40821073
+  ],
+  "image_processor_type": "CLIPImageProcessor",
+  "image_std": [
+    0.26862954,
+    0.26130258,
+    0.27577711
+  ],
+  "resample": 3,
+  "rescale_factor": 0.00392156862745098,
+  "size": {
+    "shortest_edge": 224
+  }
+}

model_index.json ADDED Viewed

	@@ -0,0 +1,38 @@

+{
+  "_class_name": "StableDiffusionPipeline",
+  "_diffusers_version": "0.27.2",
+  "_name_or_path": "stabilityai/stable-diffusion-2-1-base",
+  "feature_extractor": [
+    "transformers",
+    "CLIPImageProcessor"
+  ],
+  "image_encoder": [
+    null,
+    null
+  ],
+  "requires_safety_checker": false,
+  "safety_checker": [
+    null,
+    null
+  ],
+  "scheduler": [
+    "diffusers",
+    "DPMSolverMultistepScheduler"
+  ],
+  "text_encoder": [
+    "transformers",
+    "CLIPTextModel"
+  ],
+  "tokenizer": [
+    "transformers",
+    "CLIPTokenizer"
+  ],
+  "unet": [
+    "diffusers",
+    "UNet2DConditionModel"
+  ],
+  "vae": [
+    "diffusers",
+    "AutoencoderKL"
+  ]
+}

scheduler/scheduler_config.json ADDED Viewed

	@@ -0,0 +1,29 @@

+{
+  "_class_name": "DPMSolverMultistepScheduler",
+  "_diffusers_version": "0.27.2",
+  "algorithm_type": "dpmsolver++",
+  "beta_end": 0.012,
+  "beta_schedule": "scaled_linear",
+  "beta_start": 0.00085,
+  "clip_sample": false,
+  "dynamic_thresholding_ratio": 0.995,
+  "euler_at_final": false,
+  "final_sigmas_type": "zero",
+  "lambda_min_clipped": -Infinity,
+  "lower_order_final": true,
+  "num_train_timesteps": 1000,
+  "prediction_type": "epsilon",
+  "rescale_betas_zero_snr": false,
+  "sample_max_value": 1.0,
+  "set_alpha_to_one": false,
+  "skip_prk_steps": true,
+  "solver_order": 2,
+  "solver_type": "midpoint",
+  "steps_offset": 1,
+  "thresholding": false,
+  "timestep_spacing": "linspace",
+  "trained_betas": null,
+  "use_karras_sigmas": false,
+  "use_lu_lambdas": false,
+  "variance_type": null
+}

text_encoder/config.json ADDED Viewed

	@@ -0,0 +1,25 @@

+{
+  "_name_or_path": "/home/s.sastry/.cache/huggingface/hub/models--stabilityai--stable-diffusion-2-1-base/snapshots/5ede9e4bf3e3fd1cb0ef2f7a3fff13ee514fdf06/text_encoder",
+  "architectures": [
+    "CLIPTextModel"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 0,
+  "dropout": 0.0,
+  "eos_token_id": 2,
+  "hidden_act": "gelu",
+  "hidden_size": 1024,
+  "initializer_factor": 1.0,
+  "initializer_range": 0.02,
+  "intermediate_size": 4096,
+  "layer_norm_eps": 1e-05,
+  "max_position_embeddings": 77,
+  "model_type": "clip_text_model",
+  "num_attention_heads": 16,
+  "num_hidden_layers": 23,
+  "pad_token_id": 1,
+  "projection_dim": 512,
+  "torch_dtype": "float32",
+  "transformers_version": "4.38.2",
+  "vocab_size": 49408
+}

text_encoder/model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:67e013543d4fac905c882e2993d86a2d454ee69dc9e8f37c0c23d33a48959d15
+size 1361596304

tokenizer/merges.txt ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer/special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,24 @@

+{
+  "bos_token": {
+    "content": "<|startoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "eos_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  },
+  "pad_token": "!",
+  "unk_token": {
+    "content": "<|endoftext|>",
+    "lstrip": false,
+    "normalized": true,
+    "rstrip": false,
+    "single_word": false
+  }
+}

tokenizer/tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,38 @@

+{
+  "add_prefix_space": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "!",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49406": {
+      "content": "<|startoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "49407": {
+      "content": "<|endoftext|>",
+      "lstrip": false,
+      "normalized": true,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "bos_token": "<|startoftext|>",
+  "clean_up_tokenization_spaces": true,
+  "do_lower_case": true,
+  "eos_token": "<|endoftext|>",
+  "errors": "replace",
+  "model_max_length": 77,
+  "pad_token": "!",
+  "tokenizer_class": "CLIPTokenizer",
+  "unk_token": "<|endoftext|>"
+}

tokenizer/vocab.json ADDED Viewed

The diff for this file is too large to render. See raw diff

unet/config.json ADDED Viewed

	@@ -0,0 +1,72 @@

+{
+  "_class_name": "UNet2DConditionModel",
+  "_diffusers_version": "0.27.2",
+  "act_fn": "silu",
+  "addition_embed_type": null,
+  "addition_embed_type_num_heads": 64,
+  "addition_time_embed_dim": null,
+  "attention_head_dim": [
+    5,
+    10,
+    20,
+    20
+  ],
+  "attention_type": "default",
+  "block_out_channels": [
+    320,
+    640,
+    1280,
+    1280
+  ],
+  "center_input_sample": false,
+  "class_embed_type": null,
+  "class_embeddings_concat": false,
+  "conv_in_kernel": 3,
+  "conv_out_kernel": 3,
+  "cross_attention_dim": 1024,
+  "cross_attention_norm": null,
+  "down_block_types": [
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "CrossAttnDownBlock2D",
+    "DownBlock2D"
+  ],
+  "downsample_padding": 1,
+  "dropout": 0.0,
+  "dual_cross_attention": false,
+  "encoder_hid_dim": null,
+  "encoder_hid_dim_type": null,
+  "flip_sin_to_cos": true,
+  "freq_shift": 0,
+  "in_channels": 4,
+  "layers_per_block": 2,
+  "mid_block_only_cross_attention": null,
+  "mid_block_scale_factor": 1,
+  "mid_block_type": "UNetMidBlock2DCrossAttn",
+  "norm_eps": 1e-05,
+  "norm_num_groups": 32,
+  "num_attention_heads": null,
+  "num_class_embeds": null,
+  "only_cross_attention": false,
+  "out_channels": 4,
+  "projection_class_embeddings_input_dim": null,
+  "resnet_out_scale_factor": 1.0,
+  "resnet_skip_time_act": false,
+  "resnet_time_scale_shift": "default",
+  "reverse_transformer_layers_per_block": null,
+  "sample_size": 96,
+  "time_cond_proj_dim": null,
+  "time_embedding_act_fn": null,
+  "time_embedding_dim": null,
+  "time_embedding_type": "positional",
+  "timestep_post_act": null,
+  "transformer_layers_per_block": 1,
+  "up_block_types": [
+    "UpBlock2D",
+    "CrossAttnUpBlock2D",
+    "CrossAttnUpBlock2D",
+    "CrossAttnUpBlock2D"
+  ],
+  "upcast_attention": false,
+  "use_linear_projection": true
+}

unet/diffusion_pytorch_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fa36f06444a813e04e204fb99d1b9a2ccc80f0683ad67829f1dc9c7701eb9cb7
+size 3463726504

vae/config.json ADDED Viewed

	@@ -0,0 +1,34 @@

+{
+  "_class_name": "AutoencoderKL",
+  "_diffusers_version": "0.27.2",
+  "_name_or_path": "/home/s.sastry/.cache/huggingface/hub/models--stabilityai--stable-diffusion-2-1-base/snapshots/5ede9e4bf3e3fd1cb0ef2f7a3fff13ee514fdf06/vae",
+  "act_fn": "silu",
+  "block_out_channels": [
+    128,
+    256,
+    512,
+    512
+  ],
+  "down_block_types": [
+    "DownEncoderBlock2D",
+    "DownEncoderBlock2D",
+    "DownEncoderBlock2D",
+    "DownEncoderBlock2D"
+  ],
+  "force_upcast": true,
+  "in_channels": 3,
+  "latent_channels": 4,
+  "latents_mean": null,
+  "latents_std": null,
+  "layers_per_block": 2,
+  "norm_num_groups": 32,
+  "out_channels": 3,
+  "sample_size": 768,
+  "scaling_factor": 0.18215,
+  "up_block_types": [
+    "UpDecoderBlock2D",
+    "UpDecoderBlock2D",
+    "UpDecoderBlock2D",
+    "UpDecoderBlock2D"
+  ]
+}

vae/diffusion_pytorch_model.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:2aa1f43011b553a4cba7f37456465cdbd48aab7b54b9348b890e8058ea7683ec
+size 334643268