Instructions to use vladmandic/MicroDecoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use vladmandic/MicroDecoder with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("vladmandic/MicroDecoder", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Update README.md
Browse files
README.md
CHANGED
|
@@ -6,15 +6,33 @@ library_name: diffusers
|
|
| 6 |
|
| 7 |
# MicroDecoder
|
| 8 |
|
| 9 |
-
**Micro-VAE** that can be trained on *any* diffusion model in ~15min and provides low-quality reconstruction from latents in ~0.01sec.
|
| 10 |
-
Inference includes noise correction based on current timestep,
|
| 11 |
Intended use-case is live-preview during generative model inference.
|
| 12 |
|
| 13 |
-
-
|
| 14 |
- Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
|
| 15 |
|
| 16 |
-
|
|
|
|
| 17 |
```shell
|
| 18 |
sd_vae_micro_train.py --dim 256 --epochs 350 --resolution 512 --scale 8 --lr 0.0003 --folder ~/generative/Input/vae/ --vae AutoencoderKLQwenImage21 --repo Qwen/Qwen-Image-2.1 --output AutoencoderKLQwenImage21.safetensors
|
| 19 |
sd_vae_micro_train.py --dim 256 --epochs 350 --resolution 512 --scale 4 --lr 0.0003 --folder ~/generative/Input/vae/ --vae AutoencoderKLFlux2 --repo black-forest-labs/FLUX.2-klein-9B --output AutoencoderKLFlux2.safetensors
|
| 20 |
```
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 6 |
|
| 7 |
# MicroDecoder
|
| 8 |
|
| 9 |
+
**Micro-VAE** that can be trained on *any* diffusion model in **~15min** and provides low-quality reconstruction from latents to RGB in **~0.01sec**.
|
| 10 |
+
Inference includes noise correction based on current timestep, intentional blurring plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels.
|
| 11 |
Intended use-case is live-preview during generative model inference.
|
| 12 |
|
| 13 |
+
- Model definition and training code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_train.py)
|
| 14 |
- Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
|
| 15 |
|
| 16 |
+
## Examples
|
| 17 |
+
|
| 18 |
```shell
|
| 19 |
sd_vae_micro_train.py --dim 256 --epochs 350 --resolution 512 --scale 8 --lr 0.0003 --folder ~/generative/Input/vae/ --vae AutoencoderKLQwenImage21 --repo Qwen/Qwen-Image-2.1 --output AutoencoderKLQwenImage21.safetensors
|
| 20 |
sd_vae_micro_train.py --dim 256 --epochs 350 --resolution 512 --scale 4 --lr 0.0003 --folder ~/generative/Input/vae/ --vae AutoencoderKLFlux2 --repo black-forest-labs/FLUX.2-klein-9B --output AutoencoderKLFlux2.safetensors
|
| 21 |
```
|
| 22 |
+
|
| 23 |
+
```log
|
| 24 |
+
Encoding Latents: 100%
|
| 25 |
+
Training MicroDecoder [350 epochs | Channels: 32 | Hidden: 256 | Scale: 4x]
|
| 26 |
+
Dataset: 360 train samples | 40 validation samples | EMA decay: 0.999
|
| 27 |
+
|
| 28 |
+
Epoch | Train Tot | Tr PSNR | Tr SSIM | Val Tot | Val PSNR | Val SSIM | Val L1 | Val LAB
|
| 29 |
+
--------------------------------------------------------------------------------------------
|
| 30 |
+
001/350 | 5.7244 | 10.38 | 0.3431 | 5.9729 | 10.56 | 0.3047 | 0.2676 | 0.4254 *
|
| 31 |
+
100/350 | 2.7690 | 25.49 | 0.8030 | 2.2011 | 26.52 | 0.8672 | 0.0376 | 0.0699 *
|
| 32 |
+
200/350 | 2.4006 | 27.04 | 0.8400 | 1.1679 | 34.24 | 0.9635 | 0.0132 | 0.0317 *
|
| 33 |
+
300/350 | 2.3950 | 28.05 | 0.8472 | 1.0002 | 35.63 | 0.9723 | 0.0113 | 0.0282 *
|
| 34 |
+
350/350 | 2.3010 | 28.51 | 0.8511 | 0.9781 | 35.83 | 0.9734 | 0.0111 | 0.0278 *
|
| 35 |
+
|
| 36 |
+
Training complete! Best validation PSNR: 35.83 dB (Epoch 350)
|
| 37 |
+
Saved best EMA model weights to: AutoencoderKLFlux2.safetensors
|
| 38 |
+
```
|