Instructions to use vladmandic/MicroDecoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use vladmandic/MicroDecoder with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("vladmandic/MicroDecoder", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
Commit Β·
ad7a32e
1
Parent(s): 9652e76
update
Browse filesSigned-off-by: Vladimir Mandic <mandic00@live.com>
- README.md +30 -18
- microdecoder-minimaxh3.safetensors +3 -0
README.md
CHANGED
|
@@ -6,20 +6,32 @@ library_name: diffusers
|
|
| 6 |
|
| 7 |
# MicroDecoder
|
| 8 |
|
| 9 |
-
**
|
| 10 |
-
|
|
|
|
|
|
|
| 11 |
Intended use-case is live-preview during generative model inference.
|
| 12 |
|
| 13 |
Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models.
|
| 14 |
|
| 15 |
-
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 16 |
- Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
|
| 17 |
|
|
|
|
|
|
|
| 18 |
Example using **MicroDecoder** with `Flux.2-Klein-9B` and compared with official final VAE processing at the end:
|
| 19 |
|
| 20 |

|
| 21 |
|
| 22 |
-
##
|
| 23 |
|
| 24 |
```shell
|
| 25 |
sd_vae_micro_train.py \
|
|
@@ -35,18 +47,18 @@ sd_vae_micro_train.py \
|
|
| 35 |
```
|
| 36 |
|
| 37 |
```log
|
| 38 |
-
|
| 39 |
-
|
| 40 |
-
|
| 41 |
-
|
| 42 |
-
Epoch
|
| 43 |
-
-------------------------------------------------------------------------------------------
|
| 44 |
-
001/
|
| 45 |
-
|
| 46 |
-
|
| 47 |
-
|
| 48 |
-
|
| 49 |
-
|
| 50 |
-
|
| 51 |
-
|
| 52 |
```
|
|
|
|
| 6 |
|
| 7 |
# MicroDecoder
|
| 8 |
|
| 9 |
+
**MicroDecoder** VAE can be trained on *any* diffusion model in **~15min** and provides low-quality reconstruction from latents to RGB in **~0.01sec**.
|
| 10 |
+
|
| 11 |
+
Inference includes noise correction based on current timestep, blurring corrections plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels.
|
| 12 |
+
|
| 13 |
Intended use-case is live-preview during generative model inference.
|
| 14 |
|
| 15 |
Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models.
|
| 16 |
|
| 17 |
+
## Models
|
| 18 |
+
|
| 19 |
+
Models included in this repo are pre-trained for following architectures:
|
| 20 |
+
- **SD, SDXL, Flux.1, Flux.2, Qwen, Qwen-21, Wan-21, MiniMax-H3**
|
| 21 |
+
|
| 22 |
+
## Code
|
| 23 |
+
|
| 24 |
+
- Model definition [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_model.py)
|
| 25 |
+
- Training code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_train.py)
|
| 26 |
- Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
|
| 27 |
|
| 28 |
+
## Example
|
| 29 |
+
|
| 30 |
Example using **MicroDecoder** with `Flux.2-Klein-9B` and compared with official final VAE processing at the end:
|
| 31 |
|
| 32 |

|
| 33 |
|
| 34 |
+
## Training
|
| 35 |
|
| 36 |
```shell
|
| 37 |
sd_vae_micro_train.py \
|
|
|
|
| 47 |
```
|
| 48 |
|
| 49 |
```log
|
| 50 |
+
MicroDecoder Train
|
| 51 |
+
Args: Namespace(folder='/home/vlado/generative/Input/vae/', max=500, resolution=512, crop='center', vae='AutoencoderKLMiniMaxH3', repo='OzzyGT/MiniMax_H3_sdnq_dynamic_4bit', subfolder='vae', dim=256, epochs=400, scale=4, batch=16, lr=0.0003, ema=0.999, val=0.1, output='microdecoder-minimaxh3.safetensors', device='cuda')
|
| 52 |
+
Base VAE: cls=<class 'diffusers.models.autoencoders.autoencoder_kl_minimax_h3.AutoencoderKLMiniMaxH3'> repo="OzzyGT/MiniMax_H3_sdnq_dynamic_4bit" subfolder="vae"
|
| 53 |
+
Init VAE: model=MicroDecoder ema=EMAModel optimizer=AdamW scheduler=CosineAnnealingLR criterion=EnhancedQualityLoss metrics=DetailedMetricsTracker
|
| 54 |
+
Epoch | tTotal | tPSNR | tSSIM | vTotal | vPSNR | vSSIM | vL1 | vLAB
|
| 55 |
+
-------------------------------------------------------------------------------------------
|
| 56 |
+
001/400 | 7.0891 | 10.49 | 0.2540 | 7.1642 | 10.75 | 0.2357 | 0.2614 | 0.4146
|
| 57 |
+
...
|
| 58 |
+
400/400 | 4.5365 | 21.03 | 0.6402 | 1.9644 | 30.61 | 0.9266 | 0.0204 | 0.0465
|
| 59 |
+
Load ββββββββββββββββββββ 400/400 100% 0:00:00 0:00:08 Images=400 Shape=torch.Size([400, 3, 512, 512])
|
| 60 |
+
Encode ββββββββββββββββββββ 400/400 100% 0:00:00 0:00:46 Train=360 Validation=40 EMA=0.999
|
| 61 |
+
Train ββββββββββββββββββββ 400/400 100% 0:00:00 0:06:33 Train(psnr=21.029 ssim=0.640 lpips=2.784 l1=0.089 lab=0.156 grad=0.073 sat=0.124 fft=0.071 total=4.536) Validate(psnr=30.607 ssim=0.927 lpips=1.242 l1=0.020 lab=0.046 grad=0.046 sat=0.057 fft=0.032 total=1.964)
|
| 62 |
+
Complete: Time=393.89 Epoch/Sec=1.02 PSNR=30.61 dB @ Epoch=400
|
| 63 |
+
Save: filename="microdecoder-minimaxh3.safetensors"
|
| 64 |
```
|
microdecoder-minimaxh3.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:eb9427a1d8f92bdd123feda388f9865738095689df0a395eb49956bcd3b0415d
|
| 3 |
+
size 9713494
|