Instructions to use vladmandic/MicroDecoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use vladmandic/MicroDecoder with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("vladmandic/MicroDecoder", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- Draw Things
- DiffusionBee
|
Download README.md from vladmandic/MicroDecoder: direct link, hf CLI and curl.
- Browser
- Download file 3.38 kB
-
https://huggingface.co/vladmandic/MicroDecoder/resolve/main/README.md
- Command line
-
hf download hf://vladmandic/MicroDecoder/README.md
-
curl -L -o README.md https://huggingface.co/vladmandic/MicroDecoder/resolve/main/README.md
3.38 kB
| license: apache-2.0 | |
| pipeline_tag: text-to-image | |
| library_name: diffusers | |
| # MicroDecoder | |
| **MicroDecoder** VAE can be trained on *any* diffusion model in **~5-15min** and provides low-quality reconstruction from latents to RGB in **~0.01sec**. | |
| Inference includes noise correction based on current timestep, blurring corrections plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels. | |
| Intended use-case is live-preview during generative model inference. | |
| Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models. | |
| ## Models | |
| Models included in this repo are pre-trained for following architectures: | |
| - **SD, SDXL, Flux.1, Flux.2, Qwen, Qwen-21, Wan-21, MiniMax-H3** | |
| ## Code | |
| - Model definition [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_model.py) | |
| - Training code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_train.py) | |
| - Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py) | |
| ## Example | |
| Example using **MicroDecoder** with `Flux.2-Klein-9B` and compared with official final VAE processing at the end: | |
|  | |
| ## Training | |
| ```shell | |
| sd_vae_micro_train.py \ | |
| --dim 256 \ | |
| --epochs 350 \ | |
| --resolution 512 \ | |
| --scale 4 \ | |
| --lr 0.0003 | |
| --folder ~/generative/Input/vae/ \ | |
| --vae AutoencoderKLQwenImage21 \ | |
| --repo Qwen/Qwen-Image-2.1 | |
| --output MicroVAE-qwen21.safetensors | |
| ``` | |
| ```log | |
| MicroDecoder Train | |
| Args: Namespace(folder='/home/vlado/generative/Input/vae/', max=500, resolution=512, crop='center', vae='AutoencoderKLMiniMaxH3', repo='OzzyGT/MiniMax_H3_sdnq_dynamic_4bit', subfolder='vae', dim=256, epochs=400, scale=4, batch=16, lr=0.0003, ema=0.999, val=0.1, output='microdecoder-minimaxh3.safetensors', device='cuda') | |
| Base VAE: cls=<class 'diffusers.models.autoencoders.autoencoder_kl_minimax_h3.AutoencoderKLMiniMaxH3'> repo="OzzyGT/MiniMax_H3_sdnq_dynamic_4bit" subfolder="vae" | |
| Init VAE: model=MicroDecoder ema=EMAModel optimizer=AdamW scheduler=CosineAnnealingLR criterion=EnhancedQualityLoss metrics=DetailedMetricsTracker | |
| Epoch | tTotal | tPSNR | tSSIM | vTotal | vPSNR | vSSIM | vL1 | vLAB | |
| ------------------------------------------------------------------------------------------- | |
| 001/400 | 7.0891 | 10.49 | 0.2540 | 7.1642 | 10.75 | 0.2357 | 0.2614 | 0.4146 | |
| ... | |
| 400/400 | 4.5365 | 21.03 | 0.6402 | 1.9644 | 30.61 | 0.9266 | 0.0204 | 0.0465 | |
| Load ββββββββββββββββββββ 400/400 100% 0:00:00 0:00:08 Images=400 Shape=torch.Size([400, 3, 512, 512]) | |
| Encode ββββββββββββββββββββ 400/400 100% 0:00:00 0:00:46 Train=360 Validation=40 EMA=0.999 | |
| Train ββββββββββββββββββββ 400/400 100% 0:00:00 0:06:33 Train(psnr=21.029 ssim=0.640 lpips=2.784 l1=0.089 lab=0.156 grad=0.073 sat=0.124 fft=0.071 total=4.536) Validate(psnr=30.607 ssim=0.927 lpips=1.242 l1=0.020 lab=0.046 grad=0.046 sat=0.057 fft=0.032 total=1.964) | |
| Complete: Time=393.89 Epoch/Sec=1.02 PSNR=30.61 dB @ Epoch=400 | |
| Save: filename="microdecoder-minimaxh3.safetensors" | |
| ``` | |