Text-to-Image
Diffusers

MicroDecoder

MicroDecoder VAE can be trained on any diffusion model in ~5-15min and provides low-quality reconstruction from latents to RGB in ~0.01sec.

Inference includes noise correction based on current timestep, blurring corrections plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels.

Intended use-case is live-preview during generative model inference.

Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models.

Models

Models included in this repo are pre-trained for following architectures:

  • SD, SDXL, Flux.1, Flux.2, Qwen, Qwen-21, Wan-21, MiniMax-H3

Code

  • Model definition here
  • Training code here
  • Example inference code here

Example

Example using MicroDecoder with Flux.2-Klein-9B and compared with official final VAE processing at the end:

MicroDecoder

Training

sd_vae_micro_train.py \
  --dim 256 \
  --epochs 350 \
  --resolution 512 \
  --scale 4 \
  --lr 0.0003
  --folder ~/generative/Input/vae/ \
  --vae AutoencoderKLQwenImage21 \
  --repo Qwen/Qwen-Image-2.1
  --output MicroVAE-qwen21.safetensors
MicroDecoder Train
Args: Namespace(folder='/home/vlado/generative/Input/vae/', max=500, resolution=512, crop='center', vae='AutoencoderKLMiniMaxH3', repo='OzzyGT/MiniMax_H3_sdnq_dynamic_4bit', subfolder='vae', dim=256, epochs=400, scale=4, batch=16, lr=0.0003, ema=0.999, val=0.1, output='microdecoder-minimaxh3.safetensors', device='cuda')
Base VAE: cls=<class 'diffusers.models.autoencoders.autoencoder_kl_minimax_h3.AutoencoderKLMiniMaxH3'> repo="OzzyGT/MiniMax_H3_sdnq_dynamic_4bit" subfolder="vae"
Init VAE: model=MicroDecoder ema=EMAModel optimizer=AdamW scheduler=CosineAnnealingLR criterion=EnhancedQualityLoss metrics=DetailedMetricsTracker
Epoch   | tTotal    | tPSNR   | tSSIM   | vTotal   | vPSNR    | vSSIM    | vL1     | vLAB
-------------------------------------------------------------------------------------------
001/400 | 7.0891    | 10.49   | 0.2540  | 7.1642   | 10.75    | 0.2357   | 0.2614  | 0.4146
...
400/400 | 4.5365    | 21.03   | 0.6402  | 1.9644   | 30.61    | 0.9266   | 0.0204  | 0.0465
Load   ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:00:08 Images=400 Shape=torch.Size([400, 3, 512, 512])
Encode ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:00:46 Train=360 Validation=40 EMA=0.999
Train  ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:06:33 Train(psnr=21.029 ssim=0.640 lpips=2.784 l1=0.089 lab=0.156 grad=0.073 sat=0.124 fft=0.071 total=4.536) Validate(psnr=30.607 ssim=0.927 lpips=1.242 l1=0.020 lab=0.046 grad=0.046 sat=0.057 fft=0.032 total=1.964)
Complete: Time=393.89 Epoch/Sec=1.02 PSNR=30.61 dB @ Epoch=400
Save: filename="microdecoder-minimaxh3.safetensors"
Downloads last month
179
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support