Text-to-Image
Diffusers
vladmandic commited on
Commit
ad7a32e
Β·
1 Parent(s): 9652e76

Signed-off-by: Vladimir Mandic <mandic00@live.com>

Files changed (2) hide show
  1. README.md +30 -18
  2. microdecoder-minimaxh3.safetensors +3 -0
README.md CHANGED
@@ -6,20 +6,32 @@ library_name: diffusers
6
 
7
  # MicroDecoder
8
 
9
- **Micro-VAE** that can be trained on *any* diffusion model in **~15min** and provides low-quality reconstruction from latents to RGB in **~0.01sec**.
10
- Inference includes noise correction based on current timestep, intentional blurring plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels.
 
 
11
  Intended use-case is live-preview during generative model inference.
12
 
13
  Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models.
14
 
15
- - Model definition and training code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_train.py)
 
 
 
 
 
 
 
 
16
  - Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
17
 
 
 
18
  Example using **MicroDecoder** with `Flux.2-Klein-9B` and compared with official final VAE processing at the end:
19
 
20
  ![MicroDecoder](https://huggingface.co/vladmandic/MicroDecoder/resolve/main/MicroDecoder.jpg)
21
 
22
- ## Example
23
 
24
  ```shell
25
  sd_vae_micro_train.py \
@@ -35,18 +47,18 @@ sd_vae_micro_train.py \
35
  ```
36
 
37
  ```log
38
- Encoding Latents: 100%
39
- Training MicroDecoder [350 epochs | Channels: 32 | Hidden: 256 | Scale: 4x]
40
- Dataset: 360 train samples | 40 validation samples | EMA decay: 0.999
41
-
42
- Epoch | Train Tot | Tr PSNR | Tr SSIM | Val Tot | Val PSNR | Val SSIM | Val L1 | Val LAB
43
- --------------------------------------------------------------------------------------------
44
- 001/350 | 5.7244 | 10.38 | 0.3431 | 5.9729 | 10.56 | 0.3047 | 0.2676 | 0.4254 *
45
- 100/350 | 2.7690 | 25.49 | 0.8030 | 2.2011 | 26.52 | 0.8672 | 0.0376 | 0.0699 *
46
- 200/350 | 2.4006 | 27.04 | 0.8400 | 1.1679 | 34.24 | 0.9635 | 0.0132 | 0.0317 *
47
- 300/350 | 2.3950 | 28.05 | 0.8472 | 1.0002 | 35.63 | 0.9723 | 0.0113 | 0.0282 *
48
- 350/350 | 2.3010 | 28.51 | 0.8511 | 0.9781 | 35.83 | 0.9734 | 0.0111 | 0.0278 *
49
-
50
- Training complete! Best validation PSNR: 35.83 dB (Epoch 350)
51
- Saved best EMA model weights to: MicroVAE-qwen21.safetensors
52
  ```
 
6
 
7
  # MicroDecoder
8
 
9
+ **MicroDecoder** VAE can be trained on *any* diffusion model in **~15min** and provides low-quality reconstruction from latents to RGB in **~0.01sec**.
10
+
11
+ Inference includes noise correction based on current timestep, blurring corrections plus upscale interpolation: all with intention of providing as fast-as-possible reconstruction that is viable and consistent with any noise levels.
12
+
13
  Intended use-case is live-preview during generative model inference.
14
 
15
  Shapes/Channels/etc are inferred from the base VAE, so no configuration changes are needed between different models.
16
 
17
+ ## Models
18
+
19
+ Models included in this repo are pre-trained for following architectures:
20
+ - **SD, SDXL, Flux.1, Flux.2, Qwen, Qwen-21, Wan-21, MiniMax-H3**
21
+
22
+ ## Code
23
+
24
+ - Model definition [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_model.py)
25
+ - Training code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro_train.py)
26
  - Example inference code [here](https://github.com/vladmandic/sdnext/blob/dev/modules/vae/sd_vae_micro.py)
27
 
28
+ ## Example
29
+
30
  Example using **MicroDecoder** with `Flux.2-Klein-9B` and compared with official final VAE processing at the end:
31
 
32
  ![MicroDecoder](https://huggingface.co/vladmandic/MicroDecoder/resolve/main/MicroDecoder.jpg)
33
 
34
+ ## Training
35
 
36
  ```shell
37
  sd_vae_micro_train.py \
 
47
  ```
48
 
49
  ```log
50
+ MicroDecoder Train
51
+ Args: Namespace(folder='/home/vlado/generative/Input/vae/', max=500, resolution=512, crop='center', vae='AutoencoderKLMiniMaxH3', repo='OzzyGT/MiniMax_H3_sdnq_dynamic_4bit', subfolder='vae', dim=256, epochs=400, scale=4, batch=16, lr=0.0003, ema=0.999, val=0.1, output='microdecoder-minimaxh3.safetensors', device='cuda')
52
+ Base VAE: cls=<class 'diffusers.models.autoencoders.autoencoder_kl_minimax_h3.AutoencoderKLMiniMaxH3'> repo="OzzyGT/MiniMax_H3_sdnq_dynamic_4bit" subfolder="vae"
53
+ Init VAE: model=MicroDecoder ema=EMAModel optimizer=AdamW scheduler=CosineAnnealingLR criterion=EnhancedQualityLoss metrics=DetailedMetricsTracker
54
+ Epoch | tTotal | tPSNR | tSSIM | vTotal | vPSNR | vSSIM | vL1 | vLAB
55
+ -------------------------------------------------------------------------------------------
56
+ 001/400 | 7.0891 | 10.49 | 0.2540 | 7.1642 | 10.75 | 0.2357 | 0.2614 | 0.4146
57
+ ...
58
+ 400/400 | 4.5365 | 21.03 | 0.6402 | 1.9644 | 30.61 | 0.9266 | 0.0204 | 0.0465
59
+ Load ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:00:08 Images=400 Shape=torch.Size([400, 3, 512, 512])
60
+ Encode ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:00:46 Train=360 Validation=40 EMA=0.999
61
+ Train ━━━━━━━━━━━━━━━━━━━━ 400/400 100% 0:00:00 0:06:33 Train(psnr=21.029 ssim=0.640 lpips=2.784 l1=0.089 lab=0.156 grad=0.073 sat=0.124 fft=0.071 total=4.536) Validate(psnr=30.607 ssim=0.927 lpips=1.242 l1=0.020 lab=0.046 grad=0.046 sat=0.057 fft=0.032 total=1.964)
62
+ Complete: Time=393.89 Epoch/Sec=1.02 PSNR=30.61 dB @ Epoch=400
63
+ Save: filename="microdecoder-minimaxh3.safetensors"
64
  ```
microdecoder-minimaxh3.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:eb9427a1d8f92bdd123feda388f9865738095689df0a395eb49956bcd3b0415d
3
+ size 9713494