MiniMax-H3 10Eros-Max beta2 (FL2VA) β W4A8
W4A8 (asym_w4a8_int8) quantization of 10Eros_Max_h3_fl2va_beta2_pruned (bf16 source from TenStrip/10Eros-Max).
- Format:
asym_w4a8_int8,group_size=16,convrot_groupsize=256, per-tensor Lloyd-Max codebook + fp8 group scales (Kijai / comfy-kitchen W4A8 layout). - Size: ~11.7 GB (68.8% smaller than the 40 GB bf16 source) β fits a 24 GB card with headroom.
- Layers: 200 2D-linear layers quantized (96% of policy-targeted bytes); norms / first / last kept higher-precision.
- Quantizer: comfyui-mixed-quantizer
--format w4a8 --group-size 16 --codebook-mode fit. - Reconstruction: relL2 β 0.073 (bound 0.25), SNR β 22.8 dB, cos β 0.9973.
Requirements
- ComfyUI β₯ v0.31.0 (native W4A8 loader) or the
comfyui_w4a8_loader.patch. - comfy-kitchen with
AsymW4A8Int8Layout(PR #90). - CUDA SM β₯ 8.0 (verified on an RTX 3090 / SM 8.6).
Verification
Loads via ComfyUI's diffusion-model loader and generates end-to-end on an RTX 3090 (image still executed in ~28 s).
Community experiment; inherits the source model's license.
- Downloads last month
- -