--- license: other license_name: see-individual-model-repos pipeline_tag: text-to-image tags: - w4a8 - quantized - int4 - comfyui - comfy-kitchen - convrot - low-vram - safetensors --- # Rebels W4A8 Collection
## Models | Model | Params | Size | Notes | |---|---|---|---| | **Flux2-Klein-4B-w4a8** | 3.9 B | **2.46 GB** | Apache-2.0. Smallest here — runs on 6 GB cards | | **Z-Image-Turbo-w4a8** | 6.2 B | **3.67 GB** | Distilled — 8 steps, CFG 1.0 | | **Flux2-Klein-9B-w4a8** | 9.1 B | **5.62 GB** | Flux License | | **Krea-2-Turbo-w4a8** | 12.8 B | **7.23 GB** | Turbo — CFG 1.0, not 4.0 | | **SCAIL-2-14B-w4a8** | ~14 B | **_.__ GB** | Wan2.1-based video, NEEDS WORK | | **Wan-Animate-2-TURBO-w4a8** | 16.4 B | **9.59 GB** | Video. **Reference image must match the driving video's opening pose** | | **Qwen-Image-2512-w4a8** | ~20 B | **14.5 GB** | Full model — normal steps and CFG, not the Flash recipe, NEEDS WORK | | **MiniMax-H3-REF2VA-w4a8** | 33.1 B | **24.5 GB** | Reference-to-video with audio. Mixed: `adaln_proj` at int8 — its 2688-wide rows can't use the w4a8 kernel | All load with the stock **Load Diffusion Model** node. Base licenses carry over — check each source model before commercial use. All load with the stock **Load Diffusion Model** node. Base licenses carry over — check each source model before commercial use. **Four-bit weights. Eight-bit math. Native ComfyUI kernels. No custom nodes.** W4A8 conversions of current diffusion and video models, quantized to run on ComfyUI's native `asym_w4a8_int8` path — where int8 tensor cores do the work instead of a dequantize-then-fp16 fallback. Roughly **0.56 bytes per parameter**, and it *runs* at that size rather than merely storing at it. ---
INT8 w4a8
INT8 w4a8
## Why W4A8 instead of GGUF GGUF is excellent and I ship plenty of it. But every GGUF forward pass unpacks weights back to fp16 before the matmul — the file is small, the math is not. W4A8 keeps compute in int8 end to end. | | GGUF Q4_K_M | W4A8 | |---|---|---| | Storage | ~0.60 B/elem | ~0.56 B/elem | | Compute path | dequant → fp16 GEMM | int8 GEMM | | Loader | ComfyUI-GGUF node | **stock Load Diffusion Model** | | Weight error (measured) | varies by tensor | ~7% relL2 | The format comes from Kijai's `AsymW4A8Int8Layout` work in comfy-kitchen. This collection is about applying it correctly to models nobody has converted yet, and being explicit about what was verified. --- ## What's inside a file Each quantized Linear stores five pieces: | Tensor | Purpose | |---|---| | `weight` | int4 codes, two per byte | | `weight_s_rel` | fp8 scale, one per group of 16 | | `weight_s_channel` | one scale per output channel | | `weight_codebook` | 16 Lloyd-Max levels, fit to the tensor | | `comfy_quant` | layout config the loader reads | Three ideas stacked: a **ConvRot Hadamard rotation** that flattens outliers so four bits go further, a **codebook** of non-uniform levels fit to the actual weight distribution instead of an even grid, and **per-group fp8 scales** preserving local dynamic range. Calibration-free — no activation dataset, so nothing in the conversion biases the model toward one kind of prompt. --- ## What I do differently **Sensitive layers are never quantized.** Timestep embeddings, conditioning projections, patch projections, final output layers and rotary tables stay high precision. On a few-step model the timestep embedder has only a handful of sigma values to distinguish — crushing it to four bits corrupts every step of the schedule. Every file is checked after conversion to confirm those layers really are stored at F16/F32, because quantizers do not preserve them automatically. **Mixed formats where the kernel demands it.** The fused W4A8 kernel accepts a ConvRot group of exactly 256, so any layer whose input dimension isn't divisible by 256 cannot use it. Rather than silently shipping a file that errors on load, those layers are written as `int8_tensorwise` — also native, no group constraint, ~1% error. Each model card states which layers took that path. **Every file is measured.** Conversion reports per-layer reconstruction error against the original bf16 weights. Anything that doesn't land where it should doesn't get uploaded. --- ## Requirements - **ComfyUI 0.30.0+** with `asym_w4a8_int8` in its native quant registry - **comfy-kitchen** installed (ships the kernels) - An NVIDIA GPU or AMD GPU. On startup ComfyUI prints its available formats. You want `asym_w4a8_int8` in the **Native ops** list — under *emulated* it still runs, without the int8 speed advantage. --- ## Usage 1. Drop the `.safetensors` in `ComfyUI/models/diffusion_models` 2. Load it with **Load Diffusion Model** — the stock node, no custom loader 3. Text encoder, VAE and sampler settings are unchanged from the base model Per-model notes (step counts, CFG, resolution) live in each model's card. Distilled models have fixed schedules that must be respected — base-model settings on them produce poor results regardless of quantization. --- ## Models *Collection in progress. Each conversion has its own repo with exact sizes, measured error, and the list of layers kept at high precision.* --- ## Licensing These are quantized derivatives. **Every original license and usage restriction carries over unchanged**, and each model repo states the license of its base model. Check the specific model's card before commercial use — several bases in this collection are not permissive. --- ## Credits - **Kijai** — the W4A8 int8-codebook layout and kernels - **Comfy-Org / comfyanonymous** — comfy-kitchen and the native quantization registry - **city96** — ComfyUI-GGUF, which taught most of us how quantized loading works in ComfyUI - Original model authors — all base licenses apply Quantized by [RealRebelAI](https://huggingface.co/realrebelai) · [GitHub](https://github.com/RealRebelAI) · [X](https://x.com/realrebelai)