---
license: other
license_name: see-individual-model-repos
pipeline_tag: text-to-image
tags:
- w4a8
- quantized
- int4
- comfyui
- comfy-kitchen
- convrot
- low-vram
- safetensors
---
# Rebels W4A8 Collection
## Models
| Model | Params | Size | Notes |
|---|---|---|---|
| **Flux2-Klein-4B-w4a8** | 3.9 B | **2.46 GB** | Apache-2.0. Smallest here — runs on 6 GB cards |
| **Z-Image-Turbo-w4a8** | 6.2 B | **3.67 GB** | Distilled — 8 steps, CFG 1.0 |
| **Flux2-Klein-9B-w4a8** | 9.1 B | **5.62 GB** | Flux License |
| **Krea-2-Turbo-w4a8** | 12.8 B | **7.23 GB** | Turbo — CFG 1.0, not 4.0 |
| **SCAIL-2-14B-w4a8** | ~14 B | **_.__ GB** | Wan2.1-based video, NEEDS WORK |
| **Wan-Animate-2-TURBO-w4a8** | 16.4 B | **9.59 GB** | Video. **Reference image must match the driving video's opening pose** |
| **Qwen-Image-2512-w4a8** | ~20 B | **14.5 GB** | Full model — normal steps and CFG, not the Flash recipe, NEEDS WORK |
| **MiniMax-H3-REF2VA-w4a8** | 33.1 B | **24.5 GB** | Reference-to-video with audio. Mixed: `adaln_proj` at int8 — its 2688-wide rows can't use the w4a8 kernel |
All load with the stock **Load Diffusion Model** node. Base licenses carry over — check each source model before commercial use.
All load with the stock **Load Diffusion Model** node. Base licenses carry over — check each source model before commercial use.
**Four-bit weights. Eight-bit math. Native ComfyUI kernels. No custom nodes.**
W4A8 conversions of current diffusion and video models, quantized to run on ComfyUI's native `asym_w4a8_int8` path — where int8 tensor cores do the work instead of a dequantize-then-fp16 fallback.
Roughly **0.56 bytes per parameter**, and it *runs* at that size rather than merely storing at it.
---
 |
 |
| INT8 |
w4a8 |
## Why W4A8 instead of GGUF
GGUF is excellent and I ship plenty of it. But every GGUF forward pass unpacks weights back to fp16 before the matmul — the file is small, the math is not. W4A8 keeps compute in int8 end to end.
| | GGUF Q4_K_M | W4A8 |
|---|---|---|
| Storage | ~0.60 B/elem | ~0.56 B/elem |
| Compute path | dequant → fp16 GEMM | int8 GEMM |
| Loader | ComfyUI-GGUF node | **stock Load Diffusion Model** |
| Weight error (measured) | varies by tensor | ~7% relL2 |
The format comes from Kijai's `AsymW4A8Int8Layout` work in comfy-kitchen. This collection is about applying it correctly to models nobody has converted yet, and being explicit about what was verified.
---
## What's inside a file
Each quantized Linear stores five pieces:
| Tensor | Purpose |
|---|---|
| `weight` | int4 codes, two per byte |
| `weight_s_rel` | fp8 scale, one per group of 16 |
| `weight_s_channel` | one scale per output channel |
| `weight_codebook` | 16 Lloyd-Max levels, fit to the tensor |
| `comfy_quant` | layout config the loader reads |
Three ideas stacked: a **ConvRot Hadamard rotation** that flattens outliers so four bits go further, a **codebook** of non-uniform levels fit to the actual weight distribution instead of an even grid, and **per-group fp8 scales** preserving local dynamic range. Calibration-free — no activation dataset, so nothing in the conversion biases the model toward one kind of prompt.
---
## What I do differently
**Sensitive layers are never quantized.** Timestep embeddings, conditioning projections, patch projections, final output layers and rotary tables stay high precision. On a few-step model the timestep embedder has only a handful of sigma values to distinguish — crushing it to four bits corrupts every step of the schedule. Every file is checked after conversion to confirm those layers really are stored at F16/F32, because quantizers do not preserve them automatically.
**Mixed formats where the kernel demands it.** The fused W4A8 kernel accepts a ConvRot group of exactly 256, so any layer whose input dimension isn't divisible by 256 cannot use it. Rather than silently shipping a file that errors on load, those layers are written as `int8_tensorwise` — also native, no group constraint, ~1% error. Each model card states which layers took that path.
**Every file is measured.** Conversion reports per-layer reconstruction error against the original bf16 weights. Anything that doesn't land where it should doesn't get uploaded.
---
## Requirements
- **ComfyUI 0.30.0+** with `asym_w4a8_int8` in its native quant registry
- **comfy-kitchen** installed (ships the kernels)
- An NVIDIA GPU or AMD GPU.
On startup ComfyUI prints its available formats. You want `asym_w4a8_int8` in the **Native ops** list — under *emulated* it still runs, without the int8 speed advantage.
---
## Usage
1. Drop the `.safetensors` in `ComfyUI/models/diffusion_models`
2. Load it with **Load Diffusion Model** — the stock node, no custom loader
3. Text encoder, VAE and sampler settings are unchanged from the base model
Per-model notes (step counts, CFG, resolution) live in each model's card. Distilled models have fixed schedules that must be respected — base-model settings on them produce poor results regardless of quantization.
---
## Models
*Collection in progress. Each conversion has its own repo with exact sizes, measured error, and the list of layers kept at high precision.*
---
## Licensing
These are quantized derivatives. **Every original license and usage restriction carries over unchanged**, and each model repo states the license of its base model. Check the specific model's card before commercial use — several bases in this collection are not permissive.
---
## Credits
- **Kijai** — the W4A8 int8-codebook layout and kernels
- **Comfy-Org / comfyanonymous** — comfy-kitchen and the native quantization registry
- **city96** — ComfyUI-GGUF, which taught most of us how quantized loading works in ComfyUI
- Original model authors — all base licenses apply
Quantized by [RealRebelAI](https://huggingface.co/realrebelai) · [GitHub](https://github.com/RealRebelAI) · [X](https://x.com/realrebelai)