File size: 8,862 Bytes
3c3ef46 bcd319e 3c3ef46 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 | ---
license: other
license_name: flux-1-dev-non-commercial-license
license_link: https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/3de623fc3c33e44ffbe2bad470d0f45bccf2eb21/LICENSE.md
base_model: black-forest-labs/FLUX.1-dev
base_model_relation: quantized
pipeline_tag: text-to-image
library_name: comfyui
tags:
- flux
- flux.1-dev
- comfyui
- quantization
- convrot
- int8
- int4
- text-to-image
inference: false
---
# FLUX.1-dev ConvRot for ComfyUI
Native ComfyUI ConvRot quantizations of **FLUX.1-dev**, with three diffusion-model choices and an optional INT8 ConvRot T5-XXL text encoder.
These are post-training quantized derivatives of [`black-forest-labs/FLUX.1-dev`](https://huggingface.co/black-forest-labs/FLUX.1-dev) at source revision `3de623fc3c33e44ffbe2bad470d0f45bccf2eb21`. They load with ComfyUI's core `UNETLoader`; no custom INT4 loader is required.
> **License:** The model weights are subject to the upstream FLUX.1-dev Non-Commercial License. Access and use require compliance with that license and the upstream acceptable-use terms. See [License](#license) before downloading.

## Files and recommended combinations
| File | Role | Size | Recommendation |
|---|---|---:|---|
| `FLUX.1-dev-w8a8-convrot.safetensors` | Whole-DiT ConvRot W8A8 | 11.981 GB | **Default DiT**: best storage/fidelity balance |
| `FLUX.1-dev-int8_convrot.safetensors` | Partial ConvRot INT8 with sensitive tensors preserved | 15.920 GB | **Quality-first DiT**: closest tested quantized variant to BF16 |
| `FLUX.1-dev-paper-w4a4-convrot.safetensors` | Paper-aligned whole-DiT ConvRot W4A4 | 6.248 GB | **Memory-first / experimental** |
| `t5xxl_flux1_int8_convrot.safetensors` | Experimental T5-XXL ConvRot W8A8 | 4.898 GB | Optional text-encoder companion |
The four uploadable files total 39.047 GB. The Original BF16 model shown in the comparison is a reference and is **not** included.
Recommended pipelines:
1. Lowest combined storage: `FLUX.1-dev-w8a8-convrot.safetensors` + `t5xxl_flux1_int8_convrot.safetensors`.
2. Default quality/storage balance: `FLUX.1-dev-w8a8-convrot.safetensors` + upstream BF16 T5-XXL.
3. Quality-first: `FLUX.1-dev-int8_convrot.safetensors` + upstream BF16 T5-XXL.
4. Memory-first: `FLUX.1-dev-paper-w4a4-convrot.safetensors` + upstream BF16 T5-XXL.
The optional INT8 T5 does not contain CLIP-L. FLUX still requires CLIP-L and a VAE.
## What W8A8 and W4A4 mean
ConvRot stores rotated quantized weights plus per-layer metadata. At runtime, the quantized weight layout dispatches the matching ComfyUI/comfy-kitchen linear operator. The current BF16 activation is rotated and dynamically quantized for the matrix multiplication; activations are not stored in the checkpoint.
- **W8A8:** INT8 weights and dynamically quantized INT8 activations for selected linear layers.
- **W4A4:** packed INT4 weights and dynamically quantized INT4 activations for selected linear layers.
- Biases, norms, embeddings, and tensors outside each policy remain at their declared higher precision.
Rotation group size is 256. The W4A4 artifact uses quantization group size 64.
## ComfyUI installation
Place the selected files as follows:
```text
ComfyUI/models/diffusion_models/FLUX.1-dev-w8a8-convrot.safetensors
# or FLUX.1-dev-int8_convrot.safetensors
# or FLUX.1-dev-paper-w4a4-convrot.safetensors
ComfyUI/models/text_encoders/t5xxl_flux1_int8_convrot.safetensors # optional
```
You must separately obtain the standard FLUX assets under their applicable licenses:
```text
ComfyUI/models/text_encoders/clip_l_flux1.safetensors
ComfyUI/models/text_encoders/t5xxl_flux1_bf16.safetensors # unless using the INT8 T5 here
ComfyUI/models/vae/ae_flux1.safetensors
```
Use a current ComfyUI and comfy-kitchen build with native ConvRot support. The validated environment was:
- ComfyUI `0.28.0`
- comfy-kitchen `0.2.20`
- comfy-cli `1.11.1`
- NVIDIA CUDA GPU
W4A4 requires SM 8.0 or newer. The first run may compile kernels and take longer.
Load the diffusion model with core `UNETLoader` and `weight_dtype=default`. Load CLIP-L and the chosen T5 with `DualCLIPLoader(type=flux)`.
A ready-to-use API/comfy-cli workflow is provided at
[`workflows/FLUX1_Dev_ConvRot_API.json`](workflows/FLUX1_Dev_ConvRot_API.json).
It defaults to Whole W8A8 + INT8 T5. Change the diffusion-model and/or T5
filename to test another combination.
A maintained copy is available in the general
[`comfy-workflows`](https://github.com/kgonia/comfy-workflows/blob/master/FLUX1_Dev_ConvRot_API.json)
repository for direct API use.
## Paired quality evaluation
Twenty 1024×1024 images were generated per pipeline with identical prompts, seeds, sampler, scheduler, guidance, steps, CLIP-L, and VAE. Metrics compare final 8-bit RGB images to Original BF16.
| Pipeline | Global PSNR vs BF16 | Mean per-pair PSNR | Low-VRAM maximum peak |
|---|---:|---:|---:|
| Whole W8A8 + INT8 T5 | 26.559 dB | 28.872 dB | 16,265 MiB |
| Whole W8A8 + BF16 T5 | 27.399 dB | 29.449 dB | 16,298 MiB |
| Partial INT8 + BF16 T5 | **27.888 dB** | **29.857 dB** | 20,351 MiB |
| Paper W4A4 + BF16 T5 | 17.926 dB | 19.041 dB | **10,797 MiB** |
The isolated T5 comparison—same Whole W8A8 DiT, BF16 T5 versus INT8 T5—measured 27.389 dB global PSNR and 30.054 dB mean per-pair PSNR.
PSNR measures paired pixel similarity, not aesthetics, prompt adherence, anatomy, typography, or human preference. Small spatial changes can sharply reduce PSNR. LPIPS and human-preference metrics were not computed.
Full records:
- [`assets/quality_metrics.json`](assets/quality_metrics.json)
- [`assets/VRAM_MATRIX.md`](assets/VRAM_MATRIX.md)
- [`assets/grid_verification.json`](assets/grid_verification.json)
## VRAM methodology
Each benchmark cell used a fresh ComfyUI process. Total physical GPU memory was sampled every 20 ms under an exclusive GPU lock. Settings were 1024×1024, batch 1, 20 steps, Euler/simple, CFG 1.0, FLUX guidance 3.5. `--lowvram` is an offloading result; hardware, runtime versions, attention backend, and workflow can materially change memory use.
The no-offloading `--highvram` maximum peaks for the four publication pipelines were 19,998 MiB, 21,342 MiB, 21,342 MiB, and 18,974 MiB in the table's order. Original BF16 hit a capacity failure in that 24 GiB environment.
## Quantization policies
- **Whole W8A8:** all 304 paper-scoped 2-D transformer-block linear weights use tensorwise INT8 ConvRot; 476 tensors remain BF16 or otherwise preserved.
- **Partial INT8:** profile-based selective INT8 ConvRot with a larger BF16-preserved set. This is the largest and highest-PSNR tested quantized DiT.
- **Paper W4A4:** 301 paper-scoped transformer-block linears use W4A4; 479 tensors remain BF16 or otherwise preserved. ComfyUI's fused `single_blocks.*.linear1` expands some paper component-level decisions to the fused tensor.
- **INT8 T5-XXL:** 168 encoder linear projections use W8A8; 51 embedding, relative-attention-bias, and normalization tensors remain BF16. This T5 policy is experimental and was not evaluated by the ConvRot paper.
See [`manifests/conversion_summaries.json`](manifests/conversion_summaries.json) and [`release_manifest.json`](release_manifest.json) for artifact-level provenance and hashes.
## Limitations
- These are post-training quantizations, not fine-tuned models.
- The W4A4 result diverges substantially more from BF16 and is labeled experimental.
- INT8 T5 image similarity does not establish long-prompt comprehension equivalence.
- Results cover 20 fixed prompts and one validated runtime/hardware environment.
- No safety classifier or content filter is included.
- Quantized checkpoints do not remove the limitations or usage restrictions of FLUX.1-dev.
## License
The quantized weights are derivatives of FLUX.1-dev and are provided under the **FLUX.1-dev Non-Commercial License**. Review and comply with the exact upstream license:
- License: https://huggingface.co/black-forest-labs/FLUX.1-dev/blob/3de623fc3c33e44ffbe2bad470d0f45bccf2eb21/LICENSE.md
- Base model: https://huggingface.co/black-forest-labs/FLUX.1-dev
The helper workflow/scripts do not grant additional rights to the model weights. See [`LICENSE_NOTICE.md`](LICENSE_NOTICE.md).
## Credits
- Base model: Black Forest Labs, [`FLUX.1-dev`](https://huggingface.co/black-forest-labs/FLUX.1-dev)
- Quantized runtime: [ComfyUI](https://github.com/comfyanonymous/ComfyUI) and [comfy-kitchen](https://github.com/Comfy-Org/comfy-kitchen)
- Conversion tooling: [Starnodes Model Converter](https://github.com/Starnodes2024/comfyui-starnodes-modelconverter)
- Method reference: *ConvRot: Rotation-Based Plug-and-Play 4-bit Quantization for Diffusion Transformers*, arXiv:2512.03673
|