ajh-code's picture
Add files using upload-large-folder tool
05adc6a verified
|
Raw
History Blame Contribute Delete
6.4 kB
---
license: other
library_name: diffusers
pipeline_tag: text-to-image
base_model: microsoft/Mage-Flow
base_model_relation: quantized
tags:
- ajh
- mage-flow
- mage-flow-nvfp4-quality-ajh
- nvfp4
- blackwell
- qwen3-vl
- text-to-image
- quantization
- quality
---
# Mage-Flow-NVFP4-Quality-AJH
**Mage-Flow-NVFP4-Quality-AJH** is a portable, runnable quantized version of
[`microsoft/Mage-Flow`](https://huggingface.co/microsoft/Mage-Flow).
Standalone Hugging Face Mage-Flow repository with all 24 image MLP projections kept in BF16 and the 24 text MLP projections stored as native NVFP4.
This repository is intended to remain discoverable as a quantized
`microsoft/Mage-Flow` derivative while avoiding any runtime dependency on
a separate local `models/` checkout. The complete transformer component,
quantized text encoder, VAE, scheduler, vendored inference code, and
native runtime are all packaged inside the repository layout.
Related releases:
- [Fast / maximum compression](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-AJH)
- [Balanced](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-Balanced-AJH)
- [Quality](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-Quality-AJH)
- [ComfyUI custom nodes](https://github.com/AJH-Code/ComfyUI-MageFlow-NVFP4-AJH)
## Showcase
![Cyborg woman gazing toward a star-filled sky](examples/cyborg_stargaze_quality.png)
Generated directly with this `quality` package, without upscaling
or post-processing.
> A 4K resolution high detail photo realistic image of the top half of a cyborg woman with dark black hair, striking blue eyes that have a very subtle glow in the iris, standing side profile, head tilted up towards the sky with a questioning expression, she has subtle gaps in her skin that hint at a robotic nature, outdoor forest night setting, sky filled with bright brilliant stars that glow against the dark setting, nebula visible
Settings: 1280×1280, 20 steps, CFG 5, static shift 6, seed
`3334072683`.
## Matched BF16 comparison
This comparison uses the same frozen prompt, seed `1795681438`,
1024×1024 resolution, 20 steps, CFG 5, and static shift 6. It compares
the complete BF16 pipeline against the complete Quality package,
including its quantized text encoder.
| Full BF16 reference | Quality: 24 NVFP4 + 24 BF16 image projections |
|---|---|
| ![Full BF16 reference](examples/comparison_bf16.png) | ![Quality NVFP4 package](examples/comparison_quality_nvfp4.png) |
| 21.79 s denoise | 18.48 s denoise |
## Transformer policy
- Native NVFP4 transformer projections: `24`
- BF16 passthrough transformer projections: `24`
- Default native runtime activation search: `amax`
- Default up activation multiplier: `1`
- Default down activation multiplier: `1`
BF16 passthrough modules:
- `transformer_blocks.0.img_mlp.net.0.proj`
- `transformer_blocks.0.img_mlp.net.2`
- `transformer_blocks.1.img_mlp.net.0.proj`
- `transformer_blocks.1.img_mlp.net.2`
- `transformer_blocks.2.img_mlp.net.0.proj`
- `transformer_blocks.2.img_mlp.net.2`
- `transformer_blocks.3.img_mlp.net.0.proj`
- `transformer_blocks.3.img_mlp.net.2`
- `transformer_blocks.4.img_mlp.net.0.proj`
- `transformer_blocks.4.img_mlp.net.2`
- `transformer_blocks.5.img_mlp.net.0.proj`
- `transformer_blocks.5.img_mlp.net.2`
- `transformer_blocks.6.img_mlp.net.0.proj`
- `transformer_blocks.6.img_mlp.net.2`
- `transformer_blocks.7.img_mlp.net.0.proj`
- `transformer_blocks.7.img_mlp.net.2`
- `transformer_blocks.8.img_mlp.net.0.proj`
- `transformer_blocks.8.img_mlp.net.2`
- `transformer_blocks.9.img_mlp.net.0.proj`
- `transformer_blocks.9.img_mlp.net.2`
- `transformer_blocks.10.img_mlp.net.0.proj`
- `transformer_blocks.10.img_mlp.net.2`
- `transformer_blocks.11.img_mlp.net.0.proj`
- `transformer_blocks.11.img_mlp.net.2`
## Notes
- Keeps every image-side MLP projection in BF16.
- Quantizes only the text-side MLP projections.
- Intended as the most conservative standalone mixed-transformer package candidate.
The Qwen3-VL text encoder uses the same mixed policy as the Fast
release: 224 NVFP4 projections in blocks 2–33, 14 FP8 projections in
blocks 1 and 34, with blocks 0 and 35 plus embeddings, norms, biases,
and the vision tower retained in BF16.
## Requirements and generation
The tested stack is Linux x86-64, NVIDIA Blackwell SM120, CUDA 13.1,
Python 3.11, PyTorch `2.13.0+cu130`, `comfy-kitchen==0.2.22`, and
`flash-attn==2.8.3`. Install into a virtual environment using the
included `requirements.txt`; do not install these packages system-wide.
```bash
python3.11 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install -r requirements.txt
CUDA_HOME=/usr/local/cuda-13.1 python -m pip install --no-build-isolation flash-attn==2.8.3
CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \
--model ajh-code/Mage-Flow-NVFP4-Quality-AJH \
--prompt 'A detailed watercolor fox reading under an old oak tree' \
--output fox.png --height 1024 --width 1024 --steps 20 --seed 1
```
The included binaries target the tested stack. Run `build_native.sh`
after changing PyTorch, CUDA, or the C++ ABI.
## Measured tradeoff
The controlled matched benchmark measured the Balanced transformer at
about `1.43x` BF16 throughput (`30%` less generation time), and the
Quality transformer at about `1.22x` BF16 throughput (`18%` less
generation time). Projected transformer plus text-checkpoint storage
is about `9.89 GB` for Balanced and `10.97 GB` for Quality, versus
`17.12 GB` for BF16. These are policy-level measurements from the
research suite, not universal hardware guarantees.
## Runtime caveats
- Requires an NVIDIA Blackwell SM120 GPU.
- Uses the packaged loader and native runtime; stock Diffusers does not
natively understand `mage_flow_nvfp4_*` transformer modules.
- The package-level runtime defaults are applied only when the caller has
not already set the corresponding `MAGE_NVFP4_*` environment variables.
- Generation and text-to-image are tested; Base, Turbo, editing, CUDA
graph compatibility, and non-SM120 GPUs are not claimed.
## Validate the package
```bash
python validate_release.py
```
## License and attribution
Mage-Flow and the vendored Mage inference source are Copyright (c) 2026
Microsoft and MIT licensed. Qwen3-VL and the mixed NVFP4/FP8 text
checkpoint are Apache-2.0 licensed. See the included license files and
third-party notices for details.