--- license: other library_name: diffusers pipeline_tag: text-to-image base_model: microsoft/Mage-Flow base_model_relation: quantized tags: - ajh - mage-flow - mage-flow-nvfp4-ajh - nvfp4 - blackwell - qwen3-vl - text-to-image - quantization --- # Mage-Flow-NVFP4-AJH **Mage-Flow-NVFP4-AJH** is a portable, runnable NVFP4 package for [`microsoft/Mage-Flow`](https://huggingface.co/microsoft/Mage-Flow). It combines a native NVFP4 Mage transformer with a complete mixed NVFP4/FP8 Qwen3-VL text encoder. ## ComfyUI Ready-to-use custom nodes are available at [`AJH-Code/ComfyUI-MageFlow-NVFP4-AJH`](https://github.com/AJH-Code/ComfyUI-MageFlow-NVFP4-AJH). The plugin provides loader and generation nodes, downloads this complete model repository, and returns a standard ComfyUI `IMAGE`. Its initial release targets Linux x86-64, Python 3.11, and NVIDIA Blackwell SM120 GPUs; follow the exact runtime requirements documented in the plugin repository. ## Showcase ![Cyborg woman gazing toward a star-filled sky, generated by Mage-Flow-NVFP4-AJH](examples/cyborg_stargaze_nvfp4.png) Generated directly with the released full NVFP4 package, without upscaling or post-processing. > A 4K resolution high detail photo realistic image of the top half of a > cyborg woman with dark black hair, striking blue eyes that have a very subtle > glow in the iris, standing side profile, head tilted up towards the sky with > a questioning expression, she has subtle gaps in her skin that hint at a > robotic nature, outdoor forest night setting, sky filled with bright > brilliant stars that glow against the dark setting, nebula visible Settings: 1280×1280, 20 steps, CFG 5, static shift 6, seed `3334072683`. This is a complete Hugging Face component-layout repository, not an overlay: ```text model_index.json transformer/ config.json diffusion_pytorch_model-00001-of-00004.safetensors diffusion_pytorch_model.safetensors.index.json text_encoder/ config.json model.safetensors vae/ scheduler/ ``` The transformer shards include every retained BF16 tensor alongside the NVFP4 module state. The VAE, scheduler, text-encoder configuration, tokenizer, and processor are also included. Running the downloaded repository does not fetch BF16 weights from the base model. ## Quantization policy Mage transformer: - 48 image/text MLP up/down projections use resident W4A4 NVFP4. - Attention, modulation, normalization, and all non-target tensors remain BF16. - The 48 original BF16 target weight matrices are absent. Their biases remain BF16 and are stored normally inside the quantized modules. - The complete transformer is stored as four indexed standard Safetensors shards. Packed NVFP4 buffers use their module state-dict names: `packed_weight`, `weight_scales`, `weight_scale`, and `bias`. Qwen3-VL text encoder: - Blocks 2–33: 224 NVFP4 projections. - Blocks 1 and 34: 14 FP8 projections. - Blocks 0 and 35, embeddings, norms, biases, and the vision tower remain BF16. - Packaged text checkpoint size: `4,031,376,064` bytes. - The packaged loader constructs the model directly from this file: 475 non-quantized tensors and 238 packed projections, with no unresolved meta tensors. ## Requirements The prebuilt runtime was tested on: - NVIDIA RTX 50-series / SM120 - Linux x86-64 - CUDA 13.1 - Python 3.11 - PyTorch `2.13.0+cu130` - `comfy-kitchen==0.2.22` - `flash-attn==2.8.3` Create a local environment: ```bash python3.11 -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip python -m pip install -r requirements.txt CUDA_HOME=/usr/local/cuda-13.1 \ python -m pip install --no-build-isolation flash-attn==2.8.3 ``` The included binaries are for the exact tested stack. Rebuild them after changing PyTorch, CUDA, or the C++ ABI: ```bash CUDA_HOME=/usr/local/cuda-13.1 \ PYTHON_BIN="$PWD/.venv/bin/python" \ ./build_native.sh ``` ## Generate an image Expose exactly one SM120 GPU: ```bash CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \ --prompt 'A detailed watercolor fox reading under an old oak tree' \ --output fox.png \ --height 1024 \ --width 1024 \ --steps 20 \ --seed 1 ``` When running from another copy of the scripts, `--model` also accepts the Hub repository id or a downloaded standard-layout directory: ```bash CUDA_VISIBLE_DEVICES=0 .venv/bin/python generate.py \ --model ajh-code/Mage-Flow-NVFP4-AJH \ --prompt 'A lighthouse poster reading "ARCTIC LOOP"' \ --output lighthouse.png ``` The command refuses to overwrite an existing output and writes a companion JSON report containing coverage, memory, timing, and environment information. ## Measured results Transformer: - Loaded packed transformer allocation: `5,628,438,016` bytes. - Measured transformer allocation saving: `2,604,638,208` bytes (`2.4258 GiB`). - Target BF16 source reads: `0`. - Representative real up-projection: `3.26x` eager and `3.25x` compiled speedup over BF16. Text encoder: - BF16 language-stack residency: `7.5453 GiB`. - Packed mixed residency: `3.0427 GiB`. - Measured saving: `4.5026 GiB`. Combined validation on an RTX 5060 Ti 16 GB: - All seven coverage, source-access, allocation, finite, pixel, latent, and text-policy gates passed. - Peak text stage with the packed transformer still resident: `10,437,891,584` bytes (`9.72 GiB`). - VAL-07 pixel cosine/NRMSE versus frozen BF16: `0.9927827428` / `0.1200022063`. - VAL-07 latent cosine/NRMSE versus frozen BF16: `0.9339298065` / `0.3618087155`. - The requested strings `ARCTIC LOOP` and `NORTHERN COAST` were both rendered correctly. ![Combined native NVFP4 VAL-07 output](examples/VAL-07_combined_native.png) The standard-layout packaged loader is validated independently of the research-tree loader. It installs 48 native Mage projections and 238 packed Qwen projections and rejects any checkpoint containing the replaced BF16 transformer targets. ## Quality and speed caveats The mixed text encoder is functional and produced visually strong downstream images, but it does not meet our unusually strict embedding-similarity gate: mean token/pooled cosine was `0.9531366898` / `0.9751400001`. All ten content screening verdicts and category lists remained unchanged. The ten-case text workload was slower with the packed runtime (`30.36 s`) than BF16 (`18.29 s`). This package therefore claims major text-encoder VRAM savings, not a text-encoding speedup. Transformer projections are materially faster, but repeated matched end-to-end timing has not been completed. Additional limitations: - Native execution is currently SM120-only. - The repository follows the Hugging Face component and Safetensors layout, but the new `mage_flow_nvfp4` runtime is not yet built into stock Diffusers. Use the included loader. - CUDA graph compatibility is not claimed. - Generation and text-to-image are tested; Base, Turbo, and editing variants are not. - A controlled photorealistic 1024×1024 comparison found that the current NVFP4 transformer can produce fuzzy hair or skin microtexture for some seeds. Using the original BF16 text encoder did not remove it; the transformer quantization is the primary associated variable. - Four transformer-only held-out cases and one combined held-out case have been evaluated. This is not a broad benchmark. ## Validate the download `MANIFEST.json` records every distributed file except itself: ```bash .venv/bin/python validate_release.py ``` Hashing the transformer shards and text checkpoint can take a little while. ## License and attribution - Mage-Flow and the vendored Mage inference source are Copyright (c) 2026 Microsoft and licensed under MIT. See `LICENSE` and `licenses/MAGE-MIT.txt`. - Qwen3-VL and the mixed NVFP4/FP8 text checkpoint are licensed under Apache-2.0. See `licenses/QWEN-APACHE-2.0.txt`. - The text checkpoint was produced by [`InsecureErasure/Qwen3-VL-4B-Instruct-NVFP4`](https://huggingface.co/InsecureErasure/Qwen3-VL-4B-Instruct-NVFP4) using learned rounding and `comfy-kitchen`.