--- license: mit library_name: mage_flow pipeline_tag: image-to-image base_model: microsoft/Mage-Flow-Edit base_model_relation: quantized tags: - xpo3 - mage-flow - mage-flow-edit - image-editing - image-to-image - nvfp4 - fp8 - blackwell - quantization ---
A 30-step native Blackwell W4A4 release of Mage-Flow Edit.
Original model · ComfyUI source weights · FP8 text encoder source · XPO3 text-to-image sibling · XPO3 Edit-Turbo sibling
--- ## Download | File | Purpose | Size | |:--|:--|--:| | `Mage-Flow-Edit-XPO3-NVFP4.safetensors` | Self-contained Edit transformer | 4.65 GB | | `qwen3vl_4b_fp8_scaled.safetensors` | Shared scaled-FP8 text encoder | 5.24 GB | | `Mage-Flow-VAE.safetensors` | Shared Mage VAE | 345 MB | The root `config.json` is included for Hugging Face model/download accounting. ## Quick start Linux, Python 3.11, CUDA 13, and an NVIDIA Blackwell SM120 GPU are required. ```bash hf download ajh-code/Mage-Flow-Edit-XPO3-NVFP4 --local-dir mage-edit-xpo3 cd mage-edit-xpo3 python3.11 -m venv .venv . .venv/bin/activate pip install -r requirements.txt python edit.py reference.jpg "Change the background; keep the subject unchanged." --output edited.png ``` The promoted profile is 30 steps, CFG 5, image-only fused GELU, calibrated FP4 bridge on all 12 blocks, direct-HND attention on steps 7-29, and exact SDPA fallback on steps 0-6. Every optimization has a CLI disable switch. ## Measured performance | GPU · 512-long-edge edit | BF16 hot denoise | XPO3 hot denoise | Speedup | |:--|--:|--:|--:| | RTX 5080 · Dog background | 4.2491 s | 2.5709 s | 1.65× | | RTX 5080 · Fruit tablecloth | 4.4839 s | 2.6815 s | 1.67× | | RTX 5060 Ti · Dog background | 8.6019 s | 4.7614 s | 1.81× | | RTX 5060 Ti · Fruit tablecloth | 9.3282 s | 5.1300 s | 1.82× | Hot denoise timings are local matched measurements, not universal end-to-end claims. The RTX 5060 Ti runs also used 31-33% less peak allocation than BF16. ## Validated scope - Same-seed dog and fruit edits were deterministic across fresh and hot runs. - The standalone package reproduced both approved RTX 5080 composed outputs pixel-exactly. - Each release-only edit recorded 552 direct routes, 168 exact fallback routes, all 12 bridge blocks, and restored every temporary patch. This is an XPO3 runtime package, not a generic portable quantization format. The bundled native libraries target Linux x86-64, Python 3.11, PyTorch 2.13.0+cu130, and SM120. MIT applies to the XPO3 package and Mage-derived runtime. The scaled-FP8 Qwen3-VL component and bundled SpargeAttn code are Apache-2.0; see `THIRD_PARTY_NOTICES.md`.