File size: 21,611 Bytes
70ac946
42e411d
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42e411d
 
 
70ac946
42e411d
11ce194
42e411d
 
0e79352
70ac946
 
42e411d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
70ac946
b24a274
 
 
70ac946
42e411d
 
 
 
 
 
 
 
 
 
 
 
 
 
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
 
b24a274
70ac946
 
 
b24a274
 
 
 
 
 
70ac946
 
 
 
 
42e411d
b24a274
42e411d
 
b24a274
 
 
70ac946
 
 
 
 
 
41c029b
 
 
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b24a274
 
 
 
 
 
 
70ac946
42e411d
 
 
 
 
 
 
 
70ac946
41c029b
b24a274
42e411d
 
 
 
 
 
 
b24a274
42e411d
 
 
 
 
 
 
b24a274
42e411d
 
 
 
 
 
 
 
 
 
 
 
 
 
41c029b
42e411d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
70ac946
b24a274
70ac946
42e411d
 
b24a274
 
 
70ac946
 
 
b24a274
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
70ac946
 
 
42e411d
 
 
70ac946
11ce194
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
f82f9a3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
42e411d
 
f82f9a3
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b24a274
 
 
 
 
 
 
 
 
70ac946
b24a274
 
 
 
 
 
 
 
70ac946
 
 
 
 
 
b24a274
70ac946
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b24a274
70ac946
 
b24a274
 
70ac946
 
 
 
 
 
b24a274
 
 
 
70ac946
 
 
 
 
 
42e411d
 
 
 
 
 
 
 
 
 
 
 
70ac946
 
 
 
 
 
 
 
b24a274
 
 
 
70ac946
 
b24a274
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
---
license: mit
pipeline_tag: text-to-image
base_model: microsoft/Mage-Flow-Turbo
base_model_relation: quantized
tags:
  - xpo3
  - comfyui
  - mage-flow
  - mage-flow-turbo
  - nvfp4
  - fp8
  - blackwell
  - qwen3-vl
  - text-to-image
  - quantization
---

<p align="center">
  <img src="xpo3-nvfp4-icon.png" width="360" alt="XPO3 NVFP4 speed mascot">
</p>

# Mage-Flow Turbo NVFP4 — V2

> **Follow Mage-Flow NVFP4 development, benchmarks, and release updates at
> [Arands.com](https://arands.com/).**

**Speed and quality through cross-precision optimization.**

### AJH Mage-Flow releases

| Release | Use it when |
|---|---|
| **Mage-Flow Turbo NVFP4 — XPO3 V2** (this release) | Recommended: the fastest validated Turbo runtime and the best current Turbo choice. |
| [Mage-Flow Turbo NVFP4 Balanced](https://huggingface.co/ajh-code/Mage-Flow-Turbo-NVFP4-Balanced-AJH) | You specifically need the earlier standard packaged Turbo variant. |
| [Mage-Flow NVFP4 Fast](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-AJH) | 20-step Mage-Flow with maximum compression. |
| [Mage-Flow NVFP4 Balanced](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-Balanced-AJH) | 20-step Mage-Flow with the balanced precision policy. |
| [Mage-Flow NVFP4 Quality](https://huggingface.co/ajh-code/Mage-Flow-NVFP4-Quality-AJH) | 20-step Mage-Flow with more BF16 image-path retention. |

Only the Balanced transformer policy was packaged for Turbo before XPO3. Fast
and Quality are non-Turbo Mage-Flow releases.

This is the V2 runtime release for the existing model files. **XPO3 is the
name of the bundled runtime profile, not a new model format or loader.** With
the current **ComfyUI-XPO3 v0.3.0** custom node, the weights run through the
promoted XPO3 runtime stack: a four-step, ComfyUI-oriented derivative of
[`microsoft/Mage-Flow-Turbo`](https://huggingface.co/microsoft/Mage-Flow-Turbo)
optimized for native Blackwell NVFP4 execution. The additional speed is a
runtime update and does not require rebuilding or redownloading the model
weights.

**Recommended AJH Mage-Flow Turbo release.** In the reviewed general-prompt
set, XPO3 follows the original BF16 Turbo closely while delivering the fastest
validated runtime in this project. Individual trajectories can occasionally
improve a detail as well: the published golden-retriever comparison avoids the
duplicate-tail artifact present in its BF16 counterpart. That is an observed
example, not a claim that XPO3 is universally better than BF16.

Microsoft reports **0.59 s/image at 1024×1024 on one A100** for the original
Mage-Flow-Turbo, at roughly **18–20 GB** peak memory. This V2 runtime measured
**0.5943 s** median at 1024×1024 on an RTX 5080 with **11.85 GiB** peak
allocated. These are comparable latency points, not an apples-to-apples
cross-hardware benchmark. See the
[upstream model card](https://huggingface.co/microsoft/Mage-Flow-Turbo).

> **Three precisions. One fused runtime.**
>
> NVFP4 supplies transformer speed, scaled FP8 compresses the Qwen3-VL text
> encoder, and BF16 is retained only where matched testing found that four-bit
> execution caused visible quality damage.

## Native W4A4 runtime—not a weight-only dump

XPO3 runs its quantized transformer projections through the native Blackwell
**W4A4** path: weights are stored as block-16 NVFP4, BF16 activations are
dynamically quantized to NVFP4 with role-specific scaling, and the matrix
multiplication executes as native NVFP4. It does not expand the packed weights
to BF16 and fall back to a BF16 GEMM.

That fast path is one part of the promoted measured runtime stack:

- 44 MLP projections and 24 offline-fused QKV groups use native W4A4;
- four quality-sensitive late image MLP projections remain BF16;
- 23 eligible up projections use fused bias-plus-GELU NVFP4 execution;
- calibrated FP4 image-MLP bridge routing is enabled on blocks `0-8`;
- official Sparge/Sage2 routing runs on steps `1-2` across all 12 blocks;
- the CFG1, batch-one direct-HND processor is used on that routed path;
- steps `0` and `3`, disabled toggles, and unsupported envelopes stay on the
  packaged exact fallback path;
- the complete pipeline stays resident, while prompt conditioning and the
  fixed policy prefix are cached.

| Result | Measured value | What it establishes |
|---|---:|---|
| RTX 5080, 1024×1024 V2 hot generation | **0.5943 s** median; **0.5818 s** best | Current V2 repeat, resident model and reusable prompt/screening cache |
| RTX 5080, 1280×1280 promoted-stack median | **0.9570 s** | Validated 3-prompt x 3-seed matrix result |
| RTX 5080, 1024×1024 fresh-prompt generation | **2.7974 s** median | Resident model; empty prompt and screening caches |
| RTX 5080, 1280×1280 fresh-prompt generation | **3.1302 s** median | Resident model; empty prompt and screening caches |
| RTX 5060 Ti, 1024×1024 promoted-stack median | **1.3012 s** | Validated 16 GB entry-level Blackwell result |
| RTX 5060 Ti, 1280×1280 promoted-stack median | **2.1308 s** | Validated 16 GB entry-level Blackwell result |
| Full promoted stack versus the packaged resident FA2 baseline | **1.177x-1.212x faster** | Current release-runtime gain, not a BF16 claim |
| Balanced NVFP4 versus BF16 in the matched pre-fusion study | **1.39× faster** | Transformer precision-policy contribution |
| Scaled-FP8 text conditioning versus BF16 | **0.9932–0.9971 cosine** | Frozen prompt-suite fidelity |
| Exact-text pooled conditioning versus BF16 | **0.9986 cosine** | Text-encoder pooled-vector fidelity |
| Balanced final image versus matched BF16 | **23.78 dB / 0.0283 MAE** | Pixel-level precision-study comparison |

The conditioning cosines describe text-encoder outputs, not an overall image
quality score. Image quality was also checked with frozen same-seed prompt
suites and direct visual review.

## Download

This repository intentionally contains the three model files useful to
ComfyUI:

| File | Place in ComfyUI | Size |
|---|---|---:|
| `Mage-Flow-Turbo-XPO3-NVFP4.safetensors` | `models/diffusion_models/` | 4.87 GB |
| `qwen3vl_4b_fp8_scaled.safetensors` | `models/text_encoders/` | 5.24 GB |
| `Mage-Flow-VAE.safetensors` | `models/vae/` | 345 MB |

The diffusion model is one file rather than four Hugging Face transformer
shards. The complete three-weight download is 10.45 GB.

Install
[`ComfyUI-XPO3`](https://github.com/AJH-Code/ComfyUI-XPO3), restart ComfyUI,
and use:

1. **Load Mage-Flow Turbo XPO3**
2. **Mage-Flow Turbo XPO3 Generate**

The custom-node repository includes an API-format starter workflow that
defaults to 1280×1280. Recommended settings are four steps, CFG 1, and static
shift 6.

The Generate node also exposes independent runtime toggles for fused GELU-up,
the FP4 image-MLP bridge, bridge-block selection, Sparge/Sage2 attention,
direct-HND, and selected attention steps and blocks.

### V1 → V2

No model-weight redownload is required. V2 is a runtime and integration
upgrade: fused GELU-up, the calibrated FP4 bridge, the validated accelerated
attention/direct-HND route, explicit per-feature fallbacks, a 1280-default
ComfyUI workflow, and a reproducible benchmark runner in ComfyUI-XPO3.

## Speed: 0.59 s at 1024, 0.96 s at 1280 on RTX 5080

Measured on 16 GB RTX 5080 and RTX 5060 Ti cards under Linux, Python 3.11,
and PyTorch `2.13.0+cu130`, with one resident model load per matrix arm.
The table distinguishes repeated interactive generation from the first image
of a new prompt:

- **Hot**: model remains resident; prompt conditioning and its fixed screening
  prefix are reusable.
- **Fresh prompt**: model remains resident, but the prompt-conditioning and
  screening caches begin empty. This is not a cold process start.

Prompt conditioning is the Qwen3-VL embedding pass used by the diffusion
model; it is independent of the random seed and can therefore be reused for
re-rolls. The screening pass checks a prompt against a fixed content-policy
prefix. Its immutable prefix KV state is also cacheable and does not change
the generated image.

| GPU | Resolution | Hot generation | Fresh prompt |
|---|---|---:|---:|
| **RTX 5080** | 1024×1024 | **0.5943 s median** (**0.5818 s** best) | **2.7974 s median** |
| **RTX 5080** | 1280×1280 | **0.9570 s median** | **3.1302 s median** |
| RTX 5060 Ti | 1024×1024 | **1.3012 s median** | — |
| RTX 5060 Ti | 1280×1280 | **2.1308 s median** | — |

The 1024 and 1280 hot figures are V2 resident runs. The `0.9570 s` 1280
figure is the more robust three-prompt × three-seed matrix median; the 1024
figure is a current V2 two-seed repeat. In the completed multi-prompt matrix,
the promoted stack delivered a consistent **1.177x-1.212x** end-to-end gain
over the previous packaged resident FA2 baseline.

For process startup on the RTX 5080, model load plus resident placement takes
about **9.46-10.14 s** before the first generation. That one-time cost is kept
separate from the generation timings above.

### Reproduce the hot/fresh-prompt protocol

The matching public runner ships with
[`ComfyUI-XPO3`](https://github.com/AJH-Code/ComfyUI-XPO3). From the plugin
directory, using the same ComfyUI Python environment and model filenames:

```bash
python benchmark.py --width 1024 --height 1024 \
  --output xpo3_1024_benchmark.json
```

It runs a fixed three-prompt/three-seed suite at four steps, CFG 1, and static
shift 6; records a fresh-cache warm-up and repeated hot generations for each
prompt; and writes timing, environment, and active-feature evidence to JSON.
Your GPU, driver, and ComfyUI environment will produce different numbers.

Peak allocated memory increased modestly versus that packaged baseline:

- 1024×1024: **11.85 GiB** peak allocated (about **+51.0 MiB**)
- 1280×1280: **12.33 GiB** peak allocated (about **+78.6 MiB**)

The complete promoted stack still remained under 16 GB on both validated
cards.

Timings are hardware- and environment-dependent.

## Promoted runtime stack and fallback rules

The current v0.3.0 custom-node runtime includes:

- fused GELU-up on eligible NVFP4 up projections;
- calibrated FP4 image-MLP bridge on image blocks `0-8`;
- official Sparge/Sage2 meansim top-k `1.0`, smooth-K, all 12 blocks on steps
  `1-2`;
- the validated CFG1, batch-one direct-HND processor on that routed path;
- exact packaged fallback on steps `0` and `3`;
- independent per-feature toggles exposed through the ComfyUI generate node.

Fallback behavior is intentional rather than best-effort:

- when fused GELU-up is disabled, the wrapper calls the exact original
  activation module;
- when the bridge is disabled, a bridge block is deselected, or an unsupported
  shape/stream/device condition is seen, that block falls back before native
  launch;
- when attention acceleration is disabled, CFG is not `1.0`, steps are not
  `4`, the pipe is unsupported, or the Sparge dependency is unavailable, the
  packaged attention path remains active;
- direct-HND is only used on the supported CFG1, batch-one routed envelope.

## Quality

XPO3 does not quantize every eligible projection merely to maximize the
compression ratio. **This Turbo release ships one selected Balanced policy.**
The Fast and Quality rows below are research controls used to choose it; they
are not additional downloadable Turbo profiles.

### 1280×1280 BF16 versus XPO3 comparison set

Each pair below uses the same prompt, seed, 1280×1280 resolution, four steps,
CFG 1, and static shift 6. **Original Mage-Flow-Turbo BF16 is on the left;
XPO3 NVFP4 is on the right.** The images are direct decodes with no upscaling
or postprocessing. Click any comparison to inspect it at full size.

| Alpine lake | Golden retriever |
|---|---|
| [![BF16 versus XPO3 alpine lake comparison](examples/bf16-vs-xpo3-1280/00-alpine-lake.jpg)](examples/bf16-vs-xpo3-1280/00-alpine-lake.jpg) | [![BF16 versus XPO3 golden retriever comparison](examples/bf16-vs-xpo3-1280/01-golden-retriever.jpg)](examples/bf16-vs-xpo3-1280/01-golden-retriever.jpg) |
| Reading nook | Perfume product |
| [![BF16 versus XPO3 reading nook comparison](examples/bf16-vs-xpo3-1280/02-reading-nook.jpg)](examples/bf16-vs-xpo3-1280/02-reading-nook.jpg) | [![BF16 versus XPO3 perfume product comparison](examples/bf16-vs-xpo3-1280/03-perfume-product.jpg)](examples/bf16-vs-xpo3-1280/03-perfume-product.jpg) |
| Vintage car | Fisherman portrait |
| [![BF16 versus XPO3 vintage car comparison](examples/bf16-vs-xpo3-1280/04-vintage-car.jpg)](examples/bf16-vs-xpo3-1280/04-vintage-car.jpg) | [![BF16 versus XPO3 fisherman portrait comparison](examples/bf16-vs-xpo3-1280/05-fisherman-portrait.jpg)](examples/bf16-vs-xpo3-1280/05-fisherman-portrait.jpg) |
| Ramen | Butterfly macro |
| [![BF16 versus XPO3 ramen comparison](examples/bf16-vs-xpo3-1280/06-ramen.jpg)](examples/bf16-vs-xpo3-1280/06-ramen.jpg) | [![BF16 versus XPO3 butterfly macro comparison](examples/bf16-vs-xpo3-1280/07-butterfly-macro.jpg)](examples/bf16-vs-xpo3-1280/07-butterfly-macro.jpg) |
| Street portrait | Astronaut greenhouse |
| [![BF16 versus XPO3 street portrait comparison](examples/bf16-vs-xpo3-1280/08-street-portrait.jpg)](examples/bf16-vs-xpo3-1280/08-street-portrait.jpg) | [![BF16 versus XPO3 astronaut greenhouse comparison](examples/bf16-vs-xpo3-1280/09-astronaut-greenhouse.jpg)](examples/bf16-vs-xpo3-1280/09-astronaut-greenhouse.jpg) |
| Watercolor fox | Glass house |
| [![BF16 versus XPO3 watercolor fox comparison](examples/bf16-vs-xpo3-1280/10-watercolor-fox.jpg)](examples/bf16-vs-xpo3-1280/10-watercolor-fox.jpg) | [![BF16 versus XPO3 glass house comparison](examples/bf16-vs-xpo3-1280/11-glass-house.jpg)](examples/bf16-vs-xpo3-1280/11-glass-house.jpg) |

Across this general-prompt set, XPO3 closely follows BF16 subject, framing,
lighting, palette, and overall quality. Small trajectory-dependent differences
remain, so XPO3 should not be interpreted as pixel-identical or lossless.

At 1024×1024 in the precision-policy study:

| Transformer profile | Denoise | Peak allocated | PSNR / MAE versus BF16 |
|---|---:|---:|---:|
| Fast: 48 NVFP4 MLP projections | 3.44 s | 6.09 GiB | 18.27 / 0.0747 |
| **Chosen Balanced profile** | **3.51 s** | **6.29 GiB** | **23.78 / 0.0283** |
| Quality: image MLPs in BF16 | 4.20 s | 7.31 GiB | 23.38 / 0.0312 |
| BF16 transformer | 4.89 s | 8.52 GiB | reference |

That study isolated transformer precision before the later resident,
single-branch, QKV-fusion, cache, and fused-GELU speed work, so its latency
numbers should not be mixed with the final warm-runtime table above.

### Representative 1280×1280 XPO3 output

![Cyborg woman looking toward a star-filled sky, generated by XPO3](examples/cyborg_xpo3.png)

This is a direct, unprocessed four-step XPO3 decode: 1280×1280, CFG 1,
static shift 6, and seed `3334910515`.

> A 4K resolution high detail photo realistic image of the top half of a
> cyborg woman with dark black hair, striking blue eyes that have a very subtle
> glow in the iris, standing side profile, head tilted up towards the sky with
> a questioning expression, she has subtle gaps in her skin that hint at a
> robotic nature, outdoor forest night setting, sky filled with bright
> brilliant stars that glow against the dark setting, nebula visible

This intentionally difficult prompt combines photorealistic anatomy, exposed
machinery, damaged skin, glowing eyes, and complex night lighting. It is shown
as one representative current result, not as evidence that every prompt is
pixel-identical to BF16.

## What XPO3 changed

### 1. Architecture-aware precision selection

Forty-four Mage transformer MLP projections run in native NVFP4. Four
late image-stream projections remain BF16:

- `transformer_blocks.9.img_mlp.net.2`
- `transformer_blocks.10.img_mlp.net.2`
- `transformer_blocks.11.img_mlp.net.0.proj`
- `transformer_blocks.11.img_mlp.net.2`

Keeping these four projections in BF16 substantially improved the 1024
portrait result while adding far less size and latency than returning the
complete image stream to BF16.

### 2. Quality-selected activation scaling

BF16 activations are dynamically quantized to block-16 NVFP4 immediately
before native W4A4 execution. Scaling is selected by projection role: the
NVFP4 MLP down paths use a `0.75` activation-scale multiplier, while the other
native projections retain their validated scaling policy. Alternative
block-scale searches, weight `4-over-6`, and bias-correction candidates were
tested and rejected because lower local reconstruction error did not
consistently produce a better final diffusion trajectory.

### 3. Offline-fused NVFP4 attention

Each image/text Q, K, and V triplet is concatenated into one native NVFP4
projection. All 24 fused QKV groups are packed in the downloadable checkpoint:

- 72 original BF16 Q/K/V projections replaced;
- approximately 0.91 GiB lower peak allocation in the experimental matched
  attention test;
- no 6.29-second load-time attention packing step;
- no full-weight BF16 expansion or BF16 GEMM fallback for these projections.

### 4. Fused MLP execution

Twenty-three NVFP4 MLP up projections use a native cuBLASLt
up-projection-plus-bias-plus-GELU path. The final quality-sensitive BF16 image
up projection remains unfused, and the fused path remains independently
toggleable from the ComfyUI node.

### 5. Calibrated image-MLP bridge

The promoted runtime adds the calibrated FP4 image-MLP bridge on image blocks
`0-8`, using the pinned per-block scales selected during the bounded bridge
validation. The bridge is independently toggleable and retains the packaged
module as its exact off path.

### 6. Step-gated Sparge/Sage2 attention

The promoted runtime adds official Sparge/Sage2 routing on steps `1-2` across
all 12 blocks, with the validated CFG1, batch-one direct-HND processor on
that route. Steps `0` and `3` remain on the packaged exact path, and the
attention route is independently toggleable.

### 7. Resident and cached generation

The transformer, FP8 text encoder, and VAE fit together on the validated
16 GB GPU. CFG 1 uses a single conditional branch. Exact prompt conditioning
is cached independently of the random seed, and the fixed policy-screen prefix
uses a bounded KV cache.

### 8. Single-file deployment

The XPO3 diffusion checkpoint embeds the transformer structure, precision
policy, fused-attention mapping, native-runtime requirement, and base-model
relationship in safetensors metadata. It contains 437 tensors and is
approximately 977 MB smaller than the earlier four-shard Balanced
transformer.

## What XPO3 is—and is not

XPO3 is a checkpoint and runtime approach focused on speed and quality through
architecture-aware cross-precision deployment and fused native execution.
It uses NVIDIA's standard NVFP4 E2M1 block-16 representation; it does not
claim to invent a new floating-point datatype.

It is also not SVDQuant or a Nunchaku checkpoint. XPO3 does not use a
high-precision low-rank outlier branch. Its quality recovery comes from
measured precision placement and scale policy, while its speed comes from
native NVFP4 execution, projection fusion, epilogue fusion, single-branch
Turbo inference, and resident caching.

## Requirements and current limitations

- Linux x86-64
- NVIDIA Blackwell SM120
- under 16 GB on the validated promoted stack; 16 GB class validated
- Python 3.11
- PyTorch `2.13.0+cu130`
- TorchVision `0.28.0+cu130`
- TorchAudio `2.11.0+cu130`
- CUDA 13.x runtime; bundled binaries built with CUDA 13.1
- `flash-attn==2.8.3`

The first prebuilt runtime does not support Windows, WSL, pre-Blackwell GPUs,
Python 3.12/3.13, or other PyTorch/CUDA ABIs.

The custom-node plugin also bundles SpargeAttn runtime code and extensions
under `runtime/spas_sage_attn/` under Apache-2.0 terms. Its included license
text is at `third_party_licenses/SpargeAttn-LICENSE` in the plugin repository.

The files use ComfyUI's normal diffusion-model, text-encoder, and VAE
directories. Mage-Flow is not currently supported by ComfyUI's generic
diffusion loader, so this release requires the XPO3 custom nodes and does not
yet connect directly to stock KSampler, LoRA, ControlNet, or native Comfy
conditioning nodes.

### Image-quality limitations

Four-step Turbo is intentionally aggressive. Complex photorealistic prompts
that combine organic anatomy with exposed mechanics—such as the cyborg stress
prompt in this card—can produce seed-sensitive neck, shoulder, torso, or small-detail
defects. A matched ten-seed comparison found the same failure class in the
original all-BF16 Mage-Flow-Turbo model; it is not an XPO3-only defect.

XPO3 can change the diffusion trajectory, so it is not pixel-identical or
lossless. The published 1280 BF16/XPO3 gallery is intended to make that trade
off inspectable rather than hidden.

## Sources and licenses

- Mage-Flow-Turbo, the VAE, and the Mage inference source derive from
  [`microsoft/Mage-Flow-Turbo`](https://huggingface.co/microsoft/Mage-Flow-Turbo)
  under the upstream MIT terms.
- The scaled-FP8 Qwen3-VL checkpoint comes from
  [`starsfriday/Mage-Flow-Edit-FP8`](https://huggingface.co/starsfriday/Mage-Flow-Edit-FP8)
  and derives from Apache-2.0 Qwen3-VL weights.
- Bundled SpargeAttn runtime code and extensions under
  `runtime/spas_sage_attn/` are Apache-2.0; see
  [`thu-ml/SpargeAttn`](https://github.com/thu-ml/SpargeAttn) and the plugin's
  `third_party_licenses/SpargeAttn-LICENSE`.
- XPO3 custom-node and packaging code is MIT licensed.

See the
[`ComfyUI-XPO3` third-party notices](https://github.com/AJH-Code/ComfyUI-XPO3/blob/main/THIRD_PARTY_NOTICES.md)
and included license texts in the custom-node repository.