SUPPORT COMING IN MTPLX v2.10. This pack needs MTPLX 2.10, which is about to release. The current MTPLX 2.9.x cannot serve it yet. Update to 2.10 when it lands and this model works out of the box in the app and the CLI.

MTPLX.COM: 2 to 3x speedup. The fastest way to run models on a Mac.

Qwen 3.8 Flash-Next Optimized Speed

Dynamic 4-bit quant with 8-bit attention. Higher quality and slightly slower. Recommended.

Qwen's 125B-A6B Flash-Next preview — the Qwen4-generation architecture with GDN hybrid MoE, Qwen Sparse Attention, and the 51B-parameter n-gram memory — running natively on MTPLX from day 0, with its native multi-token-prediction head drafting through MTPLX's speculative lane. This is the recommended build: the Qwen Sparse Attention projections are kept at 8-bit, so the attention pathway that steers long contexts keeps its precision. For the absolute fastest build, pick Bare Speed.

The 32 GB n-gram embedding table streams from SSD by default, so the model fits a 96 GB+ Apple Silicon Mac with headroom — the weights stay resident, the table does not have to.

Speeds

Measured on an M5 Max, fans verified at max, single stream, real server (mtplx serve), official Qwen 3.8 sampling (temperature 1.0, top-p 0.95, top-k 20 — sampled, not greedy).

Run tok/s
Coding task, MTP speculative decode (the default) 73.5
Same task, plain autoregressive 43.8

That is a 1.7x speculative multiplier through the product serve path, on sampled output that follows the model's own distribution.

How it is built

  • MoE experts and dense matrices at 4-bit with 64-weight groups; the Qwen Sparse Attention projections promoted to 8-bit (the quality edge over Bare Speed).
  • The GDN convolution and recurrent-state parameters, every norm, the QSA indexer, and the MTP head stay 16-bit.
  • The n-gram embedding table ships as a separate ngram-table.safetensors sidecar that MTPLX streams from SSD (resident is opt-in on very large machines). The vision tower is preserved in the weights.
Download 115.1 GB (includes the 32 GB n-gram table)
Resident weights (n-gram on SSD) ~83 GB + working set
Recommended Macs 96 GB+ unified memory
Context window 262,144 tokens
MTP depth adaptive, ceiling 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

The serving contract ships inside mtplx_runtime.json. MTPLX reads it on load. Drafts are accepted with the probability-ratio rule plus residual resampling, so the output follows the model's own distribution at any temperature.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 Flash-Next Optimized Speed".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed

Sibling: Bare Speed (flat 4-bit — the quickest build).

Base model: Qwen/Qwen3.8-Flash-Next (Qwen Community License; the upstream model card is preserved in this repo as README-upstream-qwen.md).

Downloads last month
1,032
Safetensors
Model size
25B params
Tensor type
BF16
·
U32
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Youssofal/Qwen3.8-Flash-Next-MTPLX-Optimized-Speed

Quantized
(123)
this model