MTPLX.COM: 2 to 3x speedup. The fastest way to run models on a Mac.

Qwen 3.8 27B Optimized Speed FP16

4-bit dynamic quant. Great coding speeds and good quality. Recommended. This is the M1 and M2 build.

The FP16 precision sibling of Qwen 3.8 27B Optimized Speed. M1 and M2 Macs do not run bf16 well, so this build keeps every quantized weight byte-identical to the parent and stores the remaining floating tensors (scales, biases, norms, the GDN convolution and state parameters, and the MTP head) in fp16 instead of bf16. Same layout, same tuned depth and draft settings, same context window. On an M3 or newer Mac use the parent instead.

MTPLX picks the right one for you: the app and mtplx start route M1 and M2 Macs to the FP16 builds and everything newer to the parents. It is the default MTPLX picks on an M1 or M2 Mac with 32 GB or more.

Speeds

The numbers we publish for the parent were measured on an M5 Max: 58.7 tok/s on the coding task and 35.1 to 37.3 tok/s on long xhigh reasoning, official Qwen 3.8 sampling, generation running to the model's own stop. This FP16 build has the same weights and runs the same MTPLX turbo path, so the speculative math is identical; absolute tok/s on an M1 or M2 depends on that chip. We have not published M1 or M2 numbers for it yet.

Download 20.4 GB
Peak unified memory (parent, measured on M5 Max) 23.6 GB
Context window 262,144 tokens
MTP depth 3
Sampling temperature 1.0, top-p 0.95, top-k 20 (the official Qwen 3.8 contract)

MTPLX_FP16_CONVERSION_MANIFEST.json in the repo lists every tensor that was cast and every tensor that was preserved, with sha256 for each shard. Speculation in MTPLX is exact at any temperature: drafts are accepted with the probability-ratio rule plus residual resampling.

Use it

Mac app: download at mtplx.com, pick "Qwen 3.8 27B Optimized Speed FP16".

Command line:

pip install mtplx
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed-FP16

The other two FP16 builds: Bare Speed FP16, Optimized Quality FP16.

Downloads last month
-
Safetensors
Model size
6B params
Tensor type
F16
U32
BF16
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed-FP16

Base model

Qwen/Qwen3.8-27B
Quantized
(399)
this model