Qwen

Qwen Qwen3.8 Apple silicon MLX Native MTP included Vontra 2-bit

Qwen3.8 Flash Next, MLX 2-bit with native MTP

A compact MLX affine conversion of Qwen/Qwen3.8-Flash-Next with sensitive language paths kept in BF16 and the model's own MTP block preserved.

Original model · Qwen overview · MLX-VLM · Qwen Community License 1.0

About this conversion

This is a non-sensitivity MLX conversion. It is not uniform Q2 and it is not an oQ build. Routed and shared experts plus the predictive n-gram embedding paths use affine 2-bit weights at group size 32. Attention, Gated DeltaNet, hyperconnection, routing, vision, token embedding, output head, and native MTP weights remain in BF16 or their source-compatible dtype.

Item Value
Repository Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP
Base model Qwen/Qwen3.8-Flash-Next
Source revision f5d08274bafd880402bd16f5e3e6c514136ec06c
Source precision Official BF16 checkpoint
Format MLX safetensors
Quantisation Affine 2-bit, group size 32, with explicit BF16 exclusions
2-bit modules 418
Native MTP Included, one matching Qwen4Exp draft block in BF16
Indexed tensors 2,543 total, including 32 native-MTP entries
Weight shards 20
Weight size 80.070 GB / 74.571 GiB
Configured context 262,144 tokens
Architecture qwen4_exp vision-language sparse MoE

The upstream tokenizer, current chat template, image and video processor configuration, generation configuration, licence, and native MTP configuration are included.

Use a runtime with explicit qwen4_exp and native-MTP support. This release was validated with oMLX 0.6.3rc3 build 2475, MLX 0.32.0, and MLX-VLM 0.6.3.

Download and use

python -m pip install --upgrade huggingface_hub

hf download Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP \
  --local-dir ./Qwen3.8-Flash-Next-MLX-2bit-MTP

Add the downloaded directory to a compatible oMLX model directory and refresh the model registry. Native MTP is optional. Keep it disabled by default for this release because the measured native path was slightly slower than baseline.

Apple M3 Studio performance

The validation used three measured 512-token runs per mode after warm-up. The table reports median generation throughput from the exact release checkpoint.

Runtime mode Runs Output per run Median generation speed
Native MTP disabled 3 512 tokens 29.3729 tokens/s
Native MTP enabled 3 512 tokens 28.6199 tokens/s

Native MTP changed median throughput by -2.56% in this test. The MTP telemetry sample accepted 9 of 17 reported draft proposals, an acceptance rate of 52.94%. Every MTP-off and MTP-on 512-token run produced the same output hash.

Instruction following, factual recall, arithmetic, concise response, and coherent-generation gates passed. Native MTP worked correctly and preserved greedy output, but it did not improve throughput on this checkpoint and runtime. The recommended default is MTP disabled.

The benchmark covers text generation. Results vary with prompt length, context growth, cache state, runtime version, memory pressure, and thermal conditions.

Architecture

Qwen3.8 Flash Next combines Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, hashed bigram and trigram embeddings, and a native next-token-prediction block for speculative decoding.

Architecture detail Upstream value
Language-model parameters 125B total / 6B active
N-gram embedding 51B parameters, 20,000,000 entries
Native MTP 4B parameters, one draft layer
Hidden size 2,560
Layers 48
Routed / active experts 512 / 10, plus 1 shared expert
Native context 262,144 tokens, extensible upstream to 1,000,000

For upstream evaluations, intended use, safety guidance, and the full architecture discussion, see the original model card.

Conversion and validation

  • Converted directly from the official BF16 checkpoint.
  • Quantised 418 expert and predictive-embedding modules to affine Q2 at group size 32.
  • Preserved 762 sensitive or structural matrix entries in BF16, including the complete matching native MTP block.
  • Verified all 2,543 indexed tensors and all 20 shards before upload.
  • Loaded the checkpoint in oMLX and passed deterministic instruction, factual, arithmetic, concise-writing, and coherent-generation tests.
  • Ran three 512-token measurements in each MTP mode with exact paired output parity.

This is a community conversion, not an official Qwen release.

Limitations

  • The 2-bit expert allocation is aggressive. Evaluate accuracy and visual understanding on the intended workload before deployment.
  • Native MTP was 2.56% slower in the measured test. Acceptance alone does not guarantee a speedup.
  • The configured 262,144-token context does not guarantee that every Apple-silicon system can run that length within available unified memory.
  • Text generation was benchmarked. The vision stack loaded successfully, but visual quality was not benchmarked for this card.
  • This is an MLX checkpoint for Apple silicon. It is not a GGUF, CUDA, TensorRT-LLM, or vLLM checkpoint.
  • Upstream model limitations and safety considerations still apply.

Licence and attribution

The upstream model is released under the Qwen Community License 1.0. The required licence text is included in this repository and should be reviewed before use or redistribution.

Model design, training, evaluations, and upstream documentation belong to Qwen and the original contributors. The MLX conversion, native-MTP preservation, Apple-silicon validation, and packaging are provided by Vontra.

Downloads last month
-
Safetensors
Model size
29B params
Tensor type
BF16
·
U32
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP

Quantized
(89)
this model