Instructions to use Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP") config = load_config("Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Qwen3.8 Flash Next, MLX 2-bit with native MTP
A compact MLX affine conversion of Qwen/Qwen3.8-Flash-Next with sensitive language paths kept in BF16 and the model's own MTP block preserved.
Original model · Qwen overview · MLX-VLM · Qwen Community License 1.0
About this conversion
This is a non-sensitivity MLX conversion. It is not uniform Q2 and it is not an oQ build. Routed and shared experts plus the predictive n-gram embedding paths use affine 2-bit weights at group size 32. Attention, Gated DeltaNet, hyperconnection, routing, vision, token embedding, output head, and native MTP weights remain in BF16 or their source-compatible dtype.
| Item | Value |
|---|---|
| Repository | Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP |
| Base model | Qwen/Qwen3.8-Flash-Next |
| Source revision | f5d08274bafd880402bd16f5e3e6c514136ec06c |
| Source precision | Official BF16 checkpoint |
| Format | MLX safetensors |
| Quantisation | Affine 2-bit, group size 32, with explicit BF16 exclusions |
| 2-bit modules | 418 |
| Native MTP | Included, one matching Qwen4Exp draft block in BF16 |
| Indexed tensors | 2,543 total, including 32 native-MTP entries |
| Weight shards | 20 |
| Weight size | 80.070 GB / 74.571 GiB |
| Configured context | 262,144 tokens |
| Architecture | qwen4_exp vision-language sparse MoE |
The upstream tokenizer, current chat template, image and video processor configuration, generation configuration, licence, and native MTP configuration are included.
Use a runtime with explicit
qwen4_expand native-MTP support. This release was validated with oMLX 0.6.3rc3 build 2475, MLX 0.32.0, and MLX-VLM 0.6.3.
Download and use
python -m pip install --upgrade huggingface_hub
hf download Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP \
--local-dir ./Qwen3.8-Flash-Next-MLX-2bit-MTP
Add the downloaded directory to a compatible oMLX model directory and refresh the model registry. Native MTP is optional. Keep it disabled by default for this release because the measured native path was slightly slower than baseline.
Apple M3 Studio performance
The validation used three measured 512-token runs per mode after warm-up. The table reports median generation throughput from the exact release checkpoint.
| Runtime mode | Runs | Output per run | Median generation speed |
|---|---|---|---|
| Native MTP disabled | 3 | 512 tokens | 29.3729 tokens/s |
| Native MTP enabled | 3 | 512 tokens | 28.6199 tokens/s |
Native MTP changed median throughput by -2.56% in this test. The MTP telemetry sample accepted 9 of 17 reported draft proposals, an acceptance rate of 52.94%. Every MTP-off and MTP-on 512-token run produced the same output hash.
Instruction following, factual recall, arithmetic, concise response, and coherent-generation gates passed. Native MTP worked correctly and preserved greedy output, but it did not improve throughput on this checkpoint and runtime. The recommended default is MTP disabled.
The benchmark covers text generation. Results vary with prompt length, context growth, cache state, runtime version, memory pressure, and thermal conditions.
Architecture
Qwen3.8 Flash Next combines Gated DeltaNet, Qwen Sparse Attention, sparse mixture-of-experts layers, widened gated residual streams, hashed bigram and trigram embeddings, and a native next-token-prediction block for speculative decoding.
| Architecture detail | Upstream value |
|---|---|
| Language-model parameters | 125B total / 6B active |
| N-gram embedding | 51B parameters, 20,000,000 entries |
| Native MTP | 4B parameters, one draft layer |
| Hidden size | 2,560 |
| Layers | 48 |
| Routed / active experts | 512 / 10, plus 1 shared expert |
| Native context | 262,144 tokens, extensible upstream to 1,000,000 |
For upstream evaluations, intended use, safety guidance, and the full architecture discussion, see the original model card.
Conversion and validation
- Converted directly from the official BF16 checkpoint.
- Quantised 418 expert and predictive-embedding modules to affine Q2 at group size 32.
- Preserved 762 sensitive or structural matrix entries in BF16, including the complete matching native MTP block.
- Verified all 2,543 indexed tensors and all 20 shards before upload.
- Loaded the checkpoint in oMLX and passed deterministic instruction, factual, arithmetic, concise-writing, and coherent-generation tests.
- Ran three 512-token measurements in each MTP mode with exact paired output parity.
This is a community conversion, not an official Qwen release.
Limitations
- The 2-bit expert allocation is aggressive. Evaluate accuracy and visual understanding on the intended workload before deployment.
- Native MTP was 2.56% slower in the measured test. Acceptance alone does not guarantee a speedup.
- The configured 262,144-token context does not guarantee that every Apple-silicon system can run that length within available unified memory.
- Text generation was benchmarked. The vision stack loaded successfully, but visual quality was not benchmarked for this card.
- This is an MLX checkpoint for Apple silicon. It is not a GGUF, CUDA, TensorRT-LLM, or vLLM checkpoint.
- Upstream model limitations and safety considerations still apply.
Licence and attribution
The upstream model is released under the Qwen Community License 1.0. The required licence text is included in this repository and should be reviewed before use or redistribution.
Model design, training, evaluations, and upstream documentation belong to Qwen and the original contributors. The MLX conversion, native-MTP preservation, Apple-silicon validation, and packaging are provided by Vontra.
- Downloads last month
- -
2-bit
Model tree for Vontra/Qwen3.8-Flash-Next-MLX-2bit-MTP
Base model
Qwen/Qwen3.8-Flash-Next