Qwen3.5-9B-MLX-4bit / README.md
worthant's picture
forge: regenerate the model card
458dedc verified
|
Raw
History Blame Contribute Delete
3.91 kB
---
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
thumbnail: https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/hero.png
base_model:
- Qwen/Qwen3.5-9B
base_model_relation: quantized
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: mlx
tags:
- atomic-chat
- qwen3.5
- qwen
- mlx
- apple-silicon
- quantized
---
<center>
<div style="display:flex; justify-content:center; align-items:center; gap:2%; max-width:560px; margin:0 auto;">
<a href="https://atomic.chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/pill_atomic_v3.png" alt="Atomic Chat" style="width:100%; height:auto; max-width:186px;"></a>
<a href="https://discord.gg/8wGSsvmg4V" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/pill_discord_v3.png" alt="Join Discord" style="width:100%; height:auto; max-width:184px;"></a>
<a href="https://github.com/AtomicBot-ai/Atomic-Chat" style="flex:0 1 auto; min-width:0;"><img src="https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/pill_github_v3.png" alt="GitHub" style="width:100%; height:auto; max-width:141px;"></a>
</div>
<br/>
<img src="https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/hero.png" alt="Qwen3.5 9B" style="width:100%; max-width:100%; height:auto; margin-bottom:0.6em;"/>
<div style="display:flex; justify-content:center; gap:0.5em;">
<a href="https://huggingface.co/Qwen/Qwen3.5-9B"><strong>Base model: Qwen/Qwen3.5-9B</strong></a>
</div>
</center>
**Qwen3.5 9B**, self-quantized to MLX by [Atomic Chat](https://atomic.chat). Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
## Highlights
- **9.7B parameters**: the weights this repo quantizes.
- **Context length**: 262,144 tokens (256K), as published by Qwen.
- **32 layers**: Dense decoder.
- **Modalities**: Text, Image.
- **Full imatrix ladder**: every quant is calibrated with an importance matrix.
- **Unified Vision-Language Foundation**: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
- **Efficient Hybrid Architecture**: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.
> [!NOTE]
> These MLXs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
## Model Overview
| Property | Value |
|---|---|
| Base model | `Qwen/Qwen3.5-9B` |
| Parameters | 9.7B |
| Layers | 32 |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 248,320 |
| Modalities | Text, Image |
| Architecture | Dense decoder, 16 attention heads over 4 KV heads, `Qwen3_5ForConditionalGeneration` |
| This repo | MLX weights |
## Get started
- **[Atomic Chat](https://atomic.chat):** search `AtomicChat/Qwen3.5-9B-MLX-4bit` and hit **Use this model**.
- **mlx-lm:** `mlx_lm.generate --model AtomicChat/Qwen3.5-9B-MLX-4bit --prompt "Hello" --max-tokens 512`
- **Server:** `mlx_lm.server --model AtomicChat/Qwen3.5-9B-MLX-4bit --port 8080`
## Best practices
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 0.95 |
| top_k | 20 |
| min_p | 0.0 |
| repetition_penalty | 1.0 |
Qwen's recommended sampling configuration for `Qwen/Qwen3.5-9B`.
## How these were made
1. Download `Qwen/Qwen3.5-9B` (original weights).
2. Convert and quantize with `mlx_lm.convert` on our pipeline.
## License
Original model by Qwen, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE). Quantized by Atomic Chat.