---
license: apache-2.0
license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE
thumbnail: https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/hero.png
base_model:
- Qwen/Qwen3.5-9B
base_model_relation: quantized
quantized_by: AtomicChat
pipeline_tag: text-generation
library_name: mlx
tags:
- atomic-chat
- qwen3.5
- qwen
- mlx
- apple-silicon
- quantized
---
**Qwen3.5 9B**, self-quantized to MLX by [Atomic Chat](https://atomic.chat). Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline.
## Highlights
- **9.7B parameters**: the weights this repo quantizes.
- **Context length**: 262,144 tokens (256K), as published by Qwen.
- **32 layers**: Dense decoder.
- **Modalities**: Text, Image.
- **Full imatrix ladder**: every quant is calibrated with an importance matrix.
- **Unified Vision-Language Foundation**: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks.
- **Efficient Hybrid Architecture**: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead.
> [!NOTE]
> These MLXs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model.
## Model Overview
| Property | Value |
|---|---|
| Base model | `Qwen/Qwen3.5-9B` |
| Parameters | 9.7B |
| Layers | 32 |
| Context length | 262,144 tokens (256K) |
| Vocabulary | 248,320 |
| Modalities | Text, Image |
| Architecture | Dense decoder, 16 attention heads over 4 KV heads, `Qwen3_5ForConditionalGeneration` |
| This repo | MLX weights |
## Get started
- **[Atomic Chat](https://atomic.chat):** search `AtomicChat/Qwen3.5-9B-MLX-4bit` and hit **Use this model**.
- **mlx-lm:** `mlx_lm.generate --model AtomicChat/Qwen3.5-9B-MLX-4bit --prompt "Hello" --max-tokens 512`
- **Server:** `mlx_lm.server --model AtomicChat/Qwen3.5-9B-MLX-4bit --port 8080`
## Best practices
| Parameter | Value |
|---|---|
| temperature | 1.0 |
| top_p | 0.95 |
| top_k | 20 |
| min_p | 0.0 |
| repetition_penalty | 1.0 |
Qwen's recommended sampling configuration for `Qwen/Qwen3.5-9B`.
## How these were made
1. Download `Qwen/Qwen3.5-9B` (original weights).
2. Convert and quantize with `mlx_lm.convert` on our pipeline.
## License
Original model by Qwen, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE). Quantized by Atomic Chat.