--- license: apache-2.0 license_link: https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE thumbnail: https://huggingface.co/AtomicChat/Qwen3.5-9B-MLX-4bit/resolve/main/hero.png base_model: - Qwen/Qwen3.5-9B base_model_relation: quantized quantized_by: AtomicChat pipeline_tag: text-generation library_name: mlx tags: - atomic-chat - qwen3.5 - qwen - mlx - apple-silicon - quantized ---
Atomic Chat Join Discord GitHub

Qwen3.5 9B
Base model: Qwen/Qwen3.5-9B
**Qwen3.5 9B**, self-quantized to MLX by [Atomic Chat](https://atomic.chat). Built straight from Qwen's original weights with a per-tensor importance matrix, so this is not a repack of somebody else's files. Runs fully offline. ## Highlights - **9.7B parameters**: the weights this repo quantizes. - **Context length**: 262,144 tokens (256K), as published by Qwen. - **32 layers**: Dense decoder. - **Modalities**: Text, Image. - **Full imatrix ladder**: every quant is calibrated with an importance matrix. - **Unified Vision-Language Foundation**: Early fusion training on multimodal tokens achieves cross-generational parity with Qwen3 and outperforms Qwen3-VL models across reasoning, coding, agents, and visual understanding benchmarks. - **Efficient Hybrid Architecture**: Gated Delta Networks combined with sparse Mixture-of-Experts deliver high-throughput inference with minimal latency and cost overhead. > [!NOTE] > These MLXs are **self-quantized from the original weights**, not a repack. The importance matrix keeps low-bit quants closer to the full-precision model. ## Model Overview | Property | Value | |---|---| | Base model | `Qwen/Qwen3.5-9B` | | Parameters | 9.7B | | Layers | 32 | | Context length | 262,144 tokens (256K) | | Vocabulary | 248,320 | | Modalities | Text, Image | | Architecture | Dense decoder, 16 attention heads over 4 KV heads, `Qwen3_5ForConditionalGeneration` | | This repo | MLX weights | ## Get started - **[Atomic Chat](https://atomic.chat):** search `AtomicChat/Qwen3.5-9B-MLX-4bit` and hit **Use this model**. - **mlx-lm:** `mlx_lm.generate --model AtomicChat/Qwen3.5-9B-MLX-4bit --prompt "Hello" --max-tokens 512` - **Server:** `mlx_lm.server --model AtomicChat/Qwen3.5-9B-MLX-4bit --port 8080` ## Best practices | Parameter | Value | |---|---| | temperature | 1.0 | | top_p | 0.95 | | top_k | 20 | | min_p | 0.0 | | repetition_penalty | 1.0 | Qwen's recommended sampling configuration for `Qwen/Qwen3.5-9B`. ## How these were made 1. Download `Qwen/Qwen3.5-9B` (original weights). 2. Convert and quantize with `mlx_lm.convert` on our pipeline. ## License Original model by Qwen, released under the Apache 2.0 license. Full terms: [Apache 2.0](https://huggingface.co/Qwen/Qwen3.5-9B/blob/main/LICENSE). Quantized by Atomic Chat.