Qwen3.8-27B ternary (MLX 2-bit)

Ternary (1.58-bit) Qwen3.8-27B, quantization-aware distilled, packed into MLX's native affine 2-bit container. Hybrid architecture: 48 gated-delta linear-attention layers and 16 softmax-attention layers. Quantized: in_proj_qkv, in_proj_z, out_proj, q/k/v/o_proj, mlp.*, and the vision tower. Full precision: in_proj_a, in_proj_b, conv1d, A_log, dt_bias, all norms, embed_tokens, lm_head.

Format

MLX native affine quantization, no custom kernel and no runtime shim:

bits        2
group_size  128
mode        affine
levels      {0, 1, 2}        level 3 is unused
bias        == -scale        so dequantisation is scale * (q - 1) = {-a, 0, +a}

There is no rotation anywhere in this model, so there is no signs tensor and no Hadamard transform to apply at load time.

Load

from mlx_lm import load
model, tokenizer = load("<repo>")

The container is stock MLX, but the architecture still has to be implemented in your mlx-lm / mlx-vlm build for the full model to load.

Downloads last month
-
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support