mlx-community/Ornith-1.5-9B-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of ornith-ai/Ornith-1.5-9B-MLX, a 9B reasoning model. 7 GB on disk.

What it is

Property Value
Base ornith-ai/Ornith-1.5-9B-MLX (Qwen3.5, 9B, 32 layers)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the Ornith-1.0-9B OptiQ recipe: same architecture and layer count, so the per-layer sensitivity ranking transfers directly and no per-model sweep is needed
Layer split 116 components at 4-bit, 134 at 8-bit
Group size 64
On disk 7 GB

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

pip install "mlx-optiq>=0.4.28"
import optiq  # registers the arch
from mlx_lm import load, generate

model, tok = load("mlx-community/Ornith-1.5-9B-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For an OpenAI- and Anthropic-compatible endpoint with mixed-precision KV cache:

optiq serve --model mlx-community/Ornith-1.5-9B-OptiQ-4bit

This is a reasoning model, so give it a generous token budget.

Links

Downloads last month
45
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Ornith-1.5-9B-OptiQ-4bit

Quantized
(2)
this model