mlx-community/Qwythos-27B-v1-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

OptiQ mixed-precision quant of empero-ai/Qwythos-27B-v1, a Qwen3.5-family vision-language model (image + text, long context, with a bundled MTP speculation head). 19 GB on disk.

What it is

Property Value
Base empero-ai/Qwythos-27B-v1 (Qwen3.5-VL, 64-layer hybrid linear + full attention)
Method OptiQ mixed-precision, per-layer 4/8-bit
Bit allocation Reused from the base Qwen3.5-27B OptiQ recipe: the architecture is identical, so the per-layer sensitivity ranking transfers directly and no per-model sensitivity sweep is needed
Layer split 279 layers at 4-bit, 219 at 8-bit
Achieved bits-per-weight 5.55
On disk 19 GB
MTP Speculation head preserved in optiq/mtp.safetensors for faster decode via optiq serve --mtp
Vision bf16 vision tower kept in optiq/optiq_vision.safetensors for image and video-frame input

Following the naming llama.cpp uses for its mixed quants, the "4bit" label denotes the family, not the weighted average.

Run it

Qwen3.5 plus the MTP and vision sidecars register through OptiQ, so import optiq once before loading:

pip install "mlx-optiq>=0.4.7"
import optiq  # registers the arch + MTP/vision sidecars
from mlx_lm import load, generate

model, tok = load("mlx-community/Qwythos-27B-v1-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Explain mixed-precision quantization in two sentences."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=400))

For image and video input, plus an OpenAI + Anthropic-compatible endpoint with the MTP speculation head and mixed-precision KV cache:

optiq serve --model mlx-community/Qwythos-27B-v1-OptiQ-4bit

Qwythos is a reasoning model, so give it a generous token budget.

Links

Downloads last month
1
Safetensors
Model size
27B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Qwythos-27B-v1-OptiQ-4bit

Base model

Qwen/Qwen3.5-27B
Quantized
(12)
this model