Fara1.5-27B-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the LabAll OptiQ quantsDocs

Supported loaders: mlx-optiq (text, vision, and MTP) and stock mlx-lm (text). Other front-ends load MLX weights through their own stack, so support there depends on that stack rather than on these files.

An OptiQ mixed-precision MLX quant of microsoft/Fara1.5-27B, a Qwen3.5-based computer-use / web-agent vision-language model.

  • Mixed 4/8-bit, 4.74 bits per weight (20G on disk).
  • The per-layer bit allocation is transferred from the published mlx-community/Qwen3.5-27B-OptiQ-4bit quant. Fara1.5 is a finetune of Qwen3.5-27B with identical architecture, so the OptiQ allocation matches the Qwen3.5 family exactly, with no separate sensitivity pass.
  • Vision tower kept at bf16 in optiq/optiq_vision.safetensors. The one repo loads text-only under stock mlx-lm and full image+text under OptiQ.

Running it

pip install -U optiq
optiq serve --model mlx-community/Fara1.5-27B-OptiQ-4bit

Use the OpenAI-compatible endpoint at http://localhost:8000/v1. Send an image_url part for the computer-use / vision path.

Downloads last month
83
Safetensors
Model size
27B params
Tensor type
F32
U32
BF16
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for mlx-community/Fara1.5-27B-OptiQ-4bit

Base model

Qwen/Qwen3.5-27B
Quantized
(5)
this model