How to use from the
Use from the
MLX library
# Make sure mlx-vlm is installed
# pip install --upgrade mlx-vlm

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
from mlx_vlm.utils import load_config

# Load the model
model, processor = load("mlx-community/Fara1.5-9B-OptiQ-4bit")
config = load_config("mlx-community/Fara1.5-9B-OptiQ-4bit")

# Prepare input
image = ["http://images.cocodataset.org/val2017/000000039769.jpg"]
prompt = "Describe this image."

# Apply chat template
formatted_prompt = apply_chat_template(
    processor, config, prompt, num_images=1
)

# Generate output
output = generate(model, processor, formatted_prompt, image)
print(output)

Fara1.5-9B-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). Try the Lab · All OptiQ quants · Docs

Supported loaders: mlx-optiq (text, vision, and MTP) and stock mlx-lm (text). Other front-ends load MLX weights through their own stack, so support there depends on that stack rather than on these files.

An OptiQ mixed-precision MLX quant of microsoft/Fara1.5-9B, a Qwen3.5-based computer-use / web-agent vision-language model.

  • Mixed 4/8-bit, 5.21 bits per weight (7.5G on disk).
  • The per-layer bit allocation is transferred from the published mlx-community/Qwen3.5-9B-OptiQ-4bit quant. Fara1.5 is a finetune of Qwen3.5-9B with identical architecture, so the OptiQ allocation matches the Qwen3.5 family exactly, with no separate sensitivity pass.
  • Vision tower kept at bf16 in optiq/optiq_vision.safetensors. The one repo loads text-only under stock mlx-lm and full image+text under OptiQ.

Running it

pip install -U optiq
optiq serve --model mlx-community/Fara1.5-9B-OptiQ-4bit

Use the OpenAI-compatible endpoint at http://localhost:8000/v1. Send an image_url part for the computer-use / vision path.

Downloads last month
149
Safetensors
Model size
9B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/Fara1.5-9B-OptiQ-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(11)
this model