codelion's picture
Add files using upload-large-folder tool
267d14f verified
|
Raw
History Blame Contribute Delete
1.7 kB
---
license: mit
language:
- en
tags:
- mlx
- optiq
- computer-use
- web-agent
- agent
library_name: mlx
pipeline_tag: image-text-to-text
base_model: microsoft/Fara1.5-4B
---
# Fara1.5-4B-OptiQ-4bit
> **Built with [mlx-optiq](https://mlx-optiq.com)**, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon (no PyTorch, no cloud). [Try the Lab](https://mlx-optiq.com/docs/lab/) 路 [All OptiQ quants](https://mlx-optiq.com/models) 路 [Docs](https://mlx-optiq.com/docs/)
>
> **Supported loaders:** [mlx-optiq](https://mlx-optiq.com) (text, vision, and MTP) and stock [mlx-lm](https://github.com/ml-explore/mlx-lm) (text). Other front-ends load MLX weights through their own stack, so support there depends on that stack rather than on these files.
An [OptiQ](https://mlx-optiq.com) mixed-precision MLX quant of [microsoft/Fara1.5-4B](https://huggingface.co/microsoft/Fara1.5-4B), a Qwen3.5-based computer-use / web-agent vision-language model.
- **Mixed 4/8-bit**, 5.31 bits per weight (3.9G on disk).
- The per-layer bit allocation is **transferred from the published `mlx-community/Qwen3.5-4B-OptiQ-4bit`** quant. Fara1.5 is a finetune of Qwen3.5-4B with identical architecture, so the OptiQ allocation matches the Qwen3.5 family exactly, with no separate sensitivity pass.
- **Vision tower kept at bf16** in `optiq/optiq_vision.safetensors`. The one repo loads text-only under stock `mlx-lm` and full image+text under OptiQ.
## Running it
```bash
pip install -U optiq
optiq serve --model mlx-community/Fara1.5-4B-OptiQ-4bit
```
Use the OpenAI-compatible endpoint at `http://localhost:8000/v1`. Send an `image_url` part for the computer-use / vision path.