Macaw-OptiQ-4bit / README.md
codelion's picture
Upload folder using huggingface_hub
21bd610 verified
|
Raw
History Blame Contribute Delete
2.34 kB
---
language:
- en
license: other
license_name: lfm-open-license-v1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-2.6B/raw/main/LICENSE
library_name: mlx
pipeline_tag: text-generation
base_model: badtheorylabs/Macaw
tags:
- mlx
- optiq
- quantized
- 4bit
- mixed-precision
- lfm2
- agent
- tool-calling
- on-device
- apple-silicon
---
# mlx-community/Macaw-OptiQ-4bit
> **Built with [mlx-optiq](https://mlx-optiq.com)**, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. [All OptiQ quants](https://mlx-optiq.com/models) · [Docs](https://mlx-optiq.com/docs/) · [LFM2.5 family](https://mlx-optiq.com/docs/lfm2.5)
An OptiQ mixed-precision quant of [badtheorylabs/Macaw](https://huggingface.co/badtheorylabs/Macaw), an on-device Mac assistant built on LFM2.5-2.6B. 1.93 GB on disk, down from 5.2 GB at bf16.
Macaw is a tool-calling agent, so the quant is aimed at keeping tool calls well formed rather than at raw benchmark scores.
## What it is
| Property | Value |
|---|---|
| Base | [badtheorylabs/Macaw](https://huggingface.co/badtheorylabs/Macaw) (LFM2.5-2.6B derivative) |
| Architecture | `lfm2` — hybrid conv + full attention, 30 layers |
| Method | OptiQ mixed-precision, per-layer bit allocation reused from the base family |
| On disk | 1.93 GB (bf16: 5.2 GB) |
| Context | 128k |
Macaw keeps its base architecture, so which layers tolerate fewer bits is unchanged. The per-layer allocation comes from [LFM2.5-2.6B-OptiQ-4bit](https://huggingface.co/mlx-community/LFM2.5-2.6B-OptiQ-4bit) rather than a fresh sensitivity sweep: 167 of 167 layers matched, 80 kept at 8-bit and 87 at 4-bit.
## Run it
```bash
pip install mlx-optiq
optiq serve --model mlx-community/Macaw-OptiQ-4bit
```
That gives you an OpenAI and Anthropic compatible endpoint with mixed-precision KV cache, tool-call healing and prompt caching. The base model's recommended sampling ships in `generation_config.json` and `optiq serve` applies it without any flags.
## Links
- **Project website:** [mlx-optiq.com](https://mlx-optiq.com/)
- **All OptiQ quants:** [mlx-optiq.com/models](https://mlx-optiq.com/models)
- **PyPI:** [pypi.org/project/mlx-optiq](https://pypi.org/project/mlx-optiq/)
- **Base model:** [badtheorylabs/Macaw](https://huggingface.co/badtheorylabs/Macaw)