Macaw-OptiQ-4bit / README.md
codelion's picture
Upload folder using huggingface_hub
21bd610 verified
|
Raw
History Blame Contribute Delete
2.34 kB
metadata
language:
  - en
license: other
license_name: lfm-open-license-v1.0
license_link: https://huggingface.co/LiquidAI/LFM2.5-2.6B/raw/main/LICENSE
library_name: mlx
pipeline_tag: text-generation
base_model: badtheorylabs/Macaw
tags:
  - mlx
  - optiq
  - quantized
  - 4bit
  - mixed-precision
  - lfm2
  - agent
  - tool-calling
  - on-device
  - apple-silicon

mlx-community/Macaw-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs · LFM2.5 family

An OptiQ mixed-precision quant of badtheorylabs/Macaw, an on-device Mac assistant built on LFM2.5-2.6B. 1.93 GB on disk, down from 5.2 GB at bf16.

Macaw is a tool-calling agent, so the quant is aimed at keeping tool calls well formed rather than at raw benchmark scores.

What it is

Property Value
Base badtheorylabs/Macaw (LFM2.5-2.6B derivative)
Architecture lfm2 — hybrid conv + full attention, 30 layers
Method OptiQ mixed-precision, per-layer bit allocation reused from the base family
On disk 1.93 GB (bf16: 5.2 GB)
Context 128k

Macaw keeps its base architecture, so which layers tolerate fewer bits is unchanged. The per-layer allocation comes from LFM2.5-2.6B-OptiQ-4bit rather than a fresh sensitivity sweep: 167 of 167 layers matched, 80 kept at 8-bit and 87 at 4-bit.

Run it

pip install mlx-optiq
optiq serve --model mlx-community/Macaw-OptiQ-4bit

That gives you an OpenAI and Anthropic compatible endpoint with mixed-precision KV cache, tool-call healing and prompt caching. The base model's recommended sampling ships in generation_config.json and optiq serve applies it without any flags.

Links