mlx-community/LFM2.5-230M-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

An OptiQ mixed-precision quant of LiquidAI/LFM2.5-230M, the smallest model in the LFM2.5 line. 180 MB on disk, down from 459 MB at bf16, small enough to sit in memory on any Apple Silicon Mac alongside everything else you are running.

LFM2.5 is a hybrid architecture: convolutional blocks interleaved with full attention, so only a few blocks carry a KV cache. OptiQ measures each layer's sensitivity and assigns per-layer bit-widths.

What it is

Property Value
Base LiquidAI/LFM2.5-230M
Architecture lfm2 — hybrid conv + full attention
Method OptiQ mixed-precision, sensitivity-driven (bf16 reference)
On disk 180 MB (bf16: 459 MB)
Context 128k
KV cache mixed-precision kv_config.json bundled

Capability Score

Six-metric mean, the standard OptiQ eval.

Metric Score
MMLU (5-shot, 969 samples) 34.2%
GSM8K (1000 samples) 26.0%
IFEval (full set, strict) 60.4%
BFCL-V3 simple (200 calls) 18.0%
HumanEval (164 problems, pass@1) 10.4%
HashHop (long-context retrieval) 0.0%
Capability Score (mean of 6) 24.83

At 230M parameters the strengths are instruction-following and maths for the size. Long-context multi-hop retrieval sits at the floor for a model this small, and that score is genuine rather than a harness artifact.

Run it

pip install mlx-optiq
from mlx_lm import load, generate

model, tok = load("mlx-community/LFM2.5-230M-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Name three colours."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=200))

Serve it with the bundled mixed-precision KV cache:

optiq serve --model mlx-community/LFM2.5-230M-OptiQ-4bit --kv-config kv_config.json

Tool calling

LFM2.5 writes Pythonic calls between special tokens:

<|tool_call_start|>[get_weather(city="Paris")]<|tool_call_end|>

Pass tools= to apply_chat_template and they are rendered into the system prompt.

Links

Downloads last month
-
Safetensors
Model size
50.7M params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/LFM2.5-230M-OptiQ-4bit

Quantized
(23)
this model