mlx-community/LFM2.5-350M-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

An OptiQ mixed-precision quant of LiquidAI/LFM2.5-350M, a 350M on-device model from Liquid AI. 269 MB on disk, down from 709 MB at bf16.

LFM2.5 is a hybrid architecture: 16 blocks alternating short convolutions with full attention (only 6 blocks carry a KV cache). OptiQ measures each layer's sensitivity and assigns per-layer bit-widths, so the layers that matter keep 8-bit while the rest go to 4-bit.

What it is

Property Value
Base LiquidAI/LFM2.5-350M
Architecture lfm2 — 16 blocks, 10 conv + 6 full-attention
Method OptiQ mixed-precision, sensitivity-driven (bf16 reference)
On disk 269 MB (bf16: 709 MB)
Context 128k
KV cache mixed-precision kv_config.json bundled, covering the 6 attention layers

Capability Score

Six-metric mean, the standard OptiQ eval.

Metric Score
MMLU (5-shot, 969 samples) 30.3%
GSM8K (1000 samples) 2.5%
IFEval (full set, strict) 69.5%
BFCL-V3 simple (200 calls) 42.0%
HumanEval (164 problems, pass@1) 15.2%
HashHop (long-context retrieval) 0.0%
Capability Score (mean of 6) 26.60

A 350M model sits near the floor on multi-step maths and long-context retrieval; those scores are genuine, not harness artifacts. Instruction-following and tool-calling are where a model this size is actually useful.

Run it

pip install mlx-optiq
from mlx_lm import load, generate

model, tok = load("mlx-community/LFM2.5-350M-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Name three colours."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=200))

Serve it with the bundled mixed-precision KV cache:

optiq serve --model mlx-community/LFM2.5-350M-OptiQ-4bit --kv-config kv_config.json

Tool calling

LFM2.5 writes Pythonic calls between special tokens:

<|tool_call_start|>[get_weather(city="Paris")]<|tool_call_end|>

Pass tools= to apply_chat_template and they are rendered into the system prompt.

Links

Downloads last month
-
Safetensors
Model size
76M params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/LFM2.5-350M-OptiQ-4bit

Quantized
(51)
this model