mlx-community/LFM2.5-1.2B-Instruct-OptiQ-4bit

Built with mlx-optiq, the MLX-native toolkit to quantize, fine-tune, and serve LLMs locally on Apple Silicon, no PyTorch and no cloud. All OptiQ quants · Docs

An OptiQ mixed-precision quant of LiquidAI/LFM2.5-1.2B-Instruct. 825 MB on disk, down from 2.34 GB at bf16, and it scores 79% on IFEval and 70% on GSM8K at that size.

LFM2.5 is a hybrid architecture: convolutional blocks interleaved with full attention, so only a few blocks carry a KV cache. OptiQ measures each layer's sensitivity and assigns per-layer bit-widths, keeping 8-bit where it matters and 4-bit elsewhere.

What it is

Property Value
Base LiquidAI/LFM2.5-1.2B-Instruct
Architecture lfm2 — hybrid conv + full attention
Method OptiQ mixed-precision, sensitivity-driven (bf16 reference)
On disk 825 MB (bf16: 2.34 GB)
Context 128k
KV cache mixed-precision kv_config.json bundled

Capability Score

Six-metric mean, the standard OptiQ eval.

Metric Score
MMLU (5-shot, 969 samples) 43.3%
GSM8K (1000 samples) 69.7%
IFEval (full set, strict) 79.1%
BFCL-V3 simple (200 calls) 45.0%
HumanEval (164 problems, pass@1) 48.8%
HashHop (long-context retrieval) 1.0%
Capability Score (mean of 6) 47.82

Instruction-following and grade-school maths are the strengths here, and it writes usable Python at nearly 50% pass@1 for a model under a gigabyte. Long-context multi-hop retrieval is the weak spot.

Run it

pip install mlx-optiq
from mlx_lm import load, generate

model, tok = load("mlx-community/LFM2.5-1.2B-Instruct-OptiQ-4bit")
prompt = tok.apply_chat_template(
    [{"role": "user", "content": "Write a haiku about compilers."}],
    tokenize=False, add_generation_prompt=True,
)
print(generate(model, tok, prompt=prompt, max_tokens=300))

Serve it with the bundled mixed-precision KV cache:

optiq serve --model mlx-community/LFM2.5-1.2B-Instruct-OptiQ-4bit --kv-config kv_config.json

Tool calling

LFM2.5 writes Pythonic calls between special tokens:

<|tool_call_start|>[get_weather(city="Paris")]<|tool_call_end|>

Pass tools= to apply_chat_template and they are rendered into the system prompt.

Links

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mlx-community/LFM2.5-1.2B-Instruct-OptiQ-4bit

Quantized
(74)
this model