LFM2.5-2.6B-ONNX / README.md
mlabonne's picture
Quantize and tie q4f16 input embedding (#1)
6682637
|
Raw
History Blame Contribute Delete
2.97 kB
metadata
license: other
license_name: lfm1.0
license_link: LICENSE
language:
  - ar
  - zh
  - en
  - fr
  - de
  - hi
  - id
  - it
  - ja
  - ko
  - pl
  - pt
  - ru
  - es
  - th
  - vi
pipeline_tag: text-generation
tags:
  - liquid
  - edge
  - lfm2.5
  - onnx
  - onnxruntime
  - webgpu
base_model:
  - LiquidAI/LFM2.5-2.6B
Liquid AI
Try LFM β€’ Docs β€’ LEAP β€’ Discord

LFM2.5-2.6B-ONNX

LFM2.5 is a new family of hybrid models designed for on-device deployment. It builds on the LFM2 architecture with extended pre-training and reinforcement learning.

Find more details in the original model card: https://huggingface.co/LiquidAI/LFM2.5-2.6B

Recommended Variants

Precision Size Platform Use Case
Q4 ~1.9 GB WebGPU, Server Recommended for most uses (quantized embedding)
Q4F16 ~1.5 GB WebGPU Quantized embedding and q4 weights with FP16 runtime and caches
FP16 ~2.1 GB WebGPU, Server Higher quality
Q8 ~2.1 GB Server only Balance of quality and size
  • WebGPU: Use Q4, Q4F16, or FP16 (Q8 is not supported on WebGPU).
  • Server (CPU/GPU): All variants supported.

Q4 and Q4F16 use a quantized input embedding. Q4F16 uses FP16 runtime tensors and caches while quantizing the LM head and decoder linear weights to q4.

Model Files

onnx/
β”œβ”€β”€ model.onnx              # FP32
β”œβ”€β”€ model_fp16.onnx         # FP16
β”œβ”€β”€ model_q4.onnx           # Q4, quantized embedding (WebGPU)
β”œβ”€β”€ model_q4f16.onnx        # Q4 embedding/weights, FP16 runtime and caches (WebGPU)
└── model_q8.onnx           # Q8

Python (onnxruntime)

pip install onnxruntime transformers numpy huggingface_hub
# or, for GPU:
pip install onnxruntime-gpu transformers numpy huggingface_hub
from huggingface_hub import hf_hub_download

model_id = "LiquidAI/LFM2.5-2.6B-ONNX"
# Q8 recommended for server CPU/GPU; use model_q4.onnx for WebGPU.
hf_hub_download(model_id, "onnx/model_q8.onnx")
hf_hub_download(model_id, "onnx/model_q8.onnx_data")

WebGPU (Transformers.js)

import { pipeline } from "@huggingface/transformers";

const generator = await pipeline("text-generation", "LiquidAI/LFM2.5-2.6B-ONNX", {
  device: "webgpu",
  dtype: "q4", // or "q4f16" or "fp16"
});