cpu-only-inference-models / docs /lfm2-maple-4bit-report.md
ljupco's picture
Upload folder using huggingface_hub
816f936 verified
|
Raw
History Blame Contribute Delete
7.85 kB

LFM2.5 / Maple Preview β€” CPU-Efficient Quantized Variants and a Three-Engine Exploration

Author: Ljubomir Josifovski (ljupco) β€” Harpenden, UK, 2026. Hardware: thinkpad2 (Lenovo, Intel i5-8350U, 4C/8T, AVX2, 64 GB DDR4-2400) and macbook2 (M2 Max, 96 GB). Scope: quantized GGUF variants of the LiquidAI LFM2.5 family and the DeepGrove Maple Preview, plus a three-engine porting/benchmarking exploration: stock llama.cpp (with the DeepGrove fork for Maple), ik-llama.cpp, and vllm.cpp.


1. llama.cpp (stock + the DeepGrove fork) β€” the reference engines

1.1 Maple Preview on the DeepGrove fork (deepgrove-main)

The best CPU result of the whole project is the Maple Preview model (DeepGrove's ternary 2-bit MoE, ~1B active parameters of ~20B total, 256 experts with 8 routed) running on the DeepGrove fork of llama.cpp. The fork adds the maple architecture (src/models/maple.cpp), the TQ2_0 ternary quantization, and the GGUF conversion support. On top of that upstream work we added:

  • TQ2_0 ternary Metal kernels (ggml/src/ggml-metal/ggml-metal.metal) for the M2 Max, and
  • per-machine server wrapper scripts (llama_server_maple-preview_macbook2.sh, llama_server_maple-preview_thinkpad2.sh).

The Maple Preview GGUFs referenced here are the DeepGrove team's own releases; the engine is the DeepGrove fork. We made no code changes that need publishing for this model β€” the fork is upstream; we only reference it. Measured decode (benchy tg64, 4 threads, thinkpad2): ~28.2 t/s β€” the fastest model of the four.

1.2 The LFM2.5 quantized variants (this repository's contribution)

For the LiquidAI LFM2.5 family (2.6B dense, 1.2B-Thinking, 8B-A1B MoE) we produced mixed-quant GGUF variants that keep the most sensitive tensors (attention / router / output heads) at a higher precision while the bulk of the weights are 4-bit (Q4_K / Q4_0). The variants we distribute on HuggingFace (ljupco/*) are:

Model Variant Body Heads
LFM2.5-2.6B -Q4_K_M Q4_K Q4_K
LFM2.5-1.2B-Thinking -Q4_0h Q4_0 higher-precision heads
LFM2.5-8B-A1B -Q4_0h Q4_0 higher-precision heads

These are the files benchmarked in the engine sections below. Full credit and thanks to the original model creators β€” see the Credits section.


2. ik-llama.cpp β€” the LFM2 port (lfm2-port branch)

~/ik-llama.cpp carries a from-scratch LFM2 architecture port into a llama.cpp fork: the LFM2 2.6B dense and 1.2B-Thinking and 8B-A1B MoE, including the shortconv layers, the gated-delta attention, and the NEOX RoPE handling that the LFM2 family requires.

Status: correct and complete β€” all three LFM2 models run and produce valid output. Benchmarks vs the upstream llama.cpp reference (benchy tg64, 4 threads, thinkpad2):

Model ik-llama.cpp upstream llama.cpp delta
LFM2.5-2.6B-Q4_K_M 13.08 t/s 12.40 t/s +5 %
LFM2.5-1.2B-Thinking-Q4_0h 25.30 t/s 30.40 t/s βˆ’17 %
LFM2.5-8B-A1B-Q4_0h 13.81 t/s 20.99 t/s βˆ’34 %

The 2.6B dense shows a small win; the 1.2B and 8B trails. The port is the reference used for the vllm.cpp section. No separate code release is planned for this fork beyond the branch itself; it is a research port.


3. vllm.cpp β€” the LFM2/Maple port (lfm2-maple-port branch)

~/vllm.cpp/worktrees/lfm2-maple-port (fork: ljubomirj/vllm.cpp, branch lfm2-maple-port) is a port of all four models into the vllm.cpp engine (GGUF-first, C++20, single binary), with a large speed campaign. This work did not beat the llama.cpp references β€” it is published for reference and as a record of the optimisation arc.

3.1 The port

All four models run and produce correct output: LFM2.5 2.6B / 1.2B-Thinking / 8B-A1B and Maple Preview (general.architecture dispatch in src/vllm/entrypoints/model_loader.cpp).

3.2 The speed campaign β€” what was done

  1. The iqk tier β€” vendored the llama.cpp CPU K-quant integer-dot kernels into src/vt/cpu/iqk/ (Q4_K / Q6_K decode dots at ~72 GFLOP/s, DDR4-bandwidth-bound).
  2. Op fusion β€” the LFM2 layers run at ~4 ops vs ~9 originally:
    • the gate/up row-concat merge (per-row block quants make the concat a pure byte copy);
    • the silu+down fused kernel (kMatmulBf16DSilu);
    • the post-norm+gate_up fusion (kMatmulBf16DPostNorm);
    • the in-norm+in_proj and in-norm+first-qkv fusions (the input-layernorm with its residual-add folded into the act quantization);
    • the attention bf16-alignment fix (the fused q_proj's act row must match the bf16-rounded normed activation the sibling k/v projections consume).
  3. The R8 gemv arc β€” vendored the DeepGrove 8x8 LUT gemv (src/vt/cpu/gemv_q4k/). Conclusive negative result: the DeepGrove gemv is structurally tied to the DeepGrove Q4_K sub-block-pair packing (qs[j] = {value (j%32)+64*(j/32), value +32}) and its (2s, 2s+1) scale/min/bsum grouping, which is incompatible with the standard ggml Q4_K packing (sub-blocks (s, s+4)) used by all standard-quantized models. The gemv is correct only for DeepGrove-quantized weights. This is recorded here so the next engineer does not repeat the arc.
  4. Threadpool β€” the ggml-faithful barrier+steal pool is the local optimum; the poll=0 cond-wait experiment is slower than the spin for the fast decode ops.

3.3 Results (benchy tg64, 4 threads, thinkpad2)

Model vllm.cpp (best) llama.cpp reference delta
LFM2.5-2.6B-Q4_K_M 9.0 t/s 13.08 (ik) βˆ’31 %
LFM2.5-1.2B-Thinking-Q4_0h ~14 t/s 25.30 (ik) βˆ’45 %
LFM2.5-8B-A1B-Q4_0h 9.3 t/s 13.81 (ik) βˆ’33 %
Maple Preview 13.9 t/s 28.2 (DeepGrove) βˆ’51 %

The residual gap is the DDR4-bandwidth-bound dot compute itself (the dots are at the same speed as the reference; the overhead is at the floor after the fusions).


4. Publishing

4.1 HuggingFace β€” the quantized GGUFs

The mixed-quant LFM2.5 GGUFs are published under ljupco on HuggingFace. Each repository's model card declares the base_model metadata pointing at the original LiquidAI model, so HuggingFace lists our variants on the originals' Quantizations pages.

4.2 Code

  • Maple Preview: the engine is the DeepGrove fork (upstream); our additions are the TQ2_0 ternary Metal kernels (a non-trivial, reusable contribution β€” ~400+ lines of ggml-metal support for the ternary quantization that could be upstreamed to llama.cpp) and the per-machine wrapper scripts, which remain in the deepgrove-main worktree.
  • ik-llama.cpp: the lfm2-port branch remains the reference.
  • vllm.cpp: the lfm2-maple-port branch is pushed to the ljubomirj/vllm.cpp fork for reference.

5. Credits and Acknowledgements

We are deeply grateful to the teams whose work this project builds on:

  • Liquid AI β€” for the LFM2.5 family (2.6B, 1.2B-Thinking, 8B-A1B), the gated-delta / shortconv architecture, and their open weights. We thank them for making these models available.
  • DeepGrove AI β€” for the Maple Preview model, the TQ2_0 ternary quantization, and their llama.cpp fork with the Maple architecture support. Their 8x8 gemv design and the fork itself made the Maple results possible.
  • llama.cpp / ggml β€” the core inference engine and its maintainers and contributors.
  • vllm.cpp (mudler) β€” the engine we ported into.
  • The HuggingFace / GGUF ecosystem β€” the format, the tooling, and the platform.

Without their work, none of this project would exist. Any remaining errors are ours.


Report written 2026-08-08. Benchmarks are indicative single-machine measurements; expect day-to-day noise of Β±10–20 % on the thinkpad2.