LFM2.5-1.2B-Instruct — ROCmFP4 for AMD Strix Halo (gfx1151)

the first ROCmFP4 build of any LFM2.5 checkpoint

Checked 2026-08-22 against every public GGUF of this model. All existing builds (LiquidAI's own, unsloth, and others) ship standard k-quants. ROCmFP4 is a runtime tensor format that exists only in the ROCmFPX fork of llama.cpp. Repository-content comparison only — no third-party build was run or benchmarked here.

A 4-bit ROCmFP4 quantisation of LiquidAI/LFM2.5-1.2B-Instruct for AMD Ryzen AI Max+ 395 / Radeon 8060S / gfx1151.

The file

ftype 102Q4_0_ROCMFP4_COHERENT
size 695,755,456 bytes (0.65 GiB)
architecture lfm2
tensors 148
context 128,000
token embedding Q6_K

Type histogram, read from the finished file:

ROCmFP4 x92, F32 x55, Q6_K x1

Note on tiers: LEAN and COHERENT coincide for this checkpoint

LFM2.5-1.2B-Instruct ties its output projection to token_embd.weight — there is no separate output.weight tensor. At this size both the LEAN (101) and COHERENT (102) tiers select Q6_K for that shared embedding, so the two tiers produce identical tensor typing and identical file size. Only one file is published rather than two that differ solely in their declared ftype.

Measured throughput

AMD Ryzen AI Max+ 395, Radeon 8060S (gfx1151), ROCm 7.13.0, 125 GB unified memory, idle box. llama-cli -ngl 999 -fa on -c 512 -n 64 --temp 0 --seed 1234:

generation
this file 227.1 t/s

A separate 3-repetition benchmark at -c 2048 -n 512 measured 222.5 t/s for this checkpoint with no drafter.

⚠️ DSpark speculative decoding is a NET LOSS on this hardware — do not use it

LiquidAI publishes a DSpark speculator for this model. We measured it and it makes generation slower, so no ROCmFP4 draft is published here.

config generation effect
no drafter 222.5 t/s
--spec-type draft-dspark --spec-draft-n-max 8 159.3 t/s -28.4%

Mean accepted length was 2.34 (block size 9). Across all three LFM2.5 sizes the result was consistently negative: −28.4% (1.2B), −19.1% (2.6B), −38.0% (8B-A1B).

Two causes were identified, both in the runtime rather than the weights:

  1. lfm2.cpp / lfm2moe.cpp do not populate t_layer_inp[], so draft-dspark aborts on GGML_ASSERT(t_layer_inp[il] != nullptr) out of the box. A one-line patch (res->t_layer_inp[il] = prev_cur;) makes it run.
  2. With that fixed, llama.cpp reports recurrent state rollback is not compatible with 'draft-dspark' and falls back to a checkpoint path that is not bit-exact for LFM2's recurrent state — DSpark output diverges from greedy target output (reproducible 3/3).

An off-by-one in the target-layer mapping was ruled out: forcing LLAMA_DFLASH_TARGET_LAYER_OFFSET=-1 produced a worse accepted length (2.22), confirming the converter's +1 convention is correct.

DSpark on LFM2.5 needs real recurrent-state rollback support before any draft is worth shipping.

Requirements

This file uses the ROCmFP4 tensor format, which exists only in the ROCmFPX fork of llama.cpp. Stock llama.cpp will not load it.

llama-cli -m LFM2.5-1.2B-Instruct-Q4_0_ROCMFP4_COHERENT.gguf \
  -ngl 999 -fa on -c 2048 -n 512 \
  -p "The history of mathematics begins in ancient times. One of the earliest known"

Sample output

Continuation from "The history of mathematics begins in ancient times. One of the earliest known":

The history of mathematics indeed begins in ancient times, with evidence of mathematical thought dating back thousands of years. One of the earliest known mathematical records comes from ancient Mesopotamia, where clay tablets from around 1800 BCE contain mathematical problems and solutions. These tablets show that the Babylonians were skilled in arithmetic, algebra, and

Not measured

Perplexity is not published for this build; quality evidence here is the coherence check above and the tensor-level audit. Long-context behaviour at the full 128,000-token window was not tested.

Provenance

Converted from LiquidAI/LFM2.5-1.2B-Instruct at revision df58c174f05ff733f83f8cae10ea9298224c8006 to F16 GGUF using upstream llama.cpp at e85caa81ea2b65797396018c179b87ad61fa38ab, then quantised to ftype 102 with the ROCmFPX fork (feature/dspark-v2). Licence inherited from the base model.

Downloads last month
72
GGUF
Model size
1B params
Architecture
lfm2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for kingjones777/LFM2.5-1.2B-Instruct-ROCmFP4-GGUF

Quantized
(81)
this model