LFM2.5-2.6B-CoreAI / RECIPE.md
harshav's picture
LFM2.5-2.6B Core AI int8: card, recipe, upstream license, validation evidence
f9da83b verified
|
Raw
History Blame Contribute Delete
7.18 kB

Conversion recipe

This document is the reproducibility record for the published asset. It is written from scratch and contains no code from Apple's repositories; it describes what was done precisely enough to redo it, for a reader who has access to the tooling described in "Prerequisites".

No conversion scripts are shipped in this repository. They import modules from Apple's coreai-model-zoo, which is not publicly available, so publishing them would either redistribute code that is not mine or hand you something that cannot run. A precise description is more useful than either.

Prerequisites

Component Pin Public?
LiquidAI/LFM2.5-2.6B dca1825886789bd40b94368f53b1d9ada4c94598 yes
github.com/apple/coreai-models (converter) b1cb71b8522d99408059fa0b98b8742171bcb0b8 yes
github.com/apple/coreai-models (runtime) 5ed9981303b38d5a44aa6b45509bc4f6945029f5 yes
coreai-torch c89f6a44713249a12a84beec9f3e0cf2206ecc38 yes
apple/coreai-model-zoo โ€” LFM2 overlay and runtime patch stack ebef921a1f358af66c9ff67e8c6e7d4e24efad0d no

Toolchain used: macOS 27.0 (build 26A5388g), Xcode 27.0 (27A5228h), Python 3.11.15, torch 2.9.0, coremltools 9.0.

The zoo is the blocker for both directions, and there is no way around it from here:

  • Conversion needs the zoo's overlay, because that overlay is what carries the LFM2 authoring module (models/macos/lfm2.py) โ€” the converter alone does not know this architecture.
  • Inference needs the zoo's runtime patch stack (coreai-pipelined-per-token-inputs, -static-inputs, -extra-states, coreai-prefix-cache, coreai-shared-product). An unpatched runtime at the pinned commit does not accept this asset's per-token input contract.

Quantization

Applied to the authored module before export, then exported to the Core AI dialect.

Tensor group Precision Detail
Linear / MLP weights int8 blockwise, block size 32, per-block scales
lm_head int8 blockwise 32, symmetric; the head is untied and is ~0.5 GB
Attention q,k,v,out projections fp16 overlay default is fp32; overridden
Token embedding fp16 left unquantized
Norms, RoPE tables, indices fp16 / int32 untouched

The resulting compiled storage budget, which is the check that a rebuild matched:

Int8     2,621,243,392
Float16    428,342,276
Float32             34
Int32              312
UInt32              71
UInt64               1

Graph shape: input_ids [1,1] static, position_ids dynamic, KV cache dynamic on the sequence axis, max_context_length = 4096. Decode-only; no chunked prefill entrypoint.

Two deviations from the zoo recipe's defaults were measured rather than inherited:

  1. Attention projections fp16 instead of fp32. The overlay promotes these to fp32 for GPU-delegate exactness. On this model that precision is not needed, and fp32 costs ~168 MB of reads on every decode step.
  2. Attention projections were not taken to int8. That is a further ~2.5 % throughput for one lost position in 125; the higher-fidelity option was shipped instead.

One correctness fix was required on the overlay, and it matters more than either:

The checkpoint carries no top-level rope_theta. It ships rope_parameters.rope_theta = 1e7 (the transformers โ‰ฅ 5 layout). Code that reads only the legacy key silently falls back to 1e6 โ€” a 10ร— wrong RoPE that still produces fluent short completions and only clearly breaks at long context. Both the overlay and transformers 4.x hit this. Any reproduction must read the nested key.

A second, latent one: the checkpoint spells tying tie_word_embeddings, not tie_embedding. The default is correct here, so nothing breaks on this model, but it would flip silently on an untied checkpoint.

Gates

Four separate questions, deliberately not collapsed into one number.

  1. Authoring fidelity โ€” the re-authored module vs Hugging Face Lfm2ForCausalLM, both fp32, teacher-forced. Result 21/21 top-1, cosine 1.000000. This is the gate that caught the RoPE bug.
  2. Quantization damage โ€” the quantized module vs an independent fp32 Hugging Face reference (transformers โ‰ฅ 5.2), teacher-forced over 5 sequences / 125 positions. Result 122/125 top-1, minimum per-position cosine 0.997050.
  3. Conversion fidelity โ€” the exported bundle vs its own quantized weights run eagerly, greedy, 5 prompts. Result 5/5 exact. Comparing the bundle to fp32 here would conflate quantization damage with conversion bugs, so it is compared to the thing it is supposed to equal.
  4. Throughput โ€” measured before any gate loads the model, because loading first cost ~10 % on an identical bundle.

Quality is teacher-forced throughout. Free-running text is not usable as a gate on this model: every probe prompt contains at least one step with a sub-0.05 top-2 margin, so transcripts diverge on near-ties without indicating damage.

Measurement protocol

Comparisons below ~5 % are meaningless without this. Early runs showed ~3 % spread on a byte-identical bundle.

  • Clear the Core AI specialization cache entry for this asset only, for the producing binary. The asset's own main.hash is the content key.
  • One throwaway load + short generation to absorb cold specialization.
  • 60 s settle so the SoC sheds export and compile heat.
  • 5 trials, prompt 64 tokens, generate 128, fixed seed.
  • Report between-run spread across independent runs. Within-run standard deviation of adjacent trials is repeatability, not a population statistic, and quoting it as though it bounded the mean overstates confidence badly.

Two environment notes that changed results materially:

  • COREAI_CHUNK_THRESHOLD=1.
  • Ahead-of-time compilation must name one architecture. Compiling without that builds all 20 (~8 GB each). --expect-frequent-reshapes measured 84 tok/s against 160 and 8.3 GB against 3.3 GB, so it is off.

Rejected

Attempt Outcome
int4 blockwise 32 minimum cosine 0.51โ€“0.66 โ€” a different model
int4 blockwise 32, conv projections rescued to int8 cosine 0.662, 16/21 top-1; rescuing conv does not protect the MLP bulk, which is where both the bytes and the damage are
int4 blockwise 16 quality recovers, 42 tok/s โ€” ~3ร— slower than int8, dequantization dominates
int8 token embedding throughput-neutral, โˆ’214 MB; not shipped because it is not a win
--preferred-compute neural-engine no-op; the compiled asset holds an MPSGraph delegate either way. A dynamic KV dimension is not an ANE-shaped graph
Speculative decoding, static-S verify graph exports and gates its contract, but per-position logits do not match stepped decode; not published

Reproduction is verified by the gates and the storage budget above, not by hashing. The exporter names each externalized call site with a generated UUID โ€” 391 such names in this graph โ€” so two exports of identical weights differ in a few bytes and therefore in SHA-256.