LFM2.5-350M β€” Core AI (.aimodel)

LiquidAI/LFM2.5-350M converted to Core AI .aimodel bundles for Apple silicon by visible-cx. These are derivative artifacts: Liquid AI's weights re-expressed as a Core AI graph with int8 block-32 symmetric weight quantization and a two-entrypoint (decode + chunked-prefill) function map. They load through Core AI on macOS and are not usable by PyTorch, GGUF or MLX.

This is the fastest and smallest of the dense LFM2.5 bundles published here: around half a second per guided structured-output sample on a 16 GB Mac, at 1.21 GB peak resident and 155–166 tok/s. It runs on an 8 GB machine with room to spare.

⚠️ Recommended maximum context: 8192. Needle recall is 3/3 at 8k and 1/3 at 15k, failing by fabrication β€” it returned an invented entity and an invented date rather than declining. See Context window before routing deep-context work to this model.

Contents

Path Bytes Manifest context Functions
gpu-pipelined/lfm2_5_350m_decode_int8hu_block32_sym_mf64 571,453,194 4096 main + prefill
ctx8192/gpu-pipelined/lfm2_5_350m_decode_int8hu_block32_sym_mf64 571,453,212 8192 main + prefill

Each folder holds <name>.aimodel/ (main.mlirb β‰ˆ 567 MB, main.hash, asset metadata.json), a bundle-level metadata.json, and tokenizer/ (tokenizer.json, tokenizer_config.json, generation_config.json, chat_template.jinja).

Both folders hold the same weights and the same graph. --max-ctx does not change the exported graph; it sets language.max_context_length in the bundle manifest and nothing else. No ctx16384 folder is published for this model, because its recall fails by fabrication above 8k (see Context window).

Stop token: eos_token = "<|im_end|>" in both folders. Clean self-stop on every measured sample.

Provenance

Base checkpoint LiquidAI/LFM2.5-350M
Recipe export_lfm2_multifunction.py int8hu --head-sym --chunk 64
Toolchain base apple/coreai-models @ b1cb71b8522d99408059fa0b98b8742171bcb0b8 + the coreai-model-zoo python overlay
Toolchain coreai-torch 0.4.1, coreai-core 1.0.0b2, coreai-opt 0.2.1, torch 2.9.0
Producer fingerprint coreai-core 1.0.0b2 on every inner <name>.aimodel/metadata.json
Weight format int8, per-K-block-32, symmetric; symmetric head (--head-sym)
Vocab 65,536
Export functions main (S=1 decode) + prefill (S=64 chunked prefill), function_map: {"main": ["main", "prefill"]}, weights shared

mf64 in the bundle name means multifunction with a 64-wide prefill.

The zoo ships no LFM2.5-350M recipe (models/lfm2.5/recipe.toml covers the 1.2B only), so this is a new conversion against export_lfm2_multifunction.py rather than a reproduction of a published coreai-model-zoo bundle. The larger LFM and Gemma bundles in this org do mirror zoo recipes; this one does not, and the flags above are the whole specification.

Requirements

  • Apple silicon Mac, Core AI runtime.
  • Engine contract: 2 inputs (input_ids, position_ids) β†’ logits. No static inputs, no per-step mask. Runs on the pipelined engine and on the logits-capable sequential engine, so grammar-constrained decoding is available.
  • States: keyCache / valueCache Float16, 6 Γ— 1 Γ— 8 Γ— ? Γ— 64 plus convState Float16, 10 Γ— 1 Γ— 1024 Γ— 2. Dynamic sequence dim β†’ GrowingKVCache (initial 256, doubling). convState is fixed-size and does not scale with context.
  • KV cost: 12,288 bytes per token (fp16) β€” 50 MB at 4096, 101 MB at 8192.
  • Minimum practical machine memory: 8 GB.
  • The bundle manifest declares runtime_env COREAI_CHUNK_THRESHOLD=1.

Measurements

Measured on a 16 GB Apple silicon Mac (M2 Pro, macOS 27 beta).

Guided structured output

10-sample harness, guided JSON-constrained decoding against a fixed schema, greedy, sequential engine, reset() between samples, 128-token cap. Load excluded from s/row; sample 1 excluded as a cache-warm outlier.

Cold load 3.6 s
Guided JSON parse 10/10
Enum-clean 10/10
s/row (long samples) 0.84
s/row (short samples) 0.55
Decode 155–166 tok/s
TTFT 0.16–0.79 s
Peak footprint 0.21 GB
Max RSS 1.21 GB

Judged on parse rate, enum conformance, content sanity and clean stop rather than an exact-output comparison. Its outputs are the least specific of the three dense LFM2.5 bundles in this set, which is the trade for being 4–5Γ— faster than the 2.6B.

Unguided workload

Free-form generation from a fixed prompt, 900-token cap; the model self-stopped well inside the cap at both depths.

depth prompt tokens TTFT decode generated wall peak footprint
3.4k 3,371 1.48 s 145.6 tok/s 725 6.7 s 0.15 GB
8k 7,953 2.89 s 127.5 tok/s 546 7.4 s 0.30 GB

Speed is not this model's limitation; long-range recall is.

Memory, measured rather than inferred

Max RSS counts clean mapped pages the kernel can drop for free, so it is not what the machine has to give up. Measured with an external watchdog sampling wired memory:

GiB
bundle on disk 0.532
compiled blob 0.688
graph constant 0.160
blob Γ· bundle 1.29Γ—
wired floor (32-token leg β€” a lower bound, not a plateau) 0.790
forecast (blob Γ— 1.106) 0.761
requirement (floor + 1.25 GiB in-flight floor) 2.27

The one place this project's plateau law is not conservative is here, and it is stated rather than buried. On this bundle blob Γ— 1.106 reads 3.7% LOW against a number that is itself a lower bound, so the true under-read is larger β€” plausibly nearer 7%, since 32-token legs read about 3% light on a 2.6B control. Nothing ships wrong, because the measured floor outranks the forecast and the floor is what the gate uses. But a future small bundle with no floor would be gated optimistically. The mechanism is plausible and unmeasured: the runtime's own fixed costs are a larger fraction of a 0.69 GiB blob than of a 6 GiB one, and a ratio law has no constant term to carry them.

The prompt-length sweep β€” this bundle was the instrument

This model is small enough that a slope in memory-vs-prompt-length is visible above the noise, so it was the vehicle for settling a question the larger bundles could not answer: does a per-pass workspace grow with sequence length?

Eight prompt lengths from 284 to 12,099 tokens (a 43Γ— span), the real report corpus with the question header and instruction tail held fixed and only the evidence bullets varied, so every point is the same task at a different length. Run twice in opposite order, with repeats chosen per point so every leg runs ~60 s and yields 92–121 samples (unequal sample counts manufacture a slope out of noise, since wired_peak is a max over samples). The wired ceiling was deliberately set ~3 GiB above the plateau so the criterion could not bind; no leg tripped.

sweep A sweep B
wired growth (median βˆ’ baseline) +0.0134 +0.0163
process phys_footprint (median) +0.0186 +0.0184

GiB per 1,000 prompt tokens. The footprint slope reproduces to 1% across two sweeps with different repeat counts and opposite order. 0.018 GiB per 1k tokens is 19 MiB per 1k tokens, and it lands within 18% of pure KV-first-touch arithmetic β€” so the answer is that the term is the KV, not a workspace, and it is already charged by any gate that counts KV.

Prefill is linear and the rate is constant: TTFT 0.17 β†’ 3.80 s over 284 β†’ 12,099 tokens, a flat ~3,180 tok/s. There is no super-linear term in time either.

Caveats, stated: the machine's own noise floor on wired growth is ~Β±0.8 GiB, larger than the entire slope, so the conclusion rests on the monotone process-footprint series and the wired series is corroboration. The sweep bundle was a locally staged ctx16384 clone of these same weights and graph (blob 738,512,596 B, main.hash 629b1d59663f5003…) β€” the manifest integer is the only difference from the folders published here. It is one model, one machine, one afternoon.

Context window

A needle-in-haystack probe planted three distinctive facts at 10% / 50% / 90% of the filler and asked for them back at the end. Strict scoring: a fact counts only if the distinctive entity comes back correct.

depth prompt tokens result
8k 8,099 3/3 β€” all three entities returned (the date on fact 3 was omitted)
15k 15,106 1/3 β€” fabricated

The failure mode at 15k is invention, not refusal: fact 1 survived, fact 2 vanished, and fact 3 came back with a fabricated entity and a fabricated date. The output looks correct and is not.

Recommended maximum context: 8192. Use the 1.2B or 2.6B bundles above 8k; both return 3/3 verbatim at 15k. The exact boundary between 8k and 15k is not measured β€” only the two endpoints are.

Usage

Swift Package Manager, via CoreAIKit β€” a community package, not affiliated with Apple, requiring macOS 27 beta:

.package(url: "https://github.com/john-rocky/coreai-kit", branch: "main")
// target dependency: .product(name: "CoreAIKit", package: "coreai-kit")

ModelID addresses a bundle as repo + path + revision, where path is the subtree in this repo holding one complete bundle (metadata.json + *.aimodel/ + tokenizer/). It downloads from the Hub on first use and is cached afterwards:

import CoreAIKit

let model = ModelID(
    "visible-cx/LFM2.5-350M-CoreAI",
    path: "ctx8192/gpu-pipelined/lfm2_5_350m_decode_int8hu_block32_sym_mf64")

var config = ChatSession.Configuration()
config.engineVariant = .sequential   // required for guided / grammar-constrained decoding
config.temperature = nil             // greedy

let chat = try await ChatSession(model: model, configuration: config)
for try await event in chat.streamResponse(to: "…") {
    if case .response(let delta) = event { print(delta, terminator: "") }
}

Pass revision: a Hub commit hash to pin an immutable bundle. ChatSession(bundleAt:) loads a bundle directory already on disk. Leave COREAI_CHUNK_THRESHOLD alone β€” the manifest sets it.

Integrity

Core AI .aimodel bundles are not byte-reproducible: the exporter is not deterministic even against itself. Verify by digesting the exact published bytes rather than by rebuilding. Every bundle carries main.hash, the raw 32 bytes of sha256(main.mlirb), so a downloaded bundle can be checked against itself; on the Hub the same value is recoverable from the LFS oid without fetching the file.

Status

Artifact Status
gpu-pipelined/…_mf64 (ctx 4096) QUALIFIED β€” measured: 10/10 guided parse and enum-clean, 0.84/0.55 s/row, 155–166 tok/s, requirement 2.27 GiB.
ctx8192/…_mf64 QUALIFIED AT DEPTH β€” 8192 is this model's ceiling. Same weights and graph; measured at depth: 3/3 needle recall at 8,099 tokens, 0.30 GB peak footprint. Do not widen the manifest beyond 8192.

No PyTorch-reference oracle or PSNR gate has been run against these bundles.

License

LiquidAI/LFM2.5-350M is released under the LFM Open License v1.0 (lfm1.0), and upstream declares it as license: other + license_name: lfm1.0. These bundles are a derivative and the same licence and obligations travel with them β€” see the upstream licence. Redistribute the licence with the files. The contribution here is the conversion and the qualification evidence, not the weights.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for visible-cx/LFM2.5-350M-CoreAI

Quantized
(64)
this model