File size: 1,348 Bytes
df7c496
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
---
license: other
license_name: lfm1.0
license_link: https://huggingface.co/LiquidAI/LFM2-350M/blob/main/LICENSE
base_model: LiquidAI/LFM2-350M
tags: [webgpu, wgsl, gptq, int4, browser, text-generation]
---

# lfm2-350m-webgpu — GPTQ-int4 weights

Custom GPTQ-int4 quants (group 32, byte-sliced nibbles, dedicated q4
lm_head, q8 embedding gather) of
[LiquidAI/LFM2-350M](https://huggingface.co/LiquidAI/LFM2-350M) and
[LiquidAI/LFM2.5-350M](https://huggingface.co/LiquidAI/LFM2.5-350M), in the
wire format of the
[lfm2-350m-webgpu](https://huggingface.co/spaces/borkiss/lfm2-350m-webgpu-demo) browser runtime
(hand-written WGSL kernels, no ONNX / transformers.js / GGUF).

- 269 MB per model; ~390 tok/s single-stream decode and ~1640 tok/s batched
  (64 streams) on an Apple M4 (10-core GPU) in Chrome.
- LFM2: ppl +3.2% vs f32 on held-out text.
- LFM2.5: GPTQ calibrated on **chat-formatted** passages (raw-prose
  calibration degrades an instruct checkpoint badly): ppl +4.9% vs f32 in its
  native chat format.
- The runtime is verified bit-exact against a PyTorch dequantized-weights
  reference, greedy token-for-token.

f16/q8 blobs are reproducible with the repo's tools/ scripts.

Try it: **https://huggingface.co/spaces/borkiss/lfm2-350m-webgpu-demo** · run the benchmark on
your GPU and share the .log from the /web/bench.html page.