mlboydaisuke's picture
LFM2.5-Encoder-350M-Policy-Linter LiteRT: int8 (iPhone-verified bit-exact) + fp16, task-level parity verified
ad40157 verified
|
Raw
History Blame Contribute Delete
2.64 kB
metadata
license: other
license_name: lfm1.0
license_link: LICENSE
base_model: LiquidAI/LFM2.5-Encoder-350M-Policy-Linter
pipeline_tag: token-classification
library_name: litert
tags:
  - litert
  - tflite
  - on-device
  - edge
  - encoder
  - zero-shot
  - policy
  - compliance
  - liquid
  - lfm2
  - lfm2.5

LFM2.5-Encoder-350M-Policy-Linter — LiteRT

LiquidAI/LFM2.5-Encoder-350M-Policy-Linter converted to LiteRT (.tflite) for on-device inference. Zero-shot policy linting: write your rules as free text and the model scores every token against every rule in one CPU pass (demo Space).

File Recipe Size
LFM2.5-Encoder-350M-Policy-Linter_wi8fc.tflite int8 dynamic-range (linears + embedding, convs float) 365 MB mobile + desktop (iPhone-verified bit-exact, 143 ms)
LFM2.5-Encoder-350M-Policy-Linter_fp16.tflite fp16 weights, float compute 713 MB desktop — phone memory limits (XNNPACK per-signature fp32 unpacking)

Signatures

lint_128 / lint_512 (S = 128 / 512, batch 1, right-padded, up to 8 rule slots):

Input Shape
input_ids int32 [1, S] prompt tokens: Policy:\n- <rule 1>\n- <rule 2>…\n\nText:\n<doc>
attention_mask int32 [1, S] 1 = token, 0 = pad
rule_pool float32 [1, 8, S] row r = mean-pool weights over rule r's tokens (1/n each); unused rows all-zero

Output: scores float32 [1, S, 8], zeroed at padded positions. sigmoid(score[t, r]) > 0.5 flags token t under rule r; read flags only for real rules and for the document's token range. The rule_pool build mirrors the router sibling's snippet (LFM2.5-Encoder-350M-Prompt-Router) with the Policy: header.

Verification

Task-level parity vs the PyTorch reference (demo: an email-address share + a delivery-date promise against two rules): fp32, fp16 and int8 all flag the identical 10-token spans for both rules. On an iPhone 17 Pro the int8 file reproduces the desktop outputs bit-exactly (cosine 1.000000, max diff 0.0) at 143 ms per lint_512 pass (6 threads, XNNPACK).

License

LFM Open License v1.0 (see LICENSE, unchanged from the base model). Note the license's commercial-use threshold (Section 5). This repository redistributes converted Derivative Works of LiquidAI/LFM2.5-Encoder-350M-Policy-Linter with modification notices per Section 4; all credit for the model to Liquid AI.