--- license: other license_name: lfm1.0 license_link: LICENSE base_model: LiquidAI/LFM2.5-Encoder-350M pipeline_tag: text-classification library_name: onnx tags: - security - code - cwe - vulnerability-detection - lfm2 - onnx - webgpu --- # secjev encoder [LFM2.5-Encoder-350M](https://huggingface.co/LiquidAI/LFM2.5-Encoder-350M) fine-tuned to answer the 25 questions of the CWE Top 25 about a window of up to 240 lines of source code. One forward pass reads a window and returns P(yes) for all 25 questions, for example "Is a SQL query built here by concatenating or interpolating data, rather than by binding parameters?" (`cwe_89`). It reads 28 languages, from C and Go to PHP and Solidity. ``` window text ──► LFM2.5 bidirectional encoder ──► attention pool (one per question) ──► linear ──► sigmoid(logit / T) ``` It is a triage aid that ranks where to look first. It does not prove that a flaw exists. ## Files | path | what | |---|---| | `model/model.safetensors`, `model/secjev.json` | PyTorch weights (1.4 GB): the encoder plus the per-question attention pool and the linear head; labels, questions, temperature | | `onnx/model.onnx` | fp32 ONNX graph (1.4 GB): matches PyTorch to 1e-5 in the logits | | `onnx/model_q8.onnx` + `onnx/model_q8.weights` | 8-bit weights in blocks of 32 with fp32 maths (593 MB), for ONNX Runtime on CPU or CUDA, and WebGPU | | `onnx/model_q8_wasm.onnx` | the same weights file, with attention written as per-head operators, for onnxruntime-web on WebAssembly | | `onnx/bundle.json` | labels, questions, per-language limits, temperature, window rules and the file-extension map | | `*/tokenizer.json` | the base model's tokenizer | | `onnx/validate-validation.json` | the validation numbers below | Each ONNX graph takes `input_ids` (int64, `[1, n]`, n ≤ 8192, BOS first, no padding) for one window and returns `logits` (`[1, 25]`). P(yes) is `sigmoid(logit / temperature)`. Score windows one at a time, because the graphs use ONNX Runtime's `MultiHeadAttention` without a padding mask. That keeps an 8,192-token window from building a 4.3 GB score matrix. The two 8-bit graphs share `model_q8.weights`, so download it next to whichever graph you use. ## Input format The model was trained on windows of this exact shape: a header, then each line numbered and cut at 400 characters. Files are tiled every 240 lines from line 1, and a tile whose text is over 33,084 characters is halved until it fits. `bundle.json` → `window` holds these constants. `limits` says which languages each question applies to; the memory-safety questions, for example, are asked only of C and C++. ``` // ==== src/app.py:1-240 of 512 ==== 1 | import os, subprocess 2 | from flask import request ... ``` ## Usage (ONNX Runtime, Python) ```python import json import numpy as np import onnxruntime as ort from tokenizers import Tokenizer d = "onnx/" meta = json.load(open(d + "bundle.json")) tok = Tokenizer.from_file(d + "tokenizer.json") sess = ort.InferenceSession(d + "model_q8.onnx", providers=["CPUExecutionProvider"]) def window(path, lines, start, total): # the format the model was trained on body = "\n".join(f"{start + i} | {l[:400] + ' ... [long line cut]' if len(l) > 400 else l}" for i, l in enumerate(lines)) return f"// ==== {path}:{start}-{start + len(lines) - 1} of {total} ====\n" + body src = open("app.py").read().split("\n")[:240] ids = tok.encode(window("app.py", src, 1, len(src))).ids[: meta["max_length"]] logits = sess.run(None, {"input_ids": np.array([ids], dtype=np.int64)})[0][0] p = 1 / (1 + np.exp(-logits / meta["temperature"])) for i in np.argsort(-p)[:5]: cwe = meta["labels"][i] print(f"{p[i]:.3f} {cwe} {meta['questions'][cwe]}") ``` In a browser, load the graphs through the `onnxruntime-web/webgpu` entry, which runs 8-bit `MatMulNBits`. The package's default entry does not. Use `model_q8_wasm.onnx` when WebGPU is missing. Multi-threaded WebAssembly needs COOP/COEP headers. ## Training - **Real labels only.** Positives are the window before a CVE/GHSA fix, labelled on the questions its advisory names. The same place after the fix is the matching negative. Other questions on fix windows are masked. Windows from a sweep of ordinary open-source code are negatives on every question, at weight 0.25. No LLM-generated labels are used. - **Pair loss.** Within each fix, the window before the fix should score above the window after it (a logistic loss on the gap). This term made the model learn the flaw itself rather than "looks like code that gets fixed". - **Attention pooling.** Each question learns its own weighting of the window's tokens. The head learns at 5e-4 and the encoder at 3e-5, in bf16, up to 8,192 tokens. - **Selection and calibration.** The checkpoint was chosen by `auroc` on the fix validation windows. One temperature (1.6356) was fitted on held-out calibration windows. - **Data.** The permissively licensed part of the secjev dataset: 52,673 fix windows and 200,000 ordinary windows. Each fix window is seen twice per epoch. ## Evaluation Validation split: 3,060 fix windows plus 10,000 ordinary windows, 209,900 (window, question) scores in all. | | auroc (before vs after fix) | pre_above_post | before fix vs ordinary | fixed vs ordinary | ordinary flagged (p ≥ 0.5) | |---|---|---|---|---|---| | PyTorch, bf16 | 0.5830 | 0.6560 | 0.9369 | 0.8996 | 0.75% | | `model.onnx` (fp32) | 0.5832 | 0.6680 | 0.9368 | 0.8994 | 0.75% | | `model_q8.onnx` | 0.5830 | 0.6691 | 0.9370 | 0.8999 | 0.76% | | `model_q8_wasm.onnx` | 0.5830 | 0.6674 | 0.9370 | 0.8999 | 0.76% | What the columns mean: - **before fix vs ordinary (≈0.94):** the model separates code that later needed a security fix from ordinary code well. - **auroc and pre_above_post:** it separates the vulnerable version from its own fixed version, often a few changed lines, only modestly. - **fixed vs ordinary (0.90):** much of the signal is "this is the kind of code where flaws of this class live". Read a high score as "look here", not "this is vulnerable". Against fp32, the 8-bit graphs move a probability by 0.0005 on average (99th percentile 0.009, max 0.056). They keep 98.4% of the top 500 scores and 99.97% of the p ≥ 0.5 calls. Speed per window, q8 (RTX 4070 Ti SUPER, i9-14900KF): | tokens | CUDA | CPU, 32 threads | WebGPU (Chrome) | WebAssembly (Chrome) | |---|---|---|---|---| | 2,048 | 0.05 s | 1.3 s | 0.17-0.33 s | 3.0 s | | 8,192 | 0.3 s | 5.9 s | 1.5 s | 15 s | ## Limitations - It sees one window and nothing else: no callers, no configuration, no data flow across files. Questions such as authorization (`cwe_862`) or CSRF (`cwe_352`) often hinge on context outside the window. - Most training positives come from public CVE fixes, so the scores lean toward flaws that were found and fixed in open-source projects. - Expect false positives on ordinary code at roughly the rate in the table. Treat the output as a ranking for human review. ## License This model is a derivative of LiquidAI/LFM2.5-Encoder-350M and is distributed under the [LFM Open License v1.0](LICENSE). Commercial use is licensed only to organisations below US$10M in annual revenue; see section 5 of the license. Changes from the base model: the encoder weights were fine-tuned, a per-question attention pool and a 25-way linear head were added, and the result was exported to ONNX, including an 8-bit quantised copy.