secjev encoder
LFM2.5-Encoder-350M fine-tuned to
answer the 25 questions of the CWE Top 25 about a window of up to 240 lines of source code.
One forward pass reads a window and returns P(yes) for all 25 questions, for example
"Is a SQL query built here by concatenating or interpolating data, rather than by binding
parameters?" (cwe_89). It reads 28 languages, from C and Go to PHP and Solidity.
window text โโโบ LFM2.5 bidirectional encoder โโโบ attention pool (one per question) โโโบ linear โโโบ sigmoid(logit / T)
It is a triage aid that ranks where to look first. It does not prove that a flaw exists.
Files
| path | what |
|---|---|
model/model.safetensors, model/secjev.json |
PyTorch weights (1.4 GB): the encoder plus the per-question attention pool and the linear head; labels, questions, temperature |
onnx/model.onnx |
fp32 ONNX graph (1.4 GB): matches PyTorch to 1e-5 in the logits |
onnx/model_q8.onnx + onnx/model_q8.weights |
8-bit weights in blocks of 32 with fp32 maths (593 MB), for ONNX Runtime on CPU or CUDA, and WebGPU |
onnx/model_q8_wasm.onnx |
the same weights file, with attention written as per-head operators, for onnxruntime-web on WebAssembly |
onnx/bundle.json |
labels, questions, per-language limits, temperature, window rules and the file-extension map |
*/tokenizer.json |
the base model's tokenizer |
onnx/validate-validation.json |
the validation numbers below |
Each ONNX graph takes input_ids (int64, [1, n], n โค 8192, BOS first, no padding) for
one window and returns logits ([1, 25]). P(yes) is sigmoid(logit / temperature).
Score windows one at a time, because the graphs use ONNX Runtime's MultiHeadAttention
without a padding mask. That keeps an 8,192-token window from building a 4.3 GB score
matrix. The two 8-bit graphs share model_q8.weights, so download it next to whichever
graph you use.
Input format
The model was trained on windows of this exact shape: a header, then each line numbered
and cut at 400 characters. Files are tiled every 240 lines from line 1, and a tile whose
text is over 33,084 characters is halved until it fits. bundle.json โ window holds
these constants. limits says which languages each question applies to; the memory-safety
questions, for example, are asked only of C and C++.
// ==== src/app.py:1-240 of 512 ====
1 | import os, subprocess
2 | from flask import request
...
Usage (ONNX Runtime, Python)
import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer
d = "onnx/"
meta = json.load(open(d + "bundle.json"))
tok = Tokenizer.from_file(d + "tokenizer.json")
sess = ort.InferenceSession(d + "model_q8.onnx", providers=["CPUExecutionProvider"])
def window(path, lines, start, total): # the format the model was trained on
body = "\n".join(f"{start + i} | {l[:400] + ' ... [long line cut]' if len(l) > 400 else l}"
for i, l in enumerate(lines))
return f"// ==== {path}:{start}-{start + len(lines) - 1} of {total} ====\n" + body
src = open("app.py").read().split("\n")[:240]
ids = tok.encode(window("app.py", src, 1, len(src))).ids[: meta["max_length"]]
logits = sess.run(None, {"input_ids": np.array([ids], dtype=np.int64)})[0][0]
p = 1 / (1 + np.exp(-logits / meta["temperature"]))
for i in np.argsort(-p)[:5]:
cwe = meta["labels"][i]
print(f"{p[i]:.3f} {cwe} {meta['questions'][cwe]}")
In a browser, load the graphs through the onnxruntime-web/webgpu entry, which runs 8-bit
MatMulNBits. The package's default entry does not. Use model_q8_wasm.onnx when WebGPU
is missing. Multi-threaded WebAssembly needs COOP/COEP headers.
Training
- Real labels only. Positives are the window before a CVE/GHSA fix, labelled on the questions its advisory names. The same place after the fix is the matching negative. Other questions on fix windows are masked. Windows from a sweep of ordinary open-source code are negatives on every question, at weight 0.25. No LLM-generated labels are used.
- Pair loss. Within each fix, the window before the fix should score above the window after it (a logistic loss on the gap). This term made the model learn the flaw itself rather than "looks like code that gets fixed".
- Attention pooling. Each question learns its own weighting of the window's tokens. The head learns at 5e-4 and the encoder at 3e-5, in bf16, up to 8,192 tokens.
- Selection and calibration. The checkpoint was chosen by
aurocon the fix validation windows. One temperature (1.6356) was fitted on held-out calibration windows. - Data. The permissively licensed part of the secjev dataset: 52,673 fix windows and 200,000 ordinary windows. Each fix window is seen twice per epoch.
Evaluation
Validation split: 3,060 fix windows plus 10,000 ordinary windows, 209,900 (window, question) scores in all.
| auroc (before vs after fix) | pre_above_post | before fix vs ordinary | fixed vs ordinary | ordinary flagged (p โฅ 0.5) | |
|---|---|---|---|---|---|
| PyTorch, bf16 | 0.5830 | 0.6560 | 0.9369 | 0.8996 | 0.75% |
model.onnx (fp32) |
0.5832 | 0.6680 | 0.9368 | 0.8994 | 0.75% |
model_q8.onnx |
0.5830 | 0.6691 | 0.9370 | 0.8999 | 0.76% |
model_q8_wasm.onnx |
0.5830 | 0.6674 | 0.9370 | 0.8999 | 0.76% |
What the columns mean:
- before fix vs ordinary (โ0.94): the model separates code that later needed a security fix from ordinary code well.
- auroc and pre_above_post: it separates the vulnerable version from its own fixed version, often a few changed lines, only modestly.
- fixed vs ordinary (0.90): much of the signal is "this is the kind of code where flaws of this class live". Read a high score as "look here", not "this is vulnerable".
Against fp32, the 8-bit graphs move a probability by 0.0005 on average (99th percentile 0.009, max 0.056). They keep 98.4% of the top 500 scores and 99.97% of the p โฅ 0.5 calls.
Speed per window, q8 (RTX 4070 Ti SUPER, i9-14900KF):
| tokens | CUDA | CPU, 32 threads | WebGPU (Chrome) | WebAssembly (Chrome) |
|---|---|---|---|---|
| 2,048 | 0.05 s | 1.3 s | 0.17-0.33 s | 3.0 s |
| 8,192 | 0.3 s | 5.9 s | 1.5 s | 15 s |
Limitations
- It sees one window and nothing else: no callers, no configuration, no data flow across
files. Questions such as authorization (
cwe_862) or CSRF (cwe_352) often hinge on context outside the window. - Most training positives come from public CVE fixes, so the scores lean toward flaws that were found and fixed in open-source projects.
- Expect false positives on ordinary code at roughly the rate in the table. Treat the output as a ranking for human review.
License
This model is a derivative of LiquidAI/LFM2.5-Encoder-350M and is distributed under the LFM Open License v1.0. Commercial use is licensed only to organisations below US$10M in annual revenue; see section 5 of the license.
Changes from the base model: the encoder weights were fine-tuned, a per-question attention pool and a 25-way linear head were added, and the result was exported to ONNX, including an 8-bit quantised copy.
Model tree for billytesterman/secjev-encoder
Base model
LiquidAI/LFM2.5-350M-Base