secjev encoder

LFM2.5-Encoder-350M fine-tuned to answer the 25 questions of the CWE Top 25 about a window of up to 240 lines of source code. One forward pass reads a window and returns P(yes) for all 25 questions, for example "Is a SQL query built here by concatenating or interpolating data, rather than by binding parameters?" (cwe_89). It reads 28 languages, from C and Go to PHP and Solidity.

window text โ”€โ”€โ–บ LFM2.5 bidirectional encoder โ”€โ”€โ–บ attention pool (one per question) โ”€โ”€โ–บ linear โ”€โ”€โ–บ sigmoid(logit / T)

It is a triage aid that ranks where to look first. It does not prove that a flaw exists.

Files

path what
model/model.safetensors, model/secjev.json PyTorch weights (1.4 GB): the encoder plus the per-question attention pool and the linear head; labels, questions, temperature
onnx/model.onnx fp32 ONNX graph (1.4 GB): matches PyTorch to 1e-5 in the logits
onnx/model_q8.onnx + onnx/model_q8.weights 8-bit weights in blocks of 32 with fp32 maths (593 MB), for ONNX Runtime on CPU or CUDA, and WebGPU
onnx/model_q8_wasm.onnx the same weights file, with attention written as per-head operators, for onnxruntime-web on WebAssembly
onnx/bundle.json labels, questions, per-language limits, temperature, window rules and the file-extension map
*/tokenizer.json the base model's tokenizer
onnx/validate-validation.json the validation numbers below

Each ONNX graph takes input_ids (int64, [1, n], n โ‰ค 8192, BOS first, no padding) for one window and returns logits ([1, 25]). P(yes) is sigmoid(logit / temperature). Score windows one at a time, because the graphs use ONNX Runtime's MultiHeadAttention without a padding mask. That keeps an 8,192-token window from building a 4.3 GB score matrix. The two 8-bit graphs share model_q8.weights, so download it next to whichever graph you use.

Input format

The model was trained on windows of this exact shape: a header, then each line numbered and cut at 400 characters. Files are tiled every 240 lines from line 1, and a tile whose text is over 33,084 characters is halved until it fits. bundle.json โ†’ window holds these constants. limits says which languages each question applies to; the memory-safety questions, for example, are asked only of C and C++.

// ==== src/app.py:1-240 of 512 ====
1 | import os, subprocess
2 | from flask import request
...

Usage (ONNX Runtime, Python)

import json
import numpy as np
import onnxruntime as ort
from tokenizers import Tokenizer

d = "onnx/"
meta = json.load(open(d + "bundle.json"))
tok = Tokenizer.from_file(d + "tokenizer.json")
sess = ort.InferenceSession(d + "model_q8.onnx", providers=["CPUExecutionProvider"])

def window(path, lines, start, total):  # the format the model was trained on
    body = "\n".join(f"{start + i} | {l[:400] + ' ... [long line cut]' if len(l) > 400 else l}"
                     for i, l in enumerate(lines))
    return f"// ==== {path}:{start}-{start + len(lines) - 1} of {total} ====\n" + body

src = open("app.py").read().split("\n")[:240]
ids = tok.encode(window("app.py", src, 1, len(src))).ids[: meta["max_length"]]
logits = sess.run(None, {"input_ids": np.array([ids], dtype=np.int64)})[0][0]
p = 1 / (1 + np.exp(-logits / meta["temperature"]))
for i in np.argsort(-p)[:5]:
    cwe = meta["labels"][i]
    print(f"{p[i]:.3f}  {cwe}  {meta['questions'][cwe]}")

In a browser, load the graphs through the onnxruntime-web/webgpu entry, which runs 8-bit MatMulNBits. The package's default entry does not. Use model_q8_wasm.onnx when WebGPU is missing. Multi-threaded WebAssembly needs COOP/COEP headers.

Training

  • Real labels only. Positives are the window before a CVE/GHSA fix, labelled on the questions its advisory names. The same place after the fix is the matching negative. Other questions on fix windows are masked. Windows from a sweep of ordinary open-source code are negatives on every question, at weight 0.25. No LLM-generated labels are used.
  • Pair loss. Within each fix, the window before the fix should score above the window after it (a logistic loss on the gap). This term made the model learn the flaw itself rather than "looks like code that gets fixed".
  • Attention pooling. Each question learns its own weighting of the window's tokens. The head learns at 5e-4 and the encoder at 3e-5, in bf16, up to 8,192 tokens.
  • Selection and calibration. The checkpoint was chosen by auroc on the fix validation windows. One temperature (1.6356) was fitted on held-out calibration windows.
  • Data. The permissively licensed part of the secjev dataset: 52,673 fix windows and 200,000 ordinary windows. Each fix window is seen twice per epoch.

Evaluation

Validation split: 3,060 fix windows plus 10,000 ordinary windows, 209,900 (window, question) scores in all.

auroc (before vs after fix) pre_above_post before fix vs ordinary fixed vs ordinary ordinary flagged (p โ‰ฅ 0.5)
PyTorch, bf16 0.5830 0.6560 0.9369 0.8996 0.75%
model.onnx (fp32) 0.5832 0.6680 0.9368 0.8994 0.75%
model_q8.onnx 0.5830 0.6691 0.9370 0.8999 0.76%
model_q8_wasm.onnx 0.5830 0.6674 0.9370 0.8999 0.76%

What the columns mean:

  • before fix vs ordinary (โ‰ˆ0.94): the model separates code that later needed a security fix from ordinary code well.
  • auroc and pre_above_post: it separates the vulnerable version from its own fixed version, often a few changed lines, only modestly.
  • fixed vs ordinary (0.90): much of the signal is "this is the kind of code where flaws of this class live". Read a high score as "look here", not "this is vulnerable".

Against fp32, the 8-bit graphs move a probability by 0.0005 on average (99th percentile 0.009, max 0.056). They keep 98.4% of the top 500 scores and 99.97% of the p โ‰ฅ 0.5 calls.

Speed per window, q8 (RTX 4070 Ti SUPER, i9-14900KF):

tokens CUDA CPU, 32 threads WebGPU (Chrome) WebAssembly (Chrome)
2,048 0.05 s 1.3 s 0.17-0.33 s 3.0 s
8,192 0.3 s 5.9 s 1.5 s 15 s

Limitations

  • It sees one window and nothing else: no callers, no configuration, no data flow across files. Questions such as authorization (cwe_862) or CSRF (cwe_352) often hinge on context outside the window.
  • Most training positives come from public CVE fixes, so the scores lean toward flaws that were found and fixed in open-source projects.
  • Expect false positives on ordinary code at roughly the rate in the table. Treat the output as a ranking for human review.

License

This model is a derivative of LiquidAI/LFM2.5-Encoder-350M and is distributed under the LFM Open License v1.0. Commercial use is licensed only to organisations below US$10M in annual revenue; see section 5 of the license.

Changes from the base model: the encoder weights were fine-tuned, a per-question attention pool and a 25-way linear head were added, and the result was exported to ONNX, including an 8-bit quantised copy.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for billytesterman/secjev-encoder

Quantized
(4)
this model