SkillJev-2B

SkillJev-Direct v1 scores malicious AI-agent skills from SKILL.md, using OpenJev 2B, LoRA and a trained classification head. The scoring pipeline returns a calibrated score without generating text.

AI team from University of Toronto, in collaboration with TrendAI.

Use

The complete merged model is at the repository root; the original adapter and head are in adapter/. Download version v1. See the loading and scoring recipe and score.py.

Premise: SKILL.md. Fixed hypothesis:

This AI agent skill is malicious.

The 8,192-token input retains the Markdown prefix and complete hypothesis. Use scoring_config.json for calibration and the frozen review threshold; native entailment softmax is not the calibrated probability.

Results

MaliciousSkillBench, official Source-Balanced Random protocol: 6,812 train / 973 validation / 1,950 test.

Method Test recall Benign false-positive rate
SkillJev-Direct merged v1 99.13% 5.57%
Word-SVM 99.40% 40.98%
Untuned OpenJev-2B 99.47% 99.78%

Thresholds were frozen on validation, targeting 99% recall. Merged Direct has 159 fewer false positives and 4 additional misses than Word-SVM. One epoch, one seed; sources are shared across partitions. Export verification compares the merged model with the selected adapter using unchanged settings.

Scope

Markdown-only detection does not certify a complete package as safe; performance on other distributions is unverified. OpenJev is labelled MIT; third-party training texts retain their upstream terms. No skill payloads or TrendAI data are hosted here.

Downloads last month
12
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SkillJev/SkillJev-2B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(5)
this model

Dataset used to train SkillJev/SkillJev-2B