Instructions to use macmacmacmac/Sev-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use macmacmacmac/Sev-4B with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
Sev-4B: security evidence and response policies
Sev-4B scores supplied answers to questions about security logs and recovered
programs. v0.3.1-html-contrast-research continues
v0.3.0-response-policy-research. It is a Qwen3.5-4B LoRA adapter with a
decision head. It does not generate text.
On 175 policy questions from 35 source programs held out of training, this
checkpoint answers 158/175, 90.3%, up from 147/175 on v0.3.0. At the
calibration-selected alert cut it catches 35/35 required approvals and
raises 5/140 false alerts, 3.6%, under the 5% ceiling. v0.3.0 raised
14/140 and failed that ceiling. Twelve of those fourteen were allowed HTML
reads. This pass trains the contrast on the same program: writing innerText
requires approval, and parsing HTML does not. The DNS diagnostic was not rerun.
The trial's isolation-and-packing check failed. v0.3.0 remains available.
Use
Install the Sev runtime. A generic text-generation pipeline does not load the decision head.
git clone https://github.com/maceip/Sev.git
cd Sev
uv sync --extra serve
uv run python -m kev.serve \
--run macmacmacmac/Sev-4B@v0.3.1-html-contrast-research --port 8009
Send TypeSafe-shaped requests to POST /v1/systemone. Supply the observed
evidence, the applicable policy, and explicit answer choices. The model reads
supplied code without executing it. Its scores support analyst review; they do
not establish that a program ran, identify its human owner, or prove agent origin.
The checkpoint ships with temperature 1.0. The false-alert cut for the
policy panel is 1.0 on that raw approval probability: alert only when the
approval option is certain. That cut was chosen on the 140 policy calibration
questions, at a 5% false-alert budget, and then applied to development.
Reported evaluation uses
KEV_BACKEND=torch KEV_DTYPE=fp32 KEV_MERGE=1 on CUDA. Accelerated serving can
differ numerically. The API's confidence rescales the largest probability above
chance; it is not an independently verified correctness probability.
Training and lineage
| Setting | Value |
|---|---|
| Backbone | Qwen/Qwen3.5-4B-Base@1001bb4d826a52d1f399e183466143f4da7b741b |
| Immediate parent | macmacmacmac/Sev-4B@a1824aefba305fda86e3503b895a7d9b3871f79a |
| Selected run | sev-r2-response-policy-4b-v2/00-trial-0 |
| Training records / questions | 6,577 / 9,456 |
| Retained curriculum records | 6,438 |
| Added source programs / policy questions | 139 / 278 |
| Epochs / seed | 1 / 4 |
| Learning rate | 2.5e-6 |
| Batch / accumulation | 4 / 2 |
| LoRA / head | Rank 16, all targets / 256 dimensions |
| Precision | fp32 frozen weights, bf16 autocast |
| Updates / forward tokens | 823 / 1,505,379 |
| Rejected or truncated training records | 0 |
All previous curriculum records remain byte-identical. They include native Sysmon, ExCyTIn, GUIDE, public classification and authored-rule replay, and 446 SwarmTraces static-program records. The added programs have authored API-specific policies and checked labels. Their source text is real recovered code, not a native execution trace. Source-closure groups separate training, calibration and development. No locked test was read.
Training configuration, provenance,
and data lineage pin the inputs. The original
v0.1.0-research
synthetic-origin preview is a separate historical lineage. The previous
v0.2.0-swarmtraces-research
release remains available unchanged.
Matched development results
| Policy decision | v0.3.0 | v0.3.1 |
|---|---|---|
| Correct argmax answers | 147/175 | 158/175 |
| Required approvals detected at the alert cut | 25/35 | 35/35 |
| False alerts at that cut | 14/140 | 5/140 |
| Allowed HTML questions flagged | 12/21 | 0/21 |
Each alert cut was selected on calibration only, with a 5% false-alert budget. v0.3.1's cut is 1.0 on the raw approval probability. The 5 remaining false alerts are 4 console questions and 1 plain-text question. These are explicit policy decisions on a development panel, not measured detection rates in live networks.
The retained-task table and the DNS result below were measured on v0.3.0. They were not rerun for v0.3.1.
| Retained panel | v0.2.0 | v0.3.0, not rerun here |
|---|---|---|
| Manual SwarmTraces | 65/83 | 66/83 |
| Static SwarmTraces | 345/354 | 347/354 |
| Native Sysmon | 409/410 | 409/410 |
| ExCyTIn | 398/398 | 398/398 |
| Original Sysmon | 62/64 | 62/64 |
| General | 86/105 | 86/105 |
| GUIDE triage | 130/231 | 131/231 |
| GUIDE detector | 638/1,064 | 639/1,064 |
All 25 correctness-retention checks pass. Unchanged totals do not imply unchanged individual answers. GUIDE detector calibrated negative log loss worsens by 0.0116 and Brier by 0.0077. No matched continuation without the new component was run, so the isolated causal contribution of SwarmTraces is not established.
On v0.3.0 the DNS diagnostic was 21/32 and failed the zero-new-error check. v0.3.1 did not rerun it. Policy evaluation, DNS evaluation, and registered screen retain the complete comparisons.
Calibration and artifact verification
The shipped temperature minimizes question-weighted negative log loss on 2,674 calibration questions from 1,920 records. Five-fold diagnostics keep all 465 source groups intact across task families. Raw calibration ECE is 0.07626 and out-of-fold ECE is 0.03386. This is a calibration diagnostic, not fresh field validation. Development and DNS observations were excluded from fitting.
Only temperature metadata changed when assembling the serving copy. Learned head tensors, adapter and tokenizer bytes match the evaluated checkpoint. Calibration, integrity evidence, and SHA256SUMS identify the release files.
Research iteration is paused following this release. The 0.8B and 9B checkpoints are not updated by this publication.
License and attribution
Source and adapter/head weights carry Apache-2.0 notices. Preserve LICENSE, NOTICE and the Qwen BASE_LICENSE. Sev builds on Kev by Jared Palmer and Qwen3.5 by the Qwen team.
Each dataset retains its own terms. Upstream SwarmTraces reuse terms remain unverified; the model license does not relicense those artifacts. The collection includes metadata-only entries, and membership does not imply training use. The historical synthetic source's notice applies only to that source. See data provenance.
- Downloads last month
- 52
Model tree for macmacmacmac/Sev-4B
Base model
Qwen/Qwen3.5-4B-Base