# Model-card sections that do not depend on the result Drafted while training runs; merged into README.md once the numbers land. The headline framing is deliberately left out — it depends on whether the encoder actually clears the TF-IDF floor, and that gets written to match the outcome. ## Intended use **In scope.** Triage aid for threat-intelligence and detection-engineering work: - First-pass ATT&CK tagging of a threat report, for an analyst to correct - Prioritising which reports to read when triaging a backlog - Rough coverage analysis — which techniques a body of reporting talks about - A baseline to beat, with a published harness for doing so **Out of scope.** - Unreviewed labelling. Output is candidates, not conclusions. - Compliance, audit, or attestation evidence. - Detecting techniques outside the 49 covered. - Reasoning over anything other than English prose describing adversary behaviour. It is not a malware classifier and does not read binaries, logs, or code. **Defensive use only.** The model classifies adversary behaviour already described in public reporting. It generates no offensive capability. ## Limitations - **49 techniques, not the full ATT&CK matrix.** Anything outside the label set is invisible to the model — including techniques the text plainly describes. Absence of a prediction is not evidence of absence. - **Sentence-level context only.** Labels were assigned per sentence, so techniques inferable only from surrounding paragraphs are under-represented. - **Long-tailed.** `T1027` has 678 training instances; the rarest retained techniques have roughly 20. Per-technique F1 varies enormously and the head/tail table in the README is the honest view of that. - **151 source documents.** Even the leak-free split is one draw from a small pool. Treat differences of a point or two between models as noise. - **Annotation is not exhaustive.** Some unlabelled sentences do describe techniques, so measured recall is pessimistic relative to truth. - **Domain shift is untested.** Training text is vendor threat-report prose. Behaviour on incident tickets, chat logs, or non-native-English reporting is unknown and probably worse. - **Per-class thresholds are fitted.** 49 thresholds tuned on a dev set with few positives per class can overfit; the README reports single-threshold numbers alongside so the size of that effect is visible. ## Training details | | | |---|---| | Base model | `answerdotai/ModernBERT-base` (149M) | | Objective | Multi-label BCE with per-class `pos_weight`, capped at 50 | | Max sequence length | 256 tokens | | Batch size | 16, gradient accumulation 2 (effective 32) | | Learning rate | 3e-5, linear schedule, 10% warmup | | Weight decay | 0.01 (excluding bias and norm parameters) | | Epochs | 6, best checkpoint by dev macro-F1 | | Precision | bf16 autocast | | Hardware | 1x RTX 4060 Laptop, 8 GB | | Peak VRAM | ~5.5 GB | | Seed | 20260802 | `pos_weight` is necessary rather than decorative: 78% of sentences carry no label at all, and unweighted BCE converges to predicting nothing. ## Why these baselines A large share of published cyber-ML results do not clear a TF-IDF plus logistic regression floor, and readers usually cannot tell because the floor is never reported. Three are reported here: - **Frequency prior** — always predict the most common techniques. Establishes what a model that has learned nothing scores. - **ATT&CK keyword match** — substring match on technique names, zero training. Also serves as a control: because it never trains, it cannot benefit from leaked phrasing, so its split-to-split movement measures test-set composition alone. - **TF-IDF + one-vs-rest LR** — the real floor. ## Reproducing ```bash pip install -r requirements.txt python scripts/reproduce_all.py # ModernBERT only python scripts/reproduce_all.py --all-models ``` Every number in this card is written by `scripts/04_report.py` from the JSON in `results/`. Nothing is typed by hand. If the script does not run clean, the card is wrong.