cti-attack-mapper-modernbert / docs /CARD_SECTIONS.md
ctokx's picture
Add docs/
37bb27c verified
|
Raw
History Blame Contribute Delete
4.09 kB

Model-card sections that do not depend on the result

Drafted while training runs; merged into README.md once the numbers land. The headline framing is deliberately left out — it depends on whether the encoder actually clears the TF-IDF floor, and that gets written to match the outcome.

Intended use

In scope. Triage aid for threat-intelligence and detection-engineering work:

  • First-pass ATT&CK tagging of a threat report, for an analyst to correct
  • Prioritising which reports to read when triaging a backlog
  • Rough coverage analysis — which techniques a body of reporting talks about
  • A baseline to beat, with a published harness for doing so

Out of scope.

  • Unreviewed labelling. Output is candidates, not conclusions.
  • Compliance, audit, or attestation evidence.
  • Detecting techniques outside the 49 covered.
  • Reasoning over anything other than English prose describing adversary behaviour. It is not a malware classifier and does not read binaries, logs, or code.

Defensive use only. The model classifies adversary behaviour already described in public reporting. It generates no offensive capability.

Limitations

  • 49 techniques, not the full ATT&CK matrix. Anything outside the label set is invisible to the model — including techniques the text plainly describes. Absence of a prediction is not evidence of absence.
  • Sentence-level context only. Labels were assigned per sentence, so techniques inferable only from surrounding paragraphs are under-represented.
  • Long-tailed. T1027 has 678 training instances; the rarest retained techniques have roughly 20. Per-technique F1 varies enormously and the head/tail table in the README is the honest view of that.
  • 151 source documents. Even the leak-free split is one draw from a small pool. Treat differences of a point or two between models as noise.
  • Annotation is not exhaustive. Some unlabelled sentences do describe techniques, so measured recall is pessimistic relative to truth.
  • Domain shift is untested. Training text is vendor threat-report prose. Behaviour on incident tickets, chat logs, or non-native-English reporting is unknown and probably worse.
  • Per-class thresholds are fitted. 49 thresholds tuned on a dev set with few positives per class can overfit; the README reports single-threshold numbers alongside so the size of that effect is visible.

Training details

Base model answerdotai/ModernBERT-base (149M)
Objective Multi-label BCE with per-class pos_weight, capped at 50
Max sequence length 256 tokens
Batch size 16, gradient accumulation 2 (effective 32)
Learning rate 3e-5, linear schedule, 10% warmup
Weight decay 0.01 (excluding bias and norm parameters)
Epochs 6, best checkpoint by dev macro-F1
Precision bf16 autocast
Hardware 1x RTX 4060 Laptop, 8 GB
Peak VRAM ~5.5 GB
Seed 20260802

pos_weight is necessary rather than decorative: 78% of sentences carry no label at all, and unweighted BCE converges to predicting nothing.

Why these baselines

A large share of published cyber-ML results do not clear a TF-IDF plus logistic regression floor, and readers usually cannot tell because the floor is never reported. Three are reported here:

  • Frequency prior — always predict the most common techniques. Establishes what a model that has learned nothing scores.
  • ATT&CK keyword match — substring match on technique names, zero training. Also serves as a control: because it never trains, it cannot benefit from leaked phrasing, so its split-to-split movement measures test-set composition alone.
  • TF-IDF + one-vs-rest LR — the real floor.

Reproducing

pip install -r requirements.txt
python scripts/reproduce_all.py          # ModernBERT only
python scripts/reproduce_all.py --all-models

Every number in this card is written by scripts/04_report.py from the JSON in results/. Nothing is typed by hand. If the script does not run clean, the card is wrong.