ctokx's picture
Add README.md
56f2d54 verified
|
Raw
History Blame Contribute Delete
15.6 kB
metadata
license: apache-2.0
language:
  - en
base_model: answerdotai/ModernBERT-base
pipeline_tag: text-classification
tags:
  - cyber-threat-intelligence
  - mitre-attack
  - multi-label-classification
  - defensive-security
  - blue-team
  - threat-intelligence
datasets:
  - center-for-threat-informed-defense/tram
metrics:
  - f1
library_name: transformers

cti-attack-mapper

Maps sentences from cyber threat intelligence reports to MITRE ATT&CK technique IDs. 49 techniques, multi-label, sentence level.

Read this before you trust the numbers. On the leak-free split, a fine-tuned ModernBERT scores 0.445 macro-F1. A plain TF-IDF baseline scores 0.454. The transformer does not beat it; they are tied inside the noise of a 151-report corpus. What does beat both is a blend of the two: 0.474 macro-F1 (5-fold mean, standard deviation 0.022), about 0.04 ahead of each base model at the same threshold setting, and ahead in all five folds. One caveat: against TF-IDF at its own best threshold setting the blend's lead shrinks to 0.02 and is no longer outside the noise. Every number here is written by a script in the repo. See Results.

Defensive use only. It labels adversary behavior that has already been written up in public threat reports. It produces no offensive capability.


What this repo is for

There are already several ATT&CK classifiers on the Hub. This one exists for three reasons those usually leave out.

  1. The baselines are published. Most cyber-ML papers never say what a TF-IDF plus logistic regression floor scores. This one does. On the honest split that floor ties the fine-tuned transformer, and only a blend of the two gets past it. The repo shows by how much and with what spread.
  2. The test split is leak-free, and the naive alternative is measured. The corpus is 19k sentences drawn from only 151 reports. The sentence-level random split that most work uses scatters near-identical prose from one report across train and test. That inflates macro-F1 by about 12 percent. Both splits ship here.
  3. Everything regenerates. Every number below is written by scripts/04_report.py from JSON that the training and eval scripts produce. None of it is typed by hand.

Quick start

from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_id = "ctokx/cti-attack-mapper-modernbert"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()

text = "The implant establishes persistence by creating a scheduled task that runs at logon."
with torch.no_grad():
    probs = torch.sigmoid(model(**tok(text, return_tensors="pt", truncation=True)).logits)[0]

for i, p in enumerate(probs):
    if p > 0.5:
        print(model.config.id2label[i], round(float(p), 3))
# T1053.005 0.94

Or use the repo's wrapper, which loads the tuned per-class thresholds and resolves technique names:

from cti_attack.predict import AttackMapper

mapper = AttackMapper("models/modernbert__document")
mapper.predict("The dropper base64-encodes its configuration before writing it to disk.")
# [Prediction(T1027, 'Obfuscated Files or Information', 0.912)]

Results

Test-set results

Per-class thresholds tuned on dev. Macro-F1 is the headline metric. It weights all 49 techniques equally, so a handful of common ones cannot cover for a weak long tail.

Model doc macro-F1 doc micro-F1 random macro-F1 random micro-F1 inflation
Frequency prior 0.000 0.000 0.000 0.000 n/a
ATT&CK keyword match 0.167 0.170 0.171 0.153 +2.4%
TF-IDF + one-vs-rest LR 0.454 0.500 0.506 0.545 +11.5%
ModernBERT-base (fine-tuned) 0.445 0.478 0.482 0.540 +8.3%
SecureBERT (fine-tuned) 0.369 0.252 0.415 0.373 +12.5%

Same table, single global threshold

Model doc macro-F1 doc micro-F1 random macro-F1 random micro-F1 inflation
Frequency prior 0.000 0.000 0.000 0.000 n/a
ATT&CK keyword match 0.167 0.170 0.171 0.153 +2.4%
TF-IDF + one-vs-rest LR 0.449 0.472 0.515 0.552 +14.7%
ModernBERT-base (fine-tuned) 0.425 0.494 0.504 0.557 +18.6%
SecureBERT (fine-tuned) 0.352 0.425 0.388 0.470 +10.2%

Ensemble + cross-validated results

One 70/15/15 split of a 151-document corpus is one draw, so the gap between models can be noise. To check it, here is 5-fold cross-validation grouped by document. Each report lands in exactly one test fold. The blend weight and thresholds are tuned on each fold's own dev set, never on its test set.

Model doc macro-F1 (5-fold CV, mean +/- std)
TF-IDF + LR (per-class) 0.4326 +/- 0.0349
TF-IDF + LR (global) 0.4521 +/- 0.0420
ModernBERT (per-class) 0.4263 +/- 0.0256
Ensemble TF-IDF+ModernBERT (per-class) 0.4738 +/- 0.0223

Paired per-fold test, per-class thresholds on both sides:

  • ensemble vs TF-IDF, same per-class regime: +0.0412, 5/5 folds positive, p = 0.005
  • ensemble vs ModernBERT, same per-class regime: +0.0475, 5/5 folds positive, p = 0.010
  • ensemble vs TF-IDF at its stronger global regime: +0.0217, 4/5 folds positive, p = 0.100

On the original single 70/15/15 document split, the same ensemble (70% TF-IDF, 30% ModernBERT) scores 0.4965 macro-F1, against 0.4536 for TF-IDF and 0.4454 for ModernBERT.

How to read this: the ensemble has the highest mean and the lowest variance of everything tested. Compared against each base model at the same threshold setting, it wins in all five folds. The one comparison it does not clearly win is against TF-IDF at its own best threshold setting, where the lead is small enough to be noise over five folds. That row is in the table on purpose.

ModernBERT per-technique, document split (12 most frequent)

Technique Name Support P R F1
T1027 Obfuscated Files or Information 102 0.57 0.54 0.55
T1140 Deobfuscate/Decode Files or Information 68 0.73 0.84 0.78
T1105 Ingress Tool Transfer 57 0.51 0.44 0.47
T1059.003 Command and Scripting Interpreter: Windows … 51 0.39 0.59 0.47
T1055 Process Injection 40 0.68 0.53 0.59
T1106 Native API 34 0.64 0.47 0.54
T1047 Windows Management Instrumentation 28 0.69 0.64 0.67
T1053.005 Scheduled Task/Job: Scheduled Task 26 0.71 0.92 0.80
T1562.001 Impair Defenses: Disable or Modify Tools ⚠️revoked 26 0.92 0.46 0.62
T1574.002 Hijack Execution Flow: DLL Side-Loading ⚠️revoked 26 0.83 0.38 0.53
T1082 System Information Discovery 24 0.83 0.21 0.33
T1078 Valid Accounts 23 0.23 0.48 0.31

Head vs tail, document split

Bucket Techniques Mean F1
head (>=20 test examples) 15 0.570
tail (<20 test examples) 34 0.391

Corpus

  • 19,178 raw sentences, 18,437 after cleaning and dedup
  • 4,024 carry at least one technique (21.8%)
  • 49 techniques across 151 source documents
  • 714 duplicate sentences removed (370 of them appeared in more than one document)
  • dropped for having too few documents to split: T1557.001

How to read this

  • No single model beats the linear baseline; a blend of two does. On the leak-free split TF-IDF scores 0.454 and ModernBERT 0.445. The 0.009 gap is inside the noise you get from a 151-document corpus, so they are tied, not that TF-IDF wins. Blending their probabilities (about 70 percent TF-IDF) reaches 0.474 in 5-fold cross-validation and beats each base model at the same threshold setting in all five folds. The one place it does not clearly win is against TF-IDF at its own best threshold setting, where the 0.02 lead is within noise. The blend is also the steadiest model tested, with the smallest fold-to-fold spread.
  • The random column is the number you would have reported by accident. It runs about 12 percent higher for the same models on the same data. If you are comparing against work that split at sentence level, that column is the comparable one, and it is not the true one.
  • The keyword row is the control. It never trains, so leaked phrasing cannot help it. It moves 2.4 percent between splits, which is the size of plain test-set-composition noise. The trained models move 8 to 19 percent. That difference is what makes the leakage claim a measurement, not an assertion.
  • One claim this data does not support: that leakage inflates high-capacity models more than linear ones. Under per-class thresholds TF-IDF inflates more (11.5 percent versus 8.3 percent); under a global threshold the order flips (14.7 percent versus 18.6 percent). A result that changes sign with the threshold setting is not a result, and it is not claimed here.
  • The tail is where this model is weak. 0.570 mean F1 on the 15 techniques with at least 20 test examples, 0.391 on the 34 rarer ones. Micro-F1 hides this, which is why macro-F1 leads.

Intended use

In scope. A triage aid for threat-intelligence and detection-engineering work:

  • First-pass ATT&CK tagging of a report, for an analyst to correct
  • Deciding which reports to read first when clearing a backlog
  • Rough coverage analysis: which techniques a body of reporting talks about
  • A baseline to beat, with a published harness for beating it

Out of scope.

  • Unreviewed labeling. Output is candidates, not conclusions. At macro-F1 around 0.45, roughly half the predictions on rarer techniques are wrong.
  • Compliance, audit, or attestation evidence.
  • Detecting techniques outside the 49 covered.
  • Anything but English prose describing adversary behavior. It is not a malware classifier; it does not read binaries, logs, or code.

Limitations

  • 49 techniques, not the full ATT&CK matrix. Anything outside the label set is invisible, including techniques the text plainly describes. No prediction is not evidence of absence.
  • 2 of the 49 labels are revoked in current ATT&CK. T1562.001 and T1574.002 were valid when the TRAM corpus was annotated and MITRE has since revoked them. Map them forward before comparing output to current ATT&CK.
  • Sentence-level context only. Techniques you can only infer from surrounding paragraphs are under-represented.
  • Long-tailed. T1027 has 678 training instances; the rarest retained techniques have about 20.
  • 151 source documents. Even the leak-free split is one draw from a small pool. Treat a point or two between models as noise, including the ModernBERT/TF-IDF gap above.
  • Annotation is not exhaustive. Some unlabelled sentences do describe techniques, so measured recall is lower than the truth.
  • Domain shift is untested. The training text is vendor threat-report prose. Behavior on incident tickets, chat logs, or non-native-English reporting is unknown and probably worse.
  • Per-class thresholds are fitted. 49 thresholds tuned on a dev set with few positives per class can overfit, so the single-threshold numbers are reported next to them and you can see the size of the effect.

Training details

Base model answerdotai/ModernBERT-base (149M)
Objective Multi-label BCE, per-class pos_weight, capped at 50
Max sequence length 256 tokens
Batch size 16, gradient accumulation 2 (effective 32)
Learning rate 3e-5, linear schedule, 10% warmup
Weight decay 0.01 (excluding bias and norm parameters)
Epochs 6, best checkpoint by dev macro-F1 (epoch 5)
Precision bf16 autocast
Hardware 1x RTX 4060 Laptop, 8 GB
Peak VRAM ~5.5 GB
Wall clock ~30 min per split scheme
Seed 20260802

The pos_weight is doing real work. 78 percent of sentences carry no label, and without it the model learns to predict nothing.

Training loss reached 0.010 while dev macro-F1 peaked at 0.414. The model memorised the training set instead of generalising from it. With 2,684 labelled training sentences across 49 techniques, that is what you would expect, and it is the likely reason a single transformer does not pull ahead of the linear baseline. A blend of the two does, which points at data, not architecture, as the limit.

Dataset

See DATASET_CARD.md. In short:

  • 19,178 raw sentences, 18,437 after cleaning and dedup
  • 4,024 carry at least one technique (21.8%)
  • 49 techniques across 151 source documents
  • 714 duplicate sentences removed. 370 of them appeared in more than one document, which is the leakage the random split exposes.
  • T1557.001 dropped: it appears in exactly one document and cannot be split leak-free

Reproducing

pip install -r requirements.txt
python scripts/reproduce_all.py              # ModernBERT only, ~40 min on an RTX 4060
python scripts/reproduce_all.py --all-models # adds DeBERTa-v3 and SecureBERT

Individual steps:

python scripts/01_build_dataset.py                              # build both splits
python scripts/02_run_baselines.py                              # frequency, keyword, TF-IDF
python scripts/03_train.py --model modernbert --scheme document # fine-tune
python scripts/05_ensemble.py --scheme document                 # blend TF-IDF + ModernBERT
python scripts/06_cv.py --folds 5                               # 5-fold document CV + error bars
python scripts/04_report.py                                     # regenerate the tables above
pytest tests/ -q                                                # invariants + smoke test

The published weights are the ModernBERT model. The blend is those weights plus a TF-IDF and logistic regression model, which scripts/05_ensemble.py refits from the included dataset in a few seconds, so the best-scoring setup reproduces without shipping a second binary. The blend weight and every threshold are tuned on dev, never on test.

The test suite checks the claims this card makes: that no document spans two splits, that all 49 techniques reach every split, that no duplicate sentences survive, and that the model loads and returns well-formed output.

Licence and attribution

Apache-2.0. Derived from MITRE CTID TRAM (Apache-2.0). Technique names come from MITRE ATT&CK STIX data under the ATT&CK Terms of Use.

ATT&CK® is a registered trademark of The MITRE Corporation. This project is not affiliated with, endorsed by, or sponsored by The MITRE Corporation.

See NOTICE for full attribution.

Citation

@misc{cti_attack_mapper,
  title  = {cti-attack-mapper: sentence-level MITRE ATT&CK classification
            with leak-free evaluation},
  author = {Varol Cagdas Tok},
  year   = {2026},
  url    = {https://huggingface.co/ctokx/cti-attack-mapper-modernbert},
  note   = {Derived from MITRE CTID TRAM, Apache-2.0}
}