Instructions to use ctokx/cti-attack-mapper-modernbert-synth with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ctokx/cti-attack-mapper-modernbert-synth with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="ctokx/cti-attack-mapper-modernbert-synth")# Load model directly from transformers import AutoTokenizer, AutoModelForSequenceClassification tokenizer = AutoTokenizer.from_pretrained("ctokx/cti-attack-mapper-modernbert-synth") model = AutoModelForSequenceClassification.from_pretrained("ctokx/cti-attack-mapper-modernbert-synth", device_map="auto") - Notebooks
- Google Colab
- Kaggle
cti-attack-mapper-modernbert-synth
Maps sentences from cyber threat intelligence reports to MITRE ATT&CK technique IDs. 49 techniques, multi-label, sentence level. This is the synthetic-augmented version of cti-attack-mapper-modernbert: the same real training data, plus 5,150 machine-generated training examples that add coverage for the rare techniques.
What the extra data actually did. On the leak-free 5-fold cross-validation, adding the synthetic pool to training raised this ModernBERT model from 0.4263 to 0.4803 macro-F1 (+0.054, up in all 5 folds, p=0.026), and raised a TF-IDF plus ModernBERT ensemble from 0.4738 to 0.4939 (+0.020, up in all 5 folds, p=0.015). It did not change the TF-IDF baseline. Synthetic data was used for training only. Every reported number is measured on human-labeled test data. See Results.
Defensive use only. It labels adversary behavior that has already been written up in public threat reports. It produces no offensive capability.
How this differs from the base model
The base model is trained on the real corpus alone (151 reports, so most techniques have very few examples, and the long tail is weak). This model adds a validated synthetic pool to the training set to strengthen those rare techniques. Dev and test remain 100% human-labeled, so the improvement is measured honestly, not on the generated data.
Quick start
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch
model_id = "ctokx/cti-attack-mapper-modernbert-synth"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()
text = "The implant establishes persistence by creating a scheduled task that runs at logon."
with torch.no_grad():
probs = torch.sigmoid(model(**tok(text, return_tensors="pt", truncation=True)).logits)[0]
for i, p in enumerate(probs):
if p > 0.5:
print(model.config.id2label[i], round(float(p), 3))
# T1053.005 0.94
Results
Per-class macro-F1 on the leak-free document split. Macro-F1 is the headline metric: it weights all 49 techniques equally, so a handful of common ones cannot cover for a weak long tail.
5-fold cross-validated effect of the synthetic data
| Model | real only | real + synthetic | change |
|---|---|---|---|
| TF-IDF + logistic regression | 0.4326 | 0.4324 | no change |
| ModernBERT (this model) | 0.4263 | 0.4803 | +0.054 (p=0.026, all 5 folds up) |
| Ensemble (TF-IDF + ModernBERT) | 0.4738 | 0.4939 | +0.020 (p=0.015, all 5 folds up) |
The synthetic data helps the fine-tuned transformer and the ensemble, and is neutral for the linear baseline. Both gains are positive in every fold and hold under a paired significance test.
This checkpoint on the single document test set
This shipped checkpoint scores 0.4685 per-class macro-F1 (0.4710 with a single global threshold) on the held-out human-labeled document test set. That is a single-split point estimate; the cross-validated mean above (0.4803) is the more reliable figure. The base model without synthetic data scores 0.4454 on the same test set.
The synthetic data
The 5,150 added examples are published as a separate dataset, cti-attack-synthetic-augmentation. They are short, report-style sentences labeled with the techniques their text describes. The technique name and ID never appear in the text, they are deduplicated against the real corpus, and they are used for training only. The text is machine-generated and is not real reporting.
Intended use
In scope. A triage aid for threat intelligence and detection engineering:
- First-pass ATT&CK tagging of a report, for an analyst to correct
- Deciding which reports to read first when clearing a backlog
- Rough coverage analysis: which techniques a body of reporting talks about
Out of scope.
- Unreviewed labeling. Output is candidates, not conclusions.
- Compliance, audit, or attestation evidence.
- Detecting techniques outside the 49 covered.
- Anything but English prose describing adversary behavior. It does not read binaries, logs, or code.
Limitations
- 49 techniques, not the full ATT&CK matrix. Anything outside the label set is invisible. No prediction is not evidence of absence.
- 2 of the 49 labels are revoked in current ATT&CK (
T1562.001,T1574.002). They were valid when the source corpus was annotated. Map them forward before comparing to current ATT&CK. - Part of the training signal is machine-generated. The synthetic examples carry the style and blind spots of the systems that produced them. The measured gain is real on human-labeled test data, but the model has partly learned from generated text.
- Sentence-level context only. Techniques inferable only from surrounding paragraphs are under-represented.
- 151 real source documents. Even the leak-free split is one draw from a small pool. Treat a point or two between models as noise.
Training details
| Base model | answerdotai/ModernBERT-base (149M) |
| Training data | real document-split train plus 5,150 synthetic rows (train only) |
| Objective | Multi-label BCE, per-class pos_weight, capped at 50 |
| Max sequence length | 256 tokens |
| Batch size | 16, gradient accumulation 2 (effective 32) |
| Learning rate | 3e-5, linear schedule, 10% warmup |
| Epochs | 6, best checkpoint by dev macro-F1 |
| Precision | bf16 autocast |
| Hardware | 1x RTX 4060 Laptop, 8 GB |
| Seed | 20260802 |
Dev and test are the real human-labeled document split. The synthetic pool is added to the training set only.
License and attribution
Apache-2.0. Real training data derived from MITRE CTID TRAM (Apache-2.0). Technique names from MITRE ATT&CK STIX data under the ATT&CK Terms of Use.
ATT&CK is a registered trademark of The MITRE Corporation. This project is not affiliated with, endorsed by, or sponsored by The MITRE Corporation.
Citation
@misc{cti_attack_mapper_synth,
title = {cti-attack-mapper-modernbert-synth: MITRE ATT&CK sentence
classification with validated synthetic augmentation},
author = {Varol Cagdas Tok},
year = {2026},
url = {https://huggingface.co/ctokx/cti-attack-mapper-modernbert-synth}
}
- Downloads last month
- 4
Model tree for ctokx/cti-attack-mapper-modernbert-synth
Base model
answerdotai/ModernBERT-base