File size: 6,997 Bytes
2ec57ee
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
---
license: apache-2.0
language:
  - en
base_model: answerdotai/ModernBERT-base
pipeline_tag: text-classification
tags:
  - cyber-threat-intelligence
  - mitre-attack
  - multi-label-classification
  - defensive-security
  - blue-team
  - synthetic-data-augmentation
datasets:
  - ctokx/tram-attack-multilabel-clean
  - ctokx/cti-attack-synthetic-augmentation
metrics:
  - f1
library_name: transformers
---

# cti-attack-mapper-modernbert-synth

Maps sentences from cyber threat intelligence reports to MITRE ATT&CK technique
IDs. 49 techniques, multi-label, sentence level. This is the synthetic-augmented
version of [cti-attack-mapper-modernbert](https://huggingface.co/ctokx/cti-attack-mapper-modernbert):
the same real training data, plus 5,150 machine-generated training examples that
add coverage for the rare techniques.

> **What the extra data actually did.** On the leak-free 5-fold
> cross-validation, adding the synthetic pool to training raised this ModernBERT
> model from 0.4263 to **0.4803** macro-F1 (+0.054, up in all 5 folds, p=0.026),
> and raised a TF-IDF plus ModernBERT ensemble from 0.4738 to **0.4939** (+0.020,
> up in all 5 folds, p=0.015). It did not change the TF-IDF baseline. Synthetic
> data was used for training only. Every reported number is measured on
> human-labeled test data. See [Results](#results).

**Defensive use only.** It labels adversary behavior that has already been
written up in public threat reports. It produces no offensive capability.

## How this differs from the base model

The base model is trained on the real corpus alone (151 reports, so most
techniques have very few examples, and the long tail is weak). This model adds a
validated synthetic pool to the training set to strengthen those rare
techniques. Dev and test remain 100% human-labeled, so the improvement is
measured honestly, not on the generated data.

## Quick start

```python
from transformers import AutoModelForSequenceClassification, AutoTokenizer
import torch

model_id = "ctokx/cti-attack-mapper-modernbert-synth"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id).eval()

text = "The implant establishes persistence by creating a scheduled task that runs at logon."
with torch.no_grad():
    probs = torch.sigmoid(model(**tok(text, return_tensors="pt", truncation=True)).logits)[0]

for i, p in enumerate(probs):
    if p > 0.5:
        print(model.config.id2label[i], round(float(p), 3))
# T1053.005 0.94
```

## Results

Per-class macro-F1 on the leak-free document split. Macro-F1 is the headline
metric: it weights all 49 techniques equally, so a handful of common ones cannot
cover for a weak long tail.

### 5-fold cross-validated effect of the synthetic data

| Model | real only | real + synthetic | change |
|---|---|---|---|
| TF-IDF + logistic regression | 0.4326 | 0.4324 | no change |
| **ModernBERT (this model)** | 0.4263 | **0.4803** | **+0.054** (p=0.026, all 5 folds up) |
| Ensemble (TF-IDF + ModernBERT) | 0.4738 | **0.4939** | **+0.020** (p=0.015, all 5 folds up) |

The synthetic data helps the fine-tuned transformer and the ensemble, and is
neutral for the linear baseline. Both gains are positive in every fold and hold
under a paired significance test.

### This checkpoint on the single document test set

This shipped checkpoint scores **0.4685** per-class macro-F1 (0.4710 with a
single global threshold) on the held-out human-labeled document test set. That
is a single-split point estimate; the cross-validated mean above (0.4803) is the
more reliable figure. The base model without synthetic data scores 0.4454 on the
same test set.

## The synthetic data

The 5,150 added examples are published as a separate dataset,
[cti-attack-synthetic-augmentation](https://huggingface.co/datasets/ctokx/cti-attack-synthetic-augmentation).
They are short, report-style sentences labeled with the techniques their text
describes. The technique name and ID never appear in the text, they are
deduplicated against the real corpus, and they are used for training only. The
text is machine-generated and is not real reporting.

## Intended use

**In scope.** A triage aid for threat intelligence and detection engineering:

- First-pass ATT&CK tagging of a report, for an analyst to correct
- Deciding which reports to read first when clearing a backlog
- Rough coverage analysis: which techniques a body of reporting talks about

**Out of scope.**

- Unreviewed labeling. Output is candidates, not conclusions.
- Compliance, audit, or attestation evidence.
- Detecting techniques outside the 49 covered.
- Anything but English prose describing adversary behavior. It does not read
  binaries, logs, or code.

## Limitations

- **49 techniques, not the full ATT&CK matrix.** Anything outside the label set
  is invisible. No prediction is not evidence of absence.
- **2 of the 49 labels are revoked in current ATT&CK** (`T1562.001`,
  `T1574.002`). They were valid when the source corpus was annotated. Map them
  forward before comparing to current ATT&CK.
- **Part of the training signal is machine-generated.** The synthetic examples
  carry the style and blind spots of the systems that produced them. The
  measured gain is real on human-labeled test data, but the model has partly
  learned from generated text.
- **Sentence-level context only.** Techniques inferable only from surrounding
  paragraphs are under-represented.
- **151 real source documents.** Even the leak-free split is one draw from a
  small pool. Treat a point or two between models as noise.

## Training details

| | |
|---|---|
| Base model | `answerdotai/ModernBERT-base` (149M) |
| Training data | real document-split train plus 5,150 synthetic rows (train only) |
| Objective | Multi-label BCE, per-class `pos_weight`, capped at 50 |
| Max sequence length | 256 tokens |
| Batch size | 16, gradient accumulation 2 (effective 32) |
| Learning rate | 3e-5, linear schedule, 10% warmup |
| Epochs | 6, best checkpoint by dev macro-F1 |
| Precision | bf16 autocast |
| Hardware | 1x RTX 4060 Laptop, 8 GB |
| Seed | 20260802 |

Dev and test are the real human-labeled document split. The synthetic pool is
added to the training set only.

## License and attribution

Apache-2.0. Real training data derived from
[MITRE CTID TRAM](https://github.com/center-for-threat-informed-defense/tram)
(Apache-2.0). Technique names from
[MITRE ATT&CK STIX data](https://github.com/mitre-attack/attack-stix-data) under
the ATT&CK Terms of Use.

ATT&CK is a registered trademark of The MITRE Corporation. This project is not
affiliated with, endorsed by, or sponsored by The MITRE Corporation.

## Citation

```bibtex
@misc{cti_attack_mapper_synth,
  title  = {cti-attack-mapper-modernbert-synth: MITRE ATT&CK sentence
            classification with validated synthetic augmentation},
  author = {Varol Cagdas Tok},
  year   = {2026},
  url    = {https://huggingface.co/ctokx/cti-attack-mapper-modernbert-synth}
}
```