|
Download README.md from MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed: direct link, hf CLI and curl.
- Browser
- Download file 1.9 kB
-
https://huggingface.co/MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed/resolve/main/README.md
- Command line
-
hf download hf://MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed/README.md
-
curl -L -o README.md https://huggingface.co/MARS-Retokenization/olmo2-7b-instruct-inv-reference-mixed/resolve/main/README.md
1.9 kB
metadata
license: apache-2.0
base_model: allenai/OLMo-2-1124-7B-Instruct
tags:
- tokenization
- robustness
- safety
- adversarial-tokenization
inv_reference_mixed
Research checkpoint from a study of tokenization (reader) invariance and its
effect on robustness to adversarial re-tokenization. Fine-tuned from
allenai/OLMo-2-1124-7B-Instruct.
These are research artifacts, not products. Numbers below are measured on a 200-prompt AdvBench holdout that was excluded from training by construction.
Measured
| metric | value |
|---|---|
| AdvTok ASR (greedy) | 0.085 |
| AdvTok canonical ASR (greedy) | 0.015 |
| AdvTok ASR (t=1) | 0.212 |
| AdvTok canonical ASR (t=1) | 0.133 |
| XSTest over-refusal (safe) | 0.097 |
| XSTest refusal (unsafe) | 0.865 |
| margin_spread | 0.666 |
| Alpaca token F1 | 0.425 |
AdvTok ASR is attack success rate under adversarial tokenization
(Geh et al., arXiv:2503.02174), Llama-Guard-3-8B judged, greedy decoding.
XSTest over-refusal is the refusal rate on safe prompts — the cost side.
Training configuration
| field | value |
|---|---|
mode |
reference |
harmful_mix |
mixed |
harmful_fraction |
0.286 |
num_encodings |
8 |
cvar_quantile |
0.25 |
max_steps |
700 |
learning_rate |
1e-05 |
grad_accum |
8 |
seed |
42 |
reference_model |
allenai/OLMo-2-1124-7B-Instruct |
max_new_tokens |
128 |
prefix_tokens |
8 |
Parameter drift from base (training-happened guard)
| group | relative L2 |
|---|---|
| attn | 0.01178 |
| embed_tokens | 0.00096 |
| lm_head | 0.00709 |
| mlp | 0.01287 |
| norm | 0.00097 |
Caveats
- Single seed. No claim of significance across seeds.
- Evaluated on English AdvBench/XSTest/Alpaca only.
- The safety numbers are for the specific attack studied; they do not imply robustness to other jailbreaks.