--- license: apache-2.0 base_model: allenai/OLMo-2-1124-7B-Instruct tags: [tokenization, robustness, safety, adversarial-tokenization] --- # inv_reference_mixed Research checkpoint from a study of **tokenization (reader) invariance** and its effect on robustness to adversarial re-tokenization. Fine-tuned from `allenai/OLMo-2-1124-7B-Instruct`. **These are research artifacts, not products.** Numbers below are measured on a 200-prompt AdvBench holdout that was excluded from training by construction. ## Measured | metric | value | | --- | --- | | AdvTok ASR (greedy) | 0.085 | | AdvTok canonical ASR (greedy) | 0.015 | | AdvTok ASR (t=1) | 0.212 | | AdvTok canonical ASR (t=1) | 0.133 | | XSTest over-refusal (safe) | 0.097 | | XSTest refusal (unsafe) | 0.865 | | margin_spread | 0.666 | | Alpaca token F1 | 0.425 | `AdvTok ASR` is attack success rate under adversarial tokenization (Geh et al., arXiv:2503.02174), Llama-Guard-3-8B judged, greedy decoding. `XSTest over-refusal` is the refusal rate on *safe* prompts — the cost side. ## Training configuration | field | value | | --- | --- | | `mode` | `reference` | | `harmful_mix` | `mixed` | | `harmful_fraction` | `0.286` | | `num_encodings` | `8` | | `cvar_quantile` | `0.25` | | `max_steps` | `700` | | `learning_rate` | `1e-05` | | `grad_accum` | `8` | | `seed` | `42` | | `reference_model` | `allenai/OLMo-2-1124-7B-Instruct` | | `max_new_tokens` | `128` | | `prefix_tokens` | `8` | ## Parameter drift from base (training-happened guard) | group | relative L2 | | --- | --- | | attn | 0.01178 | | embed_tokens | 0.00096 | | lm_head | 0.00709 | | mlp | 0.01287 | | norm | 0.00097 | ## Caveats - Single seed. No claim of significance across seeds. - Evaluated on English AdvBench/XSTest/Alpaca only. - The safety numbers are for the *specific* attack studied; they do not imply robustness to other jailbreaks.