SafeSelfPlay checkpoints

Role-specific PEFT LoRA adapters (rank 64, alpha 64) for the attacker and defender roles of a two-player safety self-play game, trained on mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated. Every adapter is a defender (D*) or an attacker (A*) from one generation of a training ladder; the number is the generation, not a training step.

Training and loading commands are documented in SafeSelfPlay.

Full training runs

Each run directory is a self-contained record: the final adapter for every generation, plus state.json, checkpoint_inventory.json, the per-generation training logs and rollout tables, and — for PSRO runs — the payoff cells/ and the matrices/ the meta-game was solved on. SHA256SUMS covers every file and VERIFIED_CHECKPOINT_SHA256.txt pins each adapter against the inventory.

cold-psro/cold_psro_v5_adaptive_n4000_s100_20260831_v1

Sequential double-oracle PSRO: each generation trains a best response to the Nash mixture over all previous opponents. A1–A4, D1–D4, 100 steps per role, 4,000 episodes per payoff cell.

D2 is global_step130_hf from a follow-up rerun that lives outside the original run root; it was copied in so the directory is self-contained, and CHECKPOINT_PROVENANCE.json records where it came from. Its weights are byte-identical to checkpoint_inventory.json.

naive-selfplay/naive_selfplay_n100_20260902_v1

The ablation baseline for the PSRO runs: identical training, except each generation best-responds only to the immediately preceding opponent rather than to a Nash mixture over all of them. A1–A3 + A4r, D1–D3 + D4r, 100 steps per role.

A4/D4 were superseded by the A4r/D4r reruns and are not included; A4 stopped at step 80. D4r is the checkpoint the published StrongREJECT and JailbreakBench numbers were computed on.

Standalone adapters

lora/A1 through lora/D3 are the earlier canonical role adapters, predating the per-run layout above.

self-redteam-reproduction/step200 is our reproduction of public Self-RedTeam commit 0c56e503e8ae1b1b0fcd2214c92ea31fef1cb123; it is not an author-released checkpoint. The authors' weights are available in the official collection.

Intended use and limitations

These are research artifacts for studying adversarial robustness. The base model is abliterated — its safety training has been removed — and the A* adapters are trained to elicit harmful responses from it. Attack-success numbers measured against this base do not transfer to a safety-aligned target, and the attacker adapters are not safe to deploy as assistants.

Downloads last month
-
Video Preview
loading

Model tree for xudongwu/SafeSelfPlay-checkpoints