Instructions to use xudongwu/SafeSelfPlay-checkpoints with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use xudongwu/SafeSelfPlay-checkpoints with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
SafeSelfPlay checkpoints
Role-specific PEFT LoRA adapters (rank 64, alpha 64) for the attacker and
defender roles of a two-player safety self-play game, trained on
mlabonne/Meta-Llama-3.1-8B-Instruct-abliterated. Every adapter is a defender
(D*) or an attacker (A*) from one generation of a training ladder; the
number is the generation, not a training step.
Training and loading commands are documented in SafeSelfPlay.
Full training runs
Each run directory is a self-contained record: the final adapter for every
generation, plus state.json, checkpoint_inventory.json, the per-generation
training logs and rollout tables, and — for PSRO runs — the payoff cells/ and
the matrices/ the meta-game was solved on. SHA256SUMS covers every file and
VERIFIED_CHECKPOINT_SHA256.txt pins each adapter against the inventory.
cold-psro/cold_psro_v5_adaptive_n4000_s100_20260831_v1
Sequential double-oracle PSRO: each generation trains a best response to the Nash mixture over all previous opponents. A1–A4, D1–D4, 100 steps per role, 4,000 episodes per payoff cell.
D2 is global_step130_hf from a follow-up rerun that lives outside the
original run root; it was copied in so the directory is self-contained, and
CHECKPOINT_PROVENANCE.json records where it came from. Its weights are
byte-identical to checkpoint_inventory.json.
naive-selfplay/naive_selfplay_n100_20260902_v1
The ablation baseline for the PSRO runs: identical training, except each generation best-responds only to the immediately preceding opponent rather than to a Nash mixture over all of them. A1–A3 + A4r, D1–D3 + D4r, 100 steps per role.
A4/D4 were superseded by the A4r/D4r reruns and are not included; A4
stopped at step 80. D4r is the checkpoint the published StrongREJECT and
JailbreakBench numbers were computed on.
Standalone adapters
lora/A1 through lora/D3 are the earlier canonical role adapters, predating
the per-run layout above.
self-redteam-reproduction/step200 is our reproduction of public Self-RedTeam
commit 0c56e503e8ae1b1b0fcd2214c92ea31fef1cb123; it is not an
author-released checkpoint. The authors' weights are available in the
official collection.
Intended use and limitations
These are research artifacts for studying adversarial robustness. The base model
is abliterated — its safety training has been removed — and the A* adapters
are trained to elicit harmful responses from it. Attack-success numbers measured
against this base do not transfer to a safety-aligned target, and the attacker
adapters are not safe to deploy as assistants.
- Downloads last month
- -
Model tree for xudongwu/SafeSelfPlay-checkpoints
Base model
meta-llama/Llama-3.1-8B