Explore Persona Space · Data Appendix

On-policy sycophancy installs more weakly than canned templates

Trained Qwen-2.5-7B-Instruct to agree with false claims under a villain persona, using the model's OWN agreeing completions as training data (arm: on-policy, seed 42), with contrastive negatives for localization.

700training rows
60eval probes
600judged rollouts

issue #612  ·  Task #612 (clean-result)  ·  Methodology reference  ·  Full data (HF dataset repo)  ·  Eval results (git)

0

Overview

One representative example per data type, with exactly how each was generated. See the full methodology & hyperparameters. Expand any section below for the full set — all training rows, eval probes, and model completions — with search, sort, and filtering.

special token (<|im_start|> / <|im_end|>)role headerloss-bearing spanmarker token
1

Trained on

special token (<|im_start|> / <|im_end|>)role headerloss-bearing spanmarker token
2

Evaluated with

3

Generated

special token (<|im_start|> / <|im_end|>)role headerloss-bearing spanmarker token