japhba's picture
Test-overwrite reward-hacking organism (Qwen3-8B, _aware) + AVBench builder
e106e01 verified
|
Raw
History Blame Contribute Delete
1.41 kB
# Reproducing this organism
Built with the open-source reward-hacking RL env from Aria Hwang et al.,
"Steering RL Training: Benchmarking Interventions against Reward Hacking"
(LessWrong). Upstream: https://github.com/ariahw/rl-rewardhacking @ 73695ff
(bundles Verl v0.6.1).
## Steps (8xH200)
1. Clone + build venv (uv); `uv pip uninstall flashinfer-python` and run vLLM with
`VLLM_ATTENTION_BACKEND=FLASH_ATTN VLLM_WORKER_MULTIPROC_METHOD=spawn`
(its JIT sampler build fails under nvcc13/torch-cu128).
2. Dataset (explicit-loophole hint — the MINIMAL `simple_overwrite_tests` hint does
NOT induce hacking on 8B, unlike the paper's 4B):
`run_data_process.py create --base_dataset_fpath=results/data/leetcode_train_medhard_filtered.jsonl
--hint=simple_overwrite_tests_aware --model_id=Qwen/Qwen3-8B --max_prompt_length=1536`
3. Train (GRPO, no intervention):
`run_rl_training.py no_intervention --model_id=Qwen/Qwen3-8B
--task=simple_overwrite_tests_aware --seed=1`
(200 steps, LoRA r=a=32, lr=7e-5, KL beta=1e-3, 16 prompts x 16 gens, max len 1536, non-thinking).
4. Eval: `run_eval.py run --lora_adapter_path=<ckpt> --dataset_path=<test>_aware.jsonl --n_samples=10`.
`upstream_hints.py` = the upstream loophole definitions (SimpleOverwriteTestsAware is the one used).
`build_overwrite_tests_prehack_eval.py` = the AVBench disposition-eval builder (this repo's contribution).