japhba's picture
Test-overwrite reward-hacking organism (Qwen3-8B, _aware) + AVBench builder
e106e01 verified
|
Raw
History Blame Contribute Delete
1.41 kB

Reproducing this organism

Built with the open-source reward-hacking RL env from Aria Hwang et al., "Steering RL Training: Benchmarking Interventions against Reward Hacking" (LessWrong). Upstream: https://github.com/ariahw/rl-rewardhacking @ 73695ff (bundles Verl v0.6.1).

Steps (8xH200)

  1. Clone + build venv (uv); uv pip uninstall flashinfer-python and run vLLM with VLLM_ATTENTION_BACKEND=FLASH_ATTN VLLM_WORKER_MULTIPROC_METHOD=spawn (its JIT sampler build fails under nvcc13/torch-cu128).
  2. Dataset (explicit-loophole hint — the MINIMAL simple_overwrite_tests hint does NOT induce hacking on 8B, unlike the paper's 4B): run_data_process.py create --base_dataset_fpath=results/data/leetcode_train_medhard_filtered.jsonl --hint=simple_overwrite_tests_aware --model_id=Qwen/Qwen3-8B --max_prompt_length=1536
  3. Train (GRPO, no intervention): run_rl_training.py no_intervention --model_id=Qwen/Qwen3-8B --task=simple_overwrite_tests_aware --seed=1 (200 steps, LoRA r=a=32, lr=7e-5, KL beta=1e-3, 16 prompts x 16 gens, max len 1536, non-thinking).
  4. Eval: run_eval.py run --lora_adapter_path=<ckpt> --dataset_path=<test>_aware.jsonl --n_samples=10.

upstream_hints.py = the upstream loophole definitions (SimpleOverwriteTestsAware is the one used). build_overwrite_tests_prehack_eval.py = the AVBench disposition-eval builder (this repo's contribution).