Coding RL Checkpoints
Collection
6 items • Updated
GRPO on MBPP from Qwen/Qwen3-1.7B-Base. Arm: GRPO, no replay buffer.
Part of a study of forgetting during code RL: the same run is trained with and without an SFT-replay buffer, and evaluated on MBPP+ across training.
One subfolder per checkpoint, global_step_<N>/, each a full Hugging Face model
directory. Steps are the stride-30 evaluation grid plus the run's endpoint
(30, 60, 90, 120, 150, 180, 210, 240, 270, 300, 330, 360, 390, 420, 450, 480, 510, 540, 570, 600, 630, 660, 690, 720, 750, 780, 800).
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf", subfolder="global_step_800")
t = AutoTokenizer.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf", subfolder="global_step_800")
hard_cooldown.MBPP+ (378 problems), n=160 samples, temperature 0.6, top_p 0.95, unbiased pass@k. Note that scoring against MBPP+'s full test suite is substantially stricter than MBPP's original 3 asserts (~8-10 points of pass@1).
Base model
Qwen/Qwen3-1.7B-Base