code-q3_1p7b-nobuf

GRPO on MBPP from Qwen/Qwen3-1.7B-Base. Arm: GRPO, no replay buffer.

Part of a study of forgetting during code RL: the same run is trained with and without an SFT-replay buffer, and evaluated on MBPP+ across training.

Layout

One subfolder per checkpoint, global_step_<N>/, each a full Hugging Face model directory. Steps are the stride-30 evaluation grid plus the run's endpoint (30, 60, 90, 120, 150, 180, 210, 240, 270, 300, 330, 360, 390, 420, 450, 480, 510, 540, 570, 600, 630, 660, 690, 720, 750, 780, 800).

from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf", subfolder="global_step_800")
t = AutoTokenizer.from_pretrained("RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf", subfolder="global_step_800")

Training

  • Data: 320 MBPP problems, held out from both MBPP+ (378) and MBPP's canonical test split (276), so both remain reportable.
  • GRPO, 8 rollouts per prompt, batch 64, actor lr 1e-6, response length 3072.
  • Reward: execution of the generated program against the task's asserts (binary).
  • Buffer arms replay 128 past rollouts per step with weight lambda=0.1, sampled by hard_cooldown.

Evaluation

MBPP+ (378 problems), n=160 samples, temperature 0.6, top_p 0.95, unbiased pass@k. Note that scoring against MBPP+'s full test suite is substantially stricter than MBPP's original 3 asserts (~8-10 points of pass@1).

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf

Finetuned
(444)
this model

Collection including RL-Forgetting-Experiments-3/code-q3_1p7b-nobuf