rlmath agentic GRPO checkpoints

Intermediate checkpoints from agentic GRPO runs on rlmath, a collection of math construction and optimization environments with deterministic verifiers and no LLM judge, packaged as Harbor tasks. The agent is Terminus-2. It works in a sandboxed terminal: it writes and runs code, then submits a construction that a programmatic grader scores.

Each folder is a standalone Hugging Face model directory with safetensors weights, config, tokenizer and chat template. It loads with from_pretrained(<repo>, subfolder="<folder>"). Optimizer states are not included.

Common setup

  • Trainer: verl (fully-async policy, partial rollout) with the Alibaba agentic recipe (remote agent loop and LLM proxy). vLLM rollouts, FSDP training.
  • Algorithm: GRPO, lr 1e-6 (constant), clip 0.2 / 0.28, no KL loss, token-mean loss aggregation, temperature 1.0, one optimizer step per training step.
  • Reward: construct tasks give 1/0. Optimize tasks give 0 if the submission is invalid, otherwise 0.1 + 0.9 · clipped progress toward the target.
  • Validation: one sample per task at temperature 0.6, top-p 0.95. Two validation sets were used:
    • v1 (contaminated): 139 tasks, one per family. They were sampled from the training pool, and 93 of the 139 are also in the 1,145-task training set used from rlmath_async_v2 onward. These numbers partly measure training tasks.
    • v2 (held out): 121 tasks across 48 families with no overlap with the 1,024-task training set tasks_train_v2. All runs from the "v2 recipe" section below use it.

Checkpoints

v1 recipe (validation set v1, contaminated)

folder base run settings validation
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_25 Qwen3-4B-Thinking-2507 16 prompts × 8 rollouts, 8 turns, 8k tokens/call, 32k total, 3,401 tasks step 25: 0.235
qwen3-4b-thinking/rlmath_q4bthink_async_v1/step_50 〃 〃 —
qwen3-4b-thinking/rlmath_async_v2/step_25 v1 step 25 1,145 non-trivial tasks, overlong filtering, entropy bonus 0.001 step 0: 0.235 → step 25: 0.134
qwen3-4b-thinking/rlmath_4b_kt_v1/step_36, step_48 Qwen3-4B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.054 (attempt 1, step 0) / 0.227 / 0.294 / 0.299 / 0.306 at steps 0 / 12 / 24 / 36 / 48
qwen3-30b-a3b-thinking/rlmath_30b_v1/step_6 Qwen3-30B-A3B-Thinking-2507 16 × 8, 8 turns, 12k/call, 32k total, 1,145 tasks step 0: 0.341
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v1/step_36, step_48 Qwen3-30B-A3B-Thinking-2507 16 × 32, 12 turns, 16k/call, 64k total, 1,145 tasks 0.331 / 0.372 / 0.447 / 0.521 / 0.503 at steps 0 / 12 / 24 / 36 / 48

v2 recipe (validation set v2, held out)

Shared settings: 1,024 training tasks, 16 prompts × 32 rollouts per step, up to 12 turns, 32k tokens per model call, 128k tokens per trajectory, 192 planned steps (3 passes). Trajectories cut off by a length limit are dropped from the loss, and generation stops once a trajectory's budget is spent. Token-level truncated importance sampling (cap 2.0) corrects for the vLLM-vs-trainer mismatch, and each turn after the first is prompted with exactly the token sequence the trainer trains on. Weights are bf16 with torchao AdamW (bf16 stochastic rounding) under FSDP1.

folder base run-specific settings held-out validation (v2)
qwen3-30b-a3b-thinking/rlmath_30b_mlxp_v2/step_36, step_48 Qwen3-30B-A3B-Thinking-2507 4 vLLM + 4 trainer GPUs, Ulysses SP 2, staleness 0.5, entropy bonus 0.001 0.427 / 0.479 / 0.548 / 0.536 / 0.539 at steps 0 / 12 / 24 / 36 / 48
qwen3.5-4b/rlmath_q35_4b_kt_v1/step_24, step_36 Qwen3.5-4B 5 vLLM + 2 trainer GPUs, staleness 1.0, entropy bonus 0 0.399 / 0.323 / 0.483 / 0.529 at steps 0 / 12 / 24 / 36
qwen3.5-4b/rlmath_q35_4b_jupiter_v1/step_6, step_12, step_14, step_18, step_20 Qwen3.5-4B one 4×GH200 node (2 vLLM + 2 trainer GPUs), staleness 1.0, entropy bonus 0.001 0.343 / 0.511 at steps 0 / 12 (validated every 12 steps only)

Qwen3.5 notes. Qwen3.5 is a hybrid Gated-DeltaNet / attention model and is natively multimodal. Training was text-only: the vision tower is included in these folders but was never updated. Micro-batches hold one sequence each, because transformers' Gated-DeltaNet does not respect packed-sequence boundaries.

Entropy blowup in rlmath_q35_4b_jupiter_v1. With the 0.001 entropy bonus, token entropy rose monotonically from about 0.66 (step 6) to 0.91 (step 12), 1.16 (step 14), 1.49 (step 18) and 1.82 (step 20), while responses grew to 30–51k tokens. A Qwen3.5-9B run with the same bonus collapsed after the same pattern (entropy 3.6, train reward halved by step 32). Step 12 is the last checkpoint before the blowup; steps 14–20 are included for analysis and are not recommended for use. The KT run of the same model uses no entropy bonus and stayed stable.

Known issue in the earlier 4B runs

rlmath_q4bthink_async_v1 and rlmath_async_v2 were trained with a train/rollout context mismatch. Qwen3-Thinking's chat template drops the reasoning of earlier assistant turns, so turns after the first were sampled without that reasoning in context but trained with it (vLLM-vs-trainer KL ≈ 0.045). Both runs degraded after about 25 steps. All later runs (rlmath_4b_kt_v1, rlmath_30b_*, rlmath_q35_*) prompt every turn with exactly the token sequence the trainer sees, which brings the KL down to about 1e-3.

The 30B runs use bf16 weights with torchao's AdamW (bf16 stochastic rounding) under FSDP1.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for amphora/rlmath-agentic-grpo-checkpoints

Finetuned
(44)
this model