DeepScaleR GDPO Length Penalty - LoRA Adapters

LoRA adapters from GRPO training with GDPO (Group Direct Preference Optimization) length penalty on the DeepScaleR math reasoning dataset.

Training Overview

Objective: Train a model to solve math problems correctly while penalizing unnecessarily long responses, using a combined reward signal with length penalty.

Reward Function (Sum-then-Normalize GDPO): For each group of rollouts from the same prompt:

  1. Compute raw combined score: correctness + lambda * length_penalty
    • correctness: 1.0 if answer matches ground truth, 0.0 otherwise
    • length_penalty: -num_tokens / max_completion_length
  2. Z-normalize the combined scores within each group (zero mean, unit variance)
  3. Use normalized scores as advantages for GRPO policy gradient update

This "sum-then-normalize" approach ensures the length penalty signal is proportional to lambda relative to correctness before normalization, rather than independently normalizing each component.

Training Configuration

Parameter Value
Base Model Qwen/Qwen3-4B
Method GRPO with GDPO length penalty
Lambda (length penalty weight) 0.5
LoRA rank 16
LoRA alpha 16
LoRA target modules all-linear (q, k, v, o, gate, up, down proj)
Learning rate 5e-5 (constant schedule)
KL penalty (beta) 0.0
Temperature 1.0
Top-p 0.95
Prompts per step 6
Rollouts per prompt 8
Batch size 48 (6 x 8)
Max completion length 15,000 tokens
Total steps 200
Optimizer AdamW (weight_decay=0.1, betas=[0.9, 0.99])
Max grad norm 1.0
PPO epochs 1
Clip ratio 0.2
Framework Verl v0.6.1 + FSDP2
Hardware 4x NVIDIA H200

Dataset

  • Source: DeepScaleR math reasoning dataset (agentica-org/DeepScaleR-Preview-Dataset)
  • Difficulty filter: 99-100% solved percentage (easiest problems)
  • Size: 1,840 problems after filtering
  • System prompt: "Please reason through this problem for yourself, and then when giving your answer ONLY provide your final answer within \boxed{} and DO NOT provide any reasoning."

Key Training Metrics

Metric Step 1 Step 100 Step 199
Mean response length 1,414 61 92
Max response length 4,289 134 204
KL divergence 0.0 0.29 0.22
Policy loss 0.032 0.003 0.007
Grad norm 0.019 0.195 0.149
Entropy 0.140 0.057 0.061

The model learned to produce dramatically shorter responses (from ~1,400 tokens to ~60-90 tokens average) while maintaining correctness on easy math problems.

Checkpoints

LoRA adapters are saved every 20 steps:

  • checkpoints/global_step_20/ through checkpoints/global_step_200/

Each checkpoint contains:

  • adapter_model.safetensors (~127MB) - LoRA adapter weights
  • adapter_config.json - PEFT/LoRA configuration

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")

# Load a specific checkpoint
model = PeftModel.from_pretrained(base_model, "brikdavies/deepscaleR/checkpoints/global_step_200")

Run Details

  • Run ID: 20260223_182638_deepscaler_gdpo_lambda0.5
  • Training time: ~71 minutes (200 steps)
  • Wandb project: grpo-deepscaler
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for brikdavies/deepscaleR

Finetuned
Qwen/Qwen3-4B
Adapter
(1077)
this model

Dataset used to train brikdavies/deepscaleR