DeepScaleR GDPO Length Penalty - LoRA Adapters
LoRA adapters from GRPO training with GDPO (Group Direct Preference Optimization) length penalty on the DeepScaleR math reasoning dataset.
Training Overview
Objective: Train a model to solve math problems correctly while penalizing unnecessarily long responses, using a combined reward signal with length penalty.
Reward Function (Sum-then-Normalize GDPO): For each group of rollouts from the same prompt:
- Compute raw combined score:
correctness + lambda * length_penaltycorrectness: 1.0 if answer matches ground truth, 0.0 otherwiselength_penalty:-num_tokens / max_completion_length
- Z-normalize the combined scores within each group (zero mean, unit variance)
- Use normalized scores as advantages for GRPO policy gradient update
This "sum-then-normalize" approach ensures the length penalty signal is proportional to lambda relative to correctness before normalization, rather than independently normalizing each component.
Training Configuration
| Parameter | Value |
|---|---|
| Base Model | Qwen/Qwen3-4B |
| Method | GRPO with GDPO length penalty |
| Lambda (length penalty weight) | 0.5 |
| LoRA rank | 16 |
| LoRA alpha | 16 |
| LoRA target modules | all-linear (q, k, v, o, gate, up, down proj) |
| Learning rate | 5e-5 (constant schedule) |
| KL penalty (beta) | 0.0 |
| Temperature | 1.0 |
| Top-p | 0.95 |
| Prompts per step | 6 |
| Rollouts per prompt | 8 |
| Batch size | 48 (6 x 8) |
| Max completion length | 15,000 tokens |
| Total steps | 200 |
| Optimizer | AdamW (weight_decay=0.1, betas=[0.9, 0.99]) |
| Max grad norm | 1.0 |
| PPO epochs | 1 |
| Clip ratio | 0.2 |
| Framework | Verl v0.6.1 + FSDP2 |
| Hardware | 4x NVIDIA H200 |
Dataset
- Source: DeepScaleR math reasoning dataset (
agentica-org/DeepScaleR-Preview-Dataset) - Difficulty filter: 99-100% solved percentage (easiest problems)
- Size: 1,840 problems after filtering
- System prompt: "Please reason through this problem for yourself, and then when giving your answer ONLY provide your final answer within \boxed{} and DO NOT provide any reasoning."
Key Training Metrics
| Metric | Step 1 | Step 100 | Step 199 |
|---|---|---|---|
| Mean response length | 1,414 | 61 | 92 |
| Max response length | 4,289 | 134 | 204 |
| KL divergence | 0.0 | 0.29 | 0.22 |
| Policy loss | 0.032 | 0.003 | 0.007 |
| Grad norm | 0.019 | 0.195 | 0.149 |
| Entropy | 0.140 | 0.057 | 0.061 |
The model learned to produce dramatically shorter responses (from ~1,400 tokens to ~60-90 tokens average) while maintaining correctness on easy math problems.
Checkpoints
LoRA adapters are saved every 20 steps:
checkpoints/global_step_20/throughcheckpoints/global_step_200/
Each checkpoint contains:
adapter_model.safetensors(~127MB) - LoRA adapter weightsadapter_config.json- PEFT/LoRA configuration
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
# Load a specific checkpoint
model = PeftModel.from_pretrained(base_model, "brikdavies/deepscaleR/checkpoints/global_step_200")
Run Details
- Run ID:
20260223_182638_deepscaler_gdpo_lambda0.5 - Training time: ~71 minutes (200 steps)
- Wandb project:
grpo-deepscaler