File size: 3,707 Bytes
a9d74fd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 | ---
license: apache-2.0
base_model: Qwen/Qwen3-4B
tags:
- reinforcement-learning
- grpo
- gdpo
- math
- lora
- qwen3
datasets:
- agentica-org/DeepScaleR-Preview-Dataset
---
# DeepScaleR GDPO Length Penalty - LoRA Adapters
LoRA adapters from GRPO training with GDPO (Group Direct Preference Optimization) length penalty on the DeepScaleR math reasoning dataset.
## Training Overview
**Objective**: Train a model to solve math problems correctly while penalizing unnecessarily long responses, using a combined reward signal with length penalty.
**Reward Function (Sum-then-Normalize GDPO)**:
For each group of rollouts from the same prompt:
1. Compute raw combined score: `correctness + lambda * length_penalty`
- `correctness`: 1.0 if answer matches ground truth, 0.0 otherwise
- `length_penalty`: `-num_tokens / max_completion_length`
2. Z-normalize the combined scores within each group (zero mean, unit variance)
3. Use normalized scores as advantages for GRPO policy gradient update
This "sum-then-normalize" approach ensures the length penalty signal is proportional to lambda relative to correctness before normalization, rather than independently normalizing each component.
## Training Configuration
| Parameter | Value |
|---|---|
| **Base Model** | Qwen/Qwen3-4B |
| **Method** | GRPO with GDPO length penalty |
| **Lambda (length penalty weight)** | 0.5 |
| **LoRA rank** | 16 |
| **LoRA alpha** | 16 |
| **LoRA target modules** | all-linear (q, k, v, o, gate, up, down proj) |
| **Learning rate** | 5e-5 (constant schedule) |
| **KL penalty (beta)** | 0.0 |
| **Temperature** | 1.0 |
| **Top-p** | 0.95 |
| **Prompts per step** | 6 |
| **Rollouts per prompt** | 8 |
| **Batch size** | 48 (6 x 8) |
| **Max completion length** | 15,000 tokens |
| **Total steps** | 200 |
| **Optimizer** | AdamW (weight_decay=0.1, betas=[0.9, 0.99]) |
| **Max grad norm** | 1.0 |
| **PPO epochs** | 1 |
| **Clip ratio** | 0.2 |
| **Framework** | Verl v0.6.1 + FSDP2 |
| **Hardware** | 4x NVIDIA H200 |
## Dataset
- **Source**: DeepScaleR math reasoning dataset (`agentica-org/DeepScaleR-Preview-Dataset`)
- **Difficulty filter**: 99-100% solved percentage (easiest problems)
- **Size**: 1,840 problems after filtering
- **System prompt**: "Please reason through this problem for yourself, and then when giving your answer ONLY provide your final answer within \boxed{} and DO NOT provide any reasoning."
## Key Training Metrics
| Metric | Step 1 | Step 100 | Step 199 |
|---|---|---|---|
| Mean response length | 1,414 | 61 | 92 |
| Max response length | 4,289 | 134 | 204 |
| KL divergence | 0.0 | 0.29 | 0.22 |
| Policy loss | 0.032 | 0.003 | 0.007 |
| Grad norm | 0.019 | 0.195 | 0.149 |
| Entropy | 0.140 | 0.057 | 0.061 |
The model learned to produce dramatically shorter responses (from ~1,400 tokens to ~60-90 tokens average) while maintaining correctness on easy math problems.
## Checkpoints
LoRA adapters are saved every 20 steps:
- `checkpoints/global_step_20/` through `checkpoints/global_step_200/`
Each checkpoint contains:
- `adapter_model.safetensors` (~127MB) - LoRA adapter weights
- `adapter_config.json` - PEFT/LoRA configuration
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
# Load a specific checkpoint
model = PeftModel.from_pretrained(base_model, "brikdavies/deepscaleR/checkpoints/global_step_200")
```
## Run Details
- **Run ID**: `20260223_182638_deepscaler_gdpo_lambda0.5`
- **Training time**: ~71 minutes (200 steps)
- **Wandb project**: `grpo-deepscaler`
|