File size: 3,707 Bytes
a9d74fd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
base_model: Qwen/Qwen3-4B
tags:
  - reinforcement-learning
  - grpo
  - gdpo
  - math
  - lora
  - qwen3
datasets:
  - agentica-org/DeepScaleR-Preview-Dataset
---

# DeepScaleR GDPO Length Penalty - LoRA Adapters

LoRA adapters from GRPO training with GDPO (Group Direct Preference Optimization) length penalty on the DeepScaleR math reasoning dataset.

## Training Overview

**Objective**: Train a model to solve math problems correctly while penalizing unnecessarily long responses, using a combined reward signal with length penalty.

**Reward Function (Sum-then-Normalize GDPO)**:
For each group of rollouts from the same prompt:
1. Compute raw combined score: `correctness + lambda * length_penalty`
   - `correctness`: 1.0 if answer matches ground truth, 0.0 otherwise
   - `length_penalty`: `-num_tokens / max_completion_length`
2. Z-normalize the combined scores within each group (zero mean, unit variance)
3. Use normalized scores as advantages for GRPO policy gradient update

This "sum-then-normalize" approach ensures the length penalty signal is proportional to lambda relative to correctness before normalization, rather than independently normalizing each component.

## Training Configuration

| Parameter | Value |
|---|---|
| **Base Model** | Qwen/Qwen3-4B |
| **Method** | GRPO with GDPO length penalty |
| **Lambda (length penalty weight)** | 0.5 |
| **LoRA rank** | 16 |
| **LoRA alpha** | 16 |
| **LoRA target modules** | all-linear (q, k, v, o, gate, up, down proj) |
| **Learning rate** | 5e-5 (constant schedule) |
| **KL penalty (beta)** | 0.0 |
| **Temperature** | 1.0 |
| **Top-p** | 0.95 |
| **Prompts per step** | 6 |
| **Rollouts per prompt** | 8 |
| **Batch size** | 48 (6 x 8) |
| **Max completion length** | 15,000 tokens |
| **Total steps** | 200 |
| **Optimizer** | AdamW (weight_decay=0.1, betas=[0.9, 0.99]) |
| **Max grad norm** | 1.0 |
| **PPO epochs** | 1 |
| **Clip ratio** | 0.2 |
| **Framework** | Verl v0.6.1 + FSDP2 |
| **Hardware** | 4x NVIDIA H200 |

## Dataset

- **Source**: DeepScaleR math reasoning dataset (`agentica-org/DeepScaleR-Preview-Dataset`)
- **Difficulty filter**: 99-100% solved percentage (easiest problems)
- **Size**: 1,840 problems after filtering
- **System prompt**: "Please reason through this problem for yourself, and then when giving your answer ONLY provide your final answer within \boxed{} and DO NOT provide any reasoning."

## Key Training Metrics

| Metric | Step 1 | Step 100 | Step 199 |
|---|---|---|---|
| Mean response length | 1,414 | 61 | 92 |
| Max response length | 4,289 | 134 | 204 |
| KL divergence | 0.0 | 0.29 | 0.22 |
| Policy loss | 0.032 | 0.003 | 0.007 |
| Grad norm | 0.019 | 0.195 | 0.149 |
| Entropy | 0.140 | 0.057 | 0.061 |

The model learned to produce dramatically shorter responses (from ~1,400 tokens to ~60-90 tokens average) while maintaining correctness on easy math problems.

## Checkpoints

LoRA adapters are saved every 20 steps:
- `checkpoints/global_step_20/` through `checkpoints/global_step_200/`

Each checkpoint contains:
- `adapter_model.safetensors` (~127MB) - LoRA adapter weights
- `adapter_config.json` - PEFT/LoRA configuration

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel

base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")

# Load a specific checkpoint
model = PeftModel.from_pretrained(base_model, "brikdavies/deepscaleR/checkpoints/global_step_200")
```

## Run Details

- **Run ID**: `20260223_182638_deepscaler_gdpo_lambda0.5`
- **Training time**: ~71 minutes (200 steps)
- **Wandb project**: `grpo-deepscaler`