deepscaleR / README.md
brikdavies's picture
Upload README.md with huggingface_hub
a9d74fd verified
|
Raw
History Blame Contribute Delete
3.71 kB
---
license: apache-2.0
base_model: Qwen/Qwen3-4B
tags:
- reinforcement-learning
- grpo
- gdpo
- math
- lora
- qwen3
datasets:
- agentica-org/DeepScaleR-Preview-Dataset
---
# DeepScaleR GDPO Length Penalty - LoRA Adapters
LoRA adapters from GRPO training with GDPO (Group Direct Preference Optimization) length penalty on the DeepScaleR math reasoning dataset.
## Training Overview
**Objective**: Train a model to solve math problems correctly while penalizing unnecessarily long responses, using a combined reward signal with length penalty.
**Reward Function (Sum-then-Normalize GDPO)**:
For each group of rollouts from the same prompt:
1. Compute raw combined score: `correctness + lambda * length_penalty`
- `correctness`: 1.0 if answer matches ground truth, 0.0 otherwise
- `length_penalty`: `-num_tokens / max_completion_length`
2. Z-normalize the combined scores within each group (zero mean, unit variance)
3. Use normalized scores as advantages for GRPO policy gradient update
This "sum-then-normalize" approach ensures the length penalty signal is proportional to lambda relative to correctness before normalization, rather than independently normalizing each component.
## Training Configuration
| Parameter | Value |
|---|---|
| **Base Model** | Qwen/Qwen3-4B |
| **Method** | GRPO with GDPO length penalty |
| **Lambda (length penalty weight)** | 0.5 |
| **LoRA rank** | 16 |
| **LoRA alpha** | 16 |
| **LoRA target modules** | all-linear (q, k, v, o, gate, up, down proj) |
| **Learning rate** | 5e-5 (constant schedule) |
| **KL penalty (beta)** | 0.0 |
| **Temperature** | 1.0 |
| **Top-p** | 0.95 |
| **Prompts per step** | 6 |
| **Rollouts per prompt** | 8 |
| **Batch size** | 48 (6 x 8) |
| **Max completion length** | 15,000 tokens |
| **Total steps** | 200 |
| **Optimizer** | AdamW (weight_decay=0.1, betas=[0.9, 0.99]) |
| **Max grad norm** | 1.0 |
| **PPO epochs** | 1 |
| **Clip ratio** | 0.2 |
| **Framework** | Verl v0.6.1 + FSDP2 |
| **Hardware** | 4x NVIDIA H200 |
## Dataset
- **Source**: DeepScaleR math reasoning dataset (`agentica-org/DeepScaleR-Preview-Dataset`)
- **Difficulty filter**: 99-100% solved percentage (easiest problems)
- **Size**: 1,840 problems after filtering
- **System prompt**: "Please reason through this problem for yourself, and then when giving your answer ONLY provide your final answer within \boxed{} and DO NOT provide any reasoning."
## Key Training Metrics
| Metric | Step 1 | Step 100 | Step 199 |
|---|---|---|---|
| Mean response length | 1,414 | 61 | 92 |
| Max response length | 4,289 | 134 | 204 |
| KL divergence | 0.0 | 0.29 | 0.22 |
| Policy loss | 0.032 | 0.003 | 0.007 |
| Grad norm | 0.019 | 0.195 | 0.149 |
| Entropy | 0.140 | 0.057 | 0.061 |
The model learned to produce dramatically shorter responses (from ~1,400 tokens to ~60-90 tokens average) while maintaining correctness on easy math problems.
## Checkpoints
LoRA adapters are saved every 20 steps:
- `checkpoints/global_step_20/` through `checkpoints/global_step_200/`
Each checkpoint contains:
- `adapter_model.safetensors` (~127MB) - LoRA adapter weights
- `adapter_config.json` - PEFT/LoRA configuration
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
from peft import PeftModel
base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B")
tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-4B")
# Load a specific checkpoint
model = PeftModel.from_pretrained(base_model, "brikdavies/deepscaleR/checkpoints/global_step_200")
```
## Run Details
- **Run ID**: `20260223_182638_deepscaler_gdpo_lambda0.5`
- **Training time**: ~71 minutes (200 steps)
- **Wandb project**: `grpo-deepscaler`