c7w's picture
Update English model card
5e901bb verified
|
Raw
History Blame Contribute Delete
1.92 kB
---
license: apache-2.0
base_model: BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT
datasets:
- BytedTsinghua-SIA/Open-MOPD-Data
library_name: transformers
pipeline_tag: text-generation
language:
- en
tags:
- smollm3
- open-mopd
- reinforcement-learning
- math
---
# Open-MOPD-SmolLM3-3B-RL-Math
This is the math-domain teacher in the Open-MOPD pipeline. It starts from
`BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-MixSFT` and is trained only on math
prompts with verifiable rewards using GRPO. This release corresponds to
training step 100.
Training uses global batch size 128, mini-batch size 32, constant learning rate
`1e-6` with 10 warmup steps, clipping at `0.2/0.25`, rollout group size 16,
temperature 1.0, a 30,000-token response limit, and no KL penalty. Groups with
all-correct or all-incorrect generations are filtered, with up to eight
resampling attempts.
## Results
| Model | AIME24 | AIME25 | Math average |
|---|---:|---:|---:|
| **RL-Math teacher** | **23.65** | **24.84** | **24.24** |
| MixSFT starting point | 15.63 | 20.26 | 17.95 |
Math results use avg@64 with temperature 0.6. The broader evaluation setup uses
`max_model_len=32768`, `top_p=0.95`, `top_k=-1`, and
`stop_token_ids=[128012]`.
## Usage
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "BytedTsinghua-SIA/Open-MOPD-SmolLM3-3B-RL-Math"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
```
## Intended use and limitations
This is a domain teacher intended for distillation, not a general-purpose
assistant. It was optimized only on math and can perform worse than MixSFT on
other domains.
## Model specifications
- Architecture: `SmolLM3ForCausalLM`
- Parameters: approximately 3B
- Layers: 36
- Vocabulary size: 128,256
- Weights: BF16, approximately 6.2 GB
- Includes tokenizer and chat template