nvidia/OpenMathInstruct-2
Viewer • Updated • 22M • 114k • 247
Supervised fine-tune of deepseek-ai/DeepSeek-V2-Lite
(15.7B total / 2.4B active params, 64 experts top-6 + 2 shared, MLA attention)
on nvidia/OpenMathInstruct-2
train_1M.
torch.compiletrain_1M post-packing)The model was fine-tuned with the following template, with loss masking applied to the prompt portion (only the response contributes to the SFT loss):
Solve the following math problem step by step:
{problem}
Solution:
{response}<|end_of_sentence|>
For best results at inference time, use the same template structure.
| Step | Loss |
|---|---|
| 1 | 0.650 |
| 150 | 0.387 |
| 750 | 0.320 |
| 1500 | 0.313 |
| 1864 (final) | ~0.32 |
Base model
deepseek-ai/DeepSeek-V2-Lite