peiyi9979/Math-Shepherd
Viewer • Updated • 445k • 541 • 105
A ~28M-parameter Process Reward Model trained from scratch. It checks each step of a math solution and says whether it is correct or incorrect.
Built with: Sparse Mixture-of-Experts (4 experts, top-2), Grouped-Query Attention, RoPE, SwiGLU.
| Metric | Value |
|---|---|
| AUROC | 0.773 |
| Balanced accuracy | 69.99% |
| Always-guess baseline | 50.54% |
| Full-chain accuracy | 39.97% |
150k Math-Shepherd chains, split by question (no leakage), 3 epochs, AdamW, cosine schedule, dropout 0.1. The best checkpoint was chosen by validation loss.
This is a small research baseline, not a reliable verifier. Labels come from Math-Shepherd's automatic labeling, and chains longer than 512 tokens are cut off.
Math-Shepherd: Wang et al., 2023, arXiv:2312.08935