Rune-R1 (351M) β GRPO Reasoning Model
Rune-R1 is a ~351M parameter decoder-only transformer trained from scratch and aligned for math reasoning via a three-stage pipeline:
Pretrain (FineWeb-Edu) β SFT (distilled CoT format) β GRPO (RLVR on math correctness)
This repository holds the final checkpoint: the GRPO-tuned policy, starting from Rune-R1-SFT and optimized with Group Relative Policy Optimization against a verifiable, rule-based reward for math answer correctness. See Rune-R1-Base and Rune-R1-SFT for the earlier pipeline stages.
Model Description
| Developed by | samueljayasingh |
| Model type | Causal language model (text-only) |
| Base model | Rune-R1-SFT (351M, chain-of-thought SFT on top of Rune-R1-Base) |
| Fine-tuning method | GRPO (Group Relative Policy Optimization) with PPO-style clipping and a KL penalty to a frozen reference (RLVR β reinforcement learning from verifiable rewards) |
| Dataset | data/math_train.json (math word problems with verifiable final answers), evaluated on a 50-example MATH-500 held-out subset |
| Language | English |
| Tokenizer | GPT-2 (tiktoken) |
| License | Apache 2.0 |
Architecture Details
| Parameter | Value |
|---|---|
| Layers | 22 |
| Embedding dimension | 1024 |
| Attention heads / KV groups | 16 / 4 (GQA) |
| Feed-forward hidden dim | 2816 (SwiGLU) |
| Context length | 1024 tokens |
| Position embeddings | RoPE (base 10,000) |
| Normalization | RMSNorm, with QK normalization |
Intended Uses & Limitations
Intended Use
- Research into RLVR / GRPO-style reasoning fine-tuning at small model scale.
- Reference implementation for reward-verified RL post-training pipelines (pretrain β SFT β RL).
- Studying reward hacking, KL-regularization tradeoffs, and reasoning-accuracy dynamics under a small RL step budget.
Limitations
- Small model (351M) with a limited RL budget (2,000 steps) β MATH-500 accuracy remains low (0β4% across evaluation checkpoints; see table below) and should not be compared to production-scale reasoning models.
- Reward signal is a rule-based correctness check (
\boxed{}extraction + symbolic grading), so the model may still learn to produce well-formatted but incorrect reasoning that occasionally reward-hacks the verifier. - Inherits base/SFT limitations: 1024-token context, ~5B pretraining tokens, no broad safety/RLHF alignment beyond the math-correctness reward.
- Not suitable for production or user-facing deployment β this is a research artifact demonstrating the training pipeline, not a competitive reasoning model.
How to Use
import torch
import tiktoken
from rune.model import CONFIG_350M, RuneModel
ckpt = torch.load("pytorch_model.bin", map_location="cpu")
model = RuneModel(CONFIG_350M)
model.load_state_dict(ckpt)
model.eval()
enc = tiktoken.get_encoding("gpt2")
prompt = "What is 12 * 15?"
tokens = torch.tensor([enc.encode(prompt)], dtype=torch.long)
# Model responds in "<think>...reasoning...</think>\n\n\\boxed{final_answer}" format.
# See rune/generate.py in the source repo for full sampling / KV-cache generation code.
The rune package (model definition + generation utilities) is available at the Rune-R1 GitHub repository.
Hardware
Trained end to end β pretraining, SFT, and GRPO β on a single rented GPU instance:
| Component | Spec |
|---|---|
| GPU | 1x AMD MI300X |
| VRAM | 192 GB |
| vCPU | 20 |
| RAM | 240 GB |
| Boot disk | 720 GB NVMe SSD |
| Scratch disk | 5 TB NVMe SSD |
| Rate | $1.99/hr |
Training & Evaluation
Training Procedure
| Parameter | Value |
|---|---|
| Starting checkpoint | Rune-R1-SFT |
| Reference model | Frozen copy of the SFT checkpoint (KL penalty target) |
| Training steps | 2,000 |
| Rollouts per prompt (group size) | 8 |
| Inner epochs per rollout batch | 2 |
| Max new tokens (rollout) | 512 |
| Sampling temperature / top-p | 0.8 / 0.9 |
| PPO clip epsilon | 10.0 |
| KL coefficient | 0.001 |
| Learning rate | 1e-6 |
| Reward function | Rule-based: extract \boxed{} answer, symbolically grade vs. ground truth (1.0 / 0.0) |
| Eval cadence | MATH-500 (50-example subset), every 100 steps |
Evaluation Results
| Metric | Value |
|---|---|
| Final MATH-500 accuracy (step 2000) | 0% (50 examples) |
| Peak MATH-500 accuracy | 4% (steps 1600, 1900) |
| Mean reward per step (over training) | ~0.016 |
| Max single-step average reward | 0.75 |
| Steps with nonzero reward | 152 / 2001 |
MATH-500 accuracy fluctuated in the 0β4% range throughout training rather than improving monotonically, reflecting the small model size and limited RL budget rather than a fully converged reasoning model.
Citation
@misc{RuneR12026,
author = {Samuel Jayasingh},
title = {Rune-R1: A 351M Transformer Reasoning Model Trained via Pretrain-SFT-GRPO from Scratch},
year = {2026},
publisher = {Hugging Face},
howpublished = {\url{https://huggingface.co/samueljayasingh/Rune-R1}}
}
Acknowledgements
- HuggingFaceFW/fineweb-edu β pretraining corpus.
- rasbt/math_distill β distilled chain-of-thought SFT data.
- rasbt/LLMs-from-scratch β architecture and the pretrain β SFT β GRPO reasoning-from-scratch recipe this pipeline is adapted from.
- Downloads last month
- 34
Model tree for samueljayasingh/Rune-R1
Base model
samueljayasingh/rune-0.3b-base