Rune-R1 (351M) β€” GRPO Reasoning Model

Rune-R1 is a ~351M parameter decoder-only transformer trained from scratch and aligned for math reasoning via a three-stage pipeline:

Pretrain (FineWeb-Edu)  β†’  SFT (distilled CoT format)  β†’  GRPO (RLVR on math correctness)

This repository holds the final checkpoint: the GRPO-tuned policy, starting from Rune-R1-SFT and optimized with Group Relative Policy Optimization against a verifiable, rule-based reward for math answer correctness. See Rune-R1-Base and Rune-R1-SFT for the earlier pipeline stages.

Model Description

Developed by samueljayasingh
Model type Causal language model (text-only)
Base model Rune-R1-SFT (351M, chain-of-thought SFT on top of Rune-R1-Base)
Fine-tuning method GRPO (Group Relative Policy Optimization) with PPO-style clipping and a KL penalty to a frozen reference (RLVR β€” reinforcement learning from verifiable rewards)
Dataset data/math_train.json (math word problems with verifiable final answers), evaluated on a 50-example MATH-500 held-out subset
Language English
Tokenizer GPT-2 (tiktoken)
License Apache 2.0

Architecture Details

Parameter Value
Layers 22
Embedding dimension 1024
Attention heads / KV groups 16 / 4 (GQA)
Feed-forward hidden dim 2816 (SwiGLU)
Context length 1024 tokens
Position embeddings RoPE (base 10,000)
Normalization RMSNorm, with QK normalization

Intended Uses & Limitations

Intended Use

  • Research into RLVR / GRPO-style reasoning fine-tuning at small model scale.
  • Reference implementation for reward-verified RL post-training pipelines (pretrain β†’ SFT β†’ RL).
  • Studying reward hacking, KL-regularization tradeoffs, and reasoning-accuracy dynamics under a small RL step budget.

Limitations

  • Small model (351M) with a limited RL budget (2,000 steps) β€” MATH-500 accuracy remains low (0–4% across evaluation checkpoints; see table below) and should not be compared to production-scale reasoning models.
  • Reward signal is a rule-based correctness check (\boxed{} extraction + symbolic grading), so the model may still learn to produce well-formatted but incorrect reasoning that occasionally reward-hacks the verifier.
  • Inherits base/SFT limitations: 1024-token context, ~5B pretraining tokens, no broad safety/RLHF alignment beyond the math-correctness reward.
  • Not suitable for production or user-facing deployment β€” this is a research artifact demonstrating the training pipeline, not a competitive reasoning model.

How to Use

import torch
import tiktoken
from rune.model import CONFIG_350M, RuneModel

ckpt = torch.load("pytorch_model.bin", map_location="cpu")
model = RuneModel(CONFIG_350M)
model.load_state_dict(ckpt)
model.eval()

enc = tiktoken.get_encoding("gpt2")
prompt = "What is 12 * 15?"
tokens = torch.tensor([enc.encode(prompt)], dtype=torch.long)

# Model responds in "<think>...reasoning...</think>\n\n\\boxed{final_answer}" format.
# See rune/generate.py in the source repo for full sampling / KV-cache generation code.

The rune package (model definition + generation utilities) is available at the Rune-R1 GitHub repository.

Hardware

Trained end to end β€” pretraining, SFT, and GRPO β€” on a single rented GPU instance:

Component Spec
GPU 1x AMD MI300X
VRAM 192 GB
vCPU 20
RAM 240 GB
Boot disk 720 GB NVMe SSD
Scratch disk 5 TB NVMe SSD
Rate $1.99/hr

Training & Evaluation

Training Procedure

Parameter Value
Starting checkpoint Rune-R1-SFT
Reference model Frozen copy of the SFT checkpoint (KL penalty target)
Training steps 2,000
Rollouts per prompt (group size) 8
Inner epochs per rollout batch 2
Max new tokens (rollout) 512
Sampling temperature / top-p 0.8 / 0.9
PPO clip epsilon 10.0
KL coefficient 0.001
Learning rate 1e-6
Reward function Rule-based: extract \boxed{} answer, symbolically grade vs. ground truth (1.0 / 0.0)
Eval cadence MATH-500 (50-example subset), every 100 steps

Evaluation Results

Metric Value
Final MATH-500 accuracy (step 2000) 0% (50 examples)
Peak MATH-500 accuracy 4% (steps 1600, 1900)
Mean reward per step (over training) ~0.016
Max single-step average reward 0.75
Steps with nonzero reward 152 / 2001

MATH-500 accuracy fluctuated in the 0–4% range throughout training rather than improving monotonically, reflecting the small model size and limited RL budget rather than a fully converged reasoning model.

Citation

@misc{RuneR12026,
  author = {Samuel Jayasingh},
  title = {Rune-R1: A 351M Transformer Reasoning Model Trained via Pretrain-SFT-GRPO from Scratch},
  year = {2026},
  publisher = {Hugging Face},
  howpublished = {\url{https://huggingface.co/samueljayasingh/Rune-R1}}
}

Acknowledgements

Downloads last month
34
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for samueljayasingh/Rune-R1

Finetuned
(2)
this model

Datasets used to train samueljayasingh/Rune-R1