| --- |
| language: |
| - en |
| license: apache-2.0 |
| tags: |
| - llm |
| - pytorch |
| - causal-lm |
| - rune-r1 |
| - reasoning |
| - grpo |
| - rlvr |
| datasets: |
| - HuggingFaceFW/fineweb-edu |
| - rasbt/math_distill |
| metrics: |
| - accuracy |
| pipeline_tag: text-generation |
| base_model: |
| - samueljayasingh/rune-0.3b-base |
| - samueljayasingh/rune-0.3b-sft |
| --- |
| |
| # Rune-R1 (351M) β GRPO Reasoning Model |
|
|
| **Rune-R1** is a ~351M parameter decoder-only transformer trained from scratch and aligned for math reasoning via a three-stage pipeline: |
|
|
| ``` |
| Pretrain (FineWeb-Edu) β SFT (distilled CoT format) β GRPO (RLVR on math correctness) |
| ``` |
|
|
| This repository holds the final checkpoint: the GRPO-tuned policy, starting from [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) and optimized with Group Relative Policy Optimization against a verifiable, rule-based reward for math answer correctness. See [Rune-R1-Base](https://huggingface.co/samueljayasingh/Rune-R1-base) and [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) for the earlier pipeline stages. |
|
|
| ## Model Description |
|
|
| | | | |
| |---|---| |
| | **Developed by** | samueljayasingh | |
| | **Model type** | Causal language model (text-only) | |
| | **Base model** | [Rune-R1-SFT](https://huggingface.co/samueljayasingh/Rune-R1-sft) (351M, chain-of-thought SFT on top of Rune-R1-Base) | |
| | **Fine-tuning method** | GRPO (Group Relative Policy Optimization) with PPO-style clipping and a KL penalty to a frozen reference (RLVR β reinforcement learning from verifiable rewards) | |
| | **Dataset** | `data/math_train.json` (math word problems with verifiable final answers), evaluated on a 50-example MATH-500 held-out subset | |
| | **Language** | English | |
| | **Tokenizer** | GPT-2 (`tiktoken`) | |
| | **License** | Apache 2.0 | |
|
|
| ### Architecture Details |
|
|
| | Parameter | Value | |
| |---|---| |
| | Layers | 22 | |
| | Embedding dimension | 1024 | |
| | Attention heads / KV groups | 16 / 4 (GQA) | |
| | Feed-forward hidden dim | 2816 (SwiGLU) | |
| | Context length | 1024 tokens | |
| | Position embeddings | RoPE (base 10,000) | |
| | Normalization | RMSNorm, with QK normalization | |
|
|
| ## Intended Uses & Limitations |
|
|
| ### Intended Use |
| - Research into RLVR / GRPO-style reasoning fine-tuning at small model scale. |
| - Reference implementation for reward-verified RL post-training pipelines (pretrain β SFT β RL). |
| - Studying reward hacking, KL-regularization tradeoffs, and reasoning-accuracy dynamics under a small RL step budget. |
|
|
| ### Limitations |
| - Small model (351M) with a limited RL budget (2,000 steps) β MATH-500 accuracy remains low (0β4% across evaluation checkpoints; see table below) and should not be compared to production-scale reasoning models. |
| - Reward signal is a rule-based correctness check (`\boxed{}` extraction + symbolic grading), so the model may still learn to produce well-formatted but incorrect reasoning that occasionally reward-hacks the verifier. |
| - Inherits base/SFT limitations: 1024-token context, ~5B pretraining tokens, no broad safety/RLHF alignment beyond the math-correctness reward. |
| - **Not suitable for production or user-facing deployment** β this is a research artifact demonstrating the training pipeline, not a competitive reasoning model. |
|
|
| ## How to Use |
|
|
| ```python |
| import torch |
| import tiktoken |
| from rune.model import CONFIG_350M, RuneModel |
| |
| ckpt = torch.load("pytorch_model.bin", map_location="cpu") |
| model = RuneModel(CONFIG_350M) |
| model.load_state_dict(ckpt) |
| model.eval() |
| |
| enc = tiktoken.get_encoding("gpt2") |
| prompt = "What is 12 * 15?" |
| tokens = torch.tensor([enc.encode(prompt)], dtype=torch.long) |
| |
| # Model responds in "<think>...reasoning...</think>\n\n\\boxed{final_answer}" format. |
| # See rune/generate.py in the source repo for full sampling / KV-cache generation code. |
| ``` |
|
|
| The `rune` package (model definition + generation utilities) is available at the Rune-R1 GitHub repository. |
|
|
| ## Hardware |
|
|
| Trained end to end β pretraining, SFT, and GRPO β on a single rented GPU instance: |
|
|
| | Component | Spec | |
| |---|---| |
| | GPU | 1x AMD MI300X | |
| | VRAM | 192 GB | |
| | vCPU | 20 | |
| | RAM | 240 GB | |
| | Boot disk | 720 GB NVMe SSD | |
| | Scratch disk | 5 TB NVMe SSD | |
| | Rate | $1.99/hr | |
|
|
| ## Training & Evaluation |
|
|
| ### Training Procedure |
|
|
| | Parameter | Value | |
| |---|---| |
| | Starting checkpoint | Rune-R1-SFT | |
| | Reference model | Frozen copy of the SFT checkpoint (KL penalty target) | |
| | Training steps | 2,000 | |
| | Rollouts per prompt (group size) | 8 | |
| | Inner epochs per rollout batch | 2 | |
| | Max new tokens (rollout) | 512 | |
| | Sampling temperature / top-p | 0.8 / 0.9 | |
| | PPO clip epsilon | 10.0 | |
| | KL coefficient | 0.001 | |
| | Learning rate | 1e-6 | |
| | Reward function | Rule-based: extract `\boxed{}` answer, symbolically grade vs. ground truth (1.0 / 0.0) | |
| | Eval cadence | MATH-500 (50-example subset), every 100 steps | |
|
|
| ### Evaluation Results |
|
|
| | Metric | Value | |
| |---|---| |
| | Final MATH-500 accuracy (step 2000) | 0% (50 examples) | |
| | Peak MATH-500 accuracy | 4% (steps 1600, 1900) | |
| | Mean reward per step (over training) | ~0.016 | |
| | Max single-step average reward | 0.75 | |
| | Steps with nonzero reward | 152 / 2001 | |
|
|
| MATH-500 accuracy fluctuated in the 0β4% range throughout training rather than improving monotonically, reflecting the small model size and limited RL budget rather than a fully converged reasoning model. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{RuneR12026, |
| author = {Samuel Jayasingh}, |
| title = {Rune-R1: A 351M Transformer Reasoning Model Trained via Pretrain-SFT-GRPO from Scratch}, |
| year = {2026}, |
| publisher = {Hugging Face}, |
| howpublished = {\url{https://huggingface.co/samueljayasingh/Rune-R1}} |
| } |
| ``` |
|
|
| ## Acknowledgements |
|
|
| - [HuggingFaceFW/fineweb-edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu) β pretraining corpus. |
| - [rasbt/math_distill](https://huggingface.co/datasets/rasbt/math_distill) β distilled chain-of-thought SFT data. |
| - [rasbt/LLMs-from-scratch](https://github.com/rasbt/LLMs-from-scratch) β architecture and the pretrain β SFT β GRPO reasoning-from-scratch recipe this pipeline is adapted from. |