license: apache-2.0
language:
- en
library_name: pytorch
pipeline_tag: text-generation
tags:
- mathematics
- recurrent-transformer
- custom-code
- sentencepiece
datasets:
- openai/gsm8k
- AI-MO/NuminaMath-1.5
- EleutherAI/asdiv
NanoMath-R
NanoMath-R is a math model with 103.83 million parameters, written from scratch in PyTorch. It has 16 stored Transformer decoder blocks and applies the same stack twice, giving 32 block applications per token without storing a second stack of weights.
The release is intended for research on small models and direct integer arithmetic. It is not a general assistant and it is not a dependable solver for word problems, fractions, algebra, proofs, or calculations where an error could cause harm.
Model details
| Setting | Value |
|---|---|
| Stored parameters | 103,834,368 |
| Physical decoder blocks | 16 |
| Recurrent passes | 2 |
| Unrolled depth | 32 |
| Width | 768 |
| Query heads | 12 |
| Key and value heads | 4 |
| Context length | 1,024 tokens |
| Vocabulary | 4,096 SentencePiece tokens |
| Position encoding | RoPE |
| Normalization | RMSNorm |
| MLP | SwiGLU |
The model uses grouped query attention and tied input and output embeddings. Its projections have no bias, and its residual projections were initialized to zero at the start of training.
Files
model_weights.pthcontains the weights, model configuration, and training metadata.token.modelis the matching SentencePiece tokenizer.model_config.jsonandtraining_summary.jsonexpose the main checkpoint metadata without loading PyTorch.benchmark_results.jsonrecords the release checks.checksums.sha256covers every published model file.LICENSEcontains the Apache License 2.0 terms.
This checkpoint is incompatible with the NanoMath v1.0, v1.1, V2, and V3 tokenizers.
Run it
The checkpoint uses the custom classes in the
NanoMath-R source repository. It
does not load through AutoModelForCausalLM or the standard Transformers text
generation pipeline.
git clone https://github.com/agmada-asa/NanoMath-R.git
cd NanoMath-R
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements.txt
hf download agmadaasa/NanoMath-R model_weights.pth token.model --local-dir build
python3 scripts/chat.py "What is 84 / 7?"
The prompt format is:
<|user|> What is 84 / 7? <|end|>
<|assistant|>
Greedy decoding is the default. The model normally emits a reasoning span,
then <|answer|>, a numeric answer, and <|end|>.
Evaluation
Fresh release checks used greedy decoding, two recurrent passes, a KV cache, FP16 on Apple MPS, and at most 384 new tokens.
| Evaluation | Examples | Macro exact match | Micro exact match | Malformed |
|---|---|---|---|---|
| Synthetic math reserved for evaluation | 50 | 63.27% | 66.00% | 2.00% |
| SVAMP | 40 | 22.50% | 22.50% | 0.00% |
| NanoMath v1.1 continuity prompts | 6 | 50.00% | 50.00% | 0.00% |
The clearest strength is direct integer arithmetic. On the synthetic sample it scored 8/8 addition, 8/8 subtraction, 7/7 multiplication, and 7/7 exact division. The same sample scored 2/7 linear equations, 1/7 word problems, and 0/6 fraction questions. These samples are small. They describe observed behavior, not a guarantee.
Training
The packed training stream contains 233,896,960 token positions and 1,066,609 accepted documents after length filtering. Its source counts are:
| Source | Documents |
|---|---|
| Verified structural arithmetic | 750,000 |
| NuminaMath | 190,209 |
| Generated language curriculum | 117,487 |
| GSM8K | 7,404 |
| ASDiv | 1,509 |
The run completed 6,000 optimizer steps and processed 393,216,000 token positions, about 1.68 passes through the packed corpus. It used two NVIDIA T4 GPUs, FP16, sequences of 1,024 tokens, attention isolated between documents, loss weights based on message role, and activation checkpointing. Training took 41,623.7 wall seconds, or 23.12 hours of accelerator use. The checkpoint records validation loss 0.129966.
Prompt tokens had weight 0, reasoning tokens weight 1, and answer tokens weight 3. Every packed document reset its positions and could not attend to another document in the same block.
Limits and intended use
Good uses include studying recurrent depth, reproducing the reported experiments on small models, and testing direct arithmetic prompts with independent verification.
NanoMath-R is not reliable for financial, medical, engineering, or other calculations where an error could cause harm. It can produce fluent reasoning with the wrong final number. Training primarily used English, the context window is 1,024 tokens, and the release represents one training seed. The model has not received a broad safety or bias evaluation. It is also outside its intended scope for coding, factual retrieval, unrestricted chat, formal proof, and advanced mathematics.
The source datasets retain their own licenses and terms. Users should review those terms before redistributing derived artifacts.
Integrity
Checkpoint SHA-256:
6d251ab0c7595023006f7975b05436676311c18f17c8df608dc44a3e24f937b2
Tokenizer SHA-256:
195e6efabd0a9dc238f04c83807d4e2cb79bd0270866773ad0b71a44c3794ba1