| --- |
| license: apache-2.0 |
| language: |
| - en |
| library_name: pytorch |
| pipeline_tag: text-generation |
| tags: |
| - mathematics |
| - recurrent-transformer |
| - custom-code |
| - sentencepiece |
| datasets: |
| - openai/gsm8k |
| - AI-MO/NuminaMath-1.5 |
| - EleutherAI/asdiv |
| --- |
| |
| # NanoMath-R |
|
|
| NanoMath-R is a math model with 103.83 million parameters, written from scratch |
| in PyTorch. It has 16 stored Transformer decoder blocks and applies the same stack |
| twice, giving 32 block applications per token without storing a second stack of |
| weights. |
|
|
| The release is intended for research on small models and direct integer arithmetic. It |
| is not a general assistant and it is not a dependable solver for word problems, |
| fractions, algebra, proofs, or calculations where an error could cause harm. |
|
|
| ## Model details |
|
|
| | Setting | Value | |
| |---|---:| |
| | Stored parameters | 103,834,368 | |
| | Physical decoder blocks | 16 | |
| | Recurrent passes | 2 | |
| | Unrolled depth | 32 | |
| | Width | 768 | |
| | Query heads | 12 | |
| | Key and value heads | 4 | |
| | Context length | 1,024 tokens | |
| | Vocabulary | 4,096 SentencePiece tokens | |
| | Position encoding | RoPE | |
| | Normalization | RMSNorm | |
| | MLP | SwiGLU | |
|
|
| The model uses grouped query attention and tied input and output embeddings. |
| Its projections have no bias, and its residual projections were initialized to |
| zero at the start of training. |
|
|
| ## Files |
|
|
| - `model_weights.pth` contains the weights, model configuration, and training |
| metadata. |
| - `token.model` is the matching SentencePiece tokenizer. |
| - `model_config.json` and `training_summary.json` expose the main checkpoint |
| metadata without loading PyTorch. |
| - `benchmark_results.json` records the release checks. |
| - `checksums.sha256` covers every published model file. |
| - `LICENSE` contains the Apache License 2.0 terms. |
|
|
| This checkpoint is incompatible with the NanoMath v1.0, v1.1, V2, and V3 |
| tokenizers. |
|
|
| ## Run it |
|
|
| The checkpoint uses the custom classes in the |
| [NanoMath-R source repository](https://github.com/agmada-asa/NanoMath-R). It |
| does not load through `AutoModelForCausalLM` or the standard Transformers text |
| generation pipeline. |
|
|
| ```bash |
| git clone https://github.com/agmada-asa/NanoMath-R.git |
| cd NanoMath-R |
| python3 -m venv .venv |
| source .venv/bin/activate |
| pip install -r requirements.txt |
| hf download agmadaasa/NanoMath-R model_weights.pth token.model --local-dir build |
| python3 scripts/chat.py "What is 84 / 7?" |
| ``` |
|
|
| The prompt format is: |
|
|
| ```text |
| <|user|> What is 84 / 7? <|end|> |
| <|assistant|> |
| ``` |
|
|
| Greedy decoding is the default. The model normally emits a reasoning span, |
| then `<|answer|>`, a numeric answer, and `<|end|>`. |
|
|
| ## Evaluation |
|
|
| Fresh release checks used greedy decoding, two recurrent passes, a KV cache, |
| FP16 on Apple MPS, and at most 384 new tokens. |
|
|
| | Evaluation | Examples | Macro exact match | Micro exact match | Malformed | |
| |---|---:|---:|---:|---:| |
| | Synthetic math reserved for evaluation | 50 | 63.27% | 66.00% | 2.00% | |
| | SVAMP | 40 | 22.50% | 22.50% | 0.00% | |
| | NanoMath v1.1 continuity prompts | 6 | 50.00% | 50.00% | 0.00% | |
|
|
| The clearest strength is direct integer arithmetic. On the synthetic sample it |
| scored 8/8 addition, 8/8 subtraction, 7/7 multiplication, and 7/7 exact |
| division. The same sample scored 2/7 linear equations, 1/7 word problems, and |
| 0/6 fraction questions. These samples are small. They describe observed |
| behavior, not a guarantee. |
|
|
| ## Training |
|
|
| The packed training stream contains 233,896,960 token positions and 1,066,609 |
| accepted documents after length filtering. Its source counts are: |
|
|
| | Source | Documents | |
| |---|---:| |
| | Verified structural arithmetic | 750,000 | |
| | NuminaMath | 190,209 | |
| | Generated language curriculum | 117,487 | |
| | GSM8K | 7,404 | |
| | ASDiv | 1,509 | |
|
|
| The run completed 6,000 optimizer steps and processed 393,216,000 token |
| positions, about 1.68 passes through the packed corpus. It used two NVIDIA T4 |
| GPUs, FP16, sequences of 1,024 tokens, attention isolated between documents, |
| loss weights based on message role, and activation checkpointing. Training |
| took 41,623.7 wall seconds, or 23.12 hours of accelerator use. The checkpoint |
| records validation loss 0.129966. |
|
|
| Prompt tokens had weight 0, reasoning tokens weight 1, and answer tokens weight |
| 3. Every packed document reset its positions and could not attend to another |
| document in the same block. |
|
|
| ## Limits and intended use |
|
|
| Good uses include studying recurrent depth, reproducing the reported |
| experiments on small models, and testing direct arithmetic prompts with independent |
| verification. |
|
|
| NanoMath-R is not reliable for financial, medical, engineering, or other |
| calculations where an error could cause harm. It can produce fluent reasoning |
| with the wrong final number. Training primarily used English, the context |
| window is 1,024 tokens, and the release represents one training seed. The model |
| has not received a broad safety or bias evaluation. It is also outside its |
| intended scope for coding, factual retrieval, unrestricted chat, |
| formal proof, and advanced mathematics. |
|
|
| The source datasets retain their own licenses and terms. Users should review |
| those terms before redistributing derived artifacts. |
|
|
| ## Integrity |
|
|
| Checkpoint SHA-256: |
|
|
| ```text |
| 6d251ab0c7595023006f7975b05436676311c18f17c8df608dc44a3e24f937b2 |
| ``` |
|
|
| Tokenizer SHA-256: |
|
|
| ```text |
| 195e6efabd0a9dc238f04c83807d4e2cb79bd0270866773ad0b71a44c3794ba1 |
| ``` |
|
|