| --- |
| license: apache-2.0 |
| language: |
| - en |
| tags: |
| - math |
| - reasoning |
| - text-generation |
| - research-preview |
| - small-language-models |
| - sft |
| - architecture-research |
| - sonamath |
| pipeline_tag: text-generation |
| --- |
| |
| # SonaMath-0.5B |
|
|
| > ## Important (please read) |
| > |
| > **The weights in this repository are from the SFT stage only.** |
| > |
| > - They are **not** trained with our experimental RL method. |
| > - They are **not** a drop-in replacement for GRPO / group-based optimizers. |
| > - **GSM8K numbers below measure this SFT checkpoint**, not an RL algorithm. |
| > |
| > We **also** research a new RL formulation (intended as an alternative to GRPO-style group optimization) and a custom reasoning architecture for small LMs. That work is **active / unfinished** and will be reported **separately** when we have matched-compute experiments and RL checkpoints. |
| > |
| > This page is a **transparent research preview** under limited compute — not a claim that “new RL already beats GRPO.” |
|
|
| --- |
|
|
| **SonaMath** is a **research preview** of a **custom architecture** for **math reasoning in the small-language-model (SLM) regime** (**approx. 0.5B** parameters), trained under a **strict compute budget** (**approx. 3B** pretrain tokens + **light SFT**). |
|
|
| > Formerly previewed as *JunMath*; the project brand is now **SonaMath**. |
|
|
| ## What this release is / is not |
|
|
| | | | |
| |---|---| |
| | **Is** | Early **SFT** weights + tokenizer + public eval protocol | |
| | **Is** | Evidence that a **small custom model** can start learning math format/reasoning under few tokens | |
| | **Is** | A call for compute to scale data, ablations, and (later) RL experiments | |
| | **Is not** | An RL-trained model | |
| | **Is not** | A finished “GRPO killer” or published RL baseline win | |
| | **Is not** | A drop-in `transformers` `AutoModel` checkpoint (custom runtime) | |
|
|
| ## How to use (inference) |
|
|
| > **Important:** SonaMath uses a **custom architecture**. |
| > It does **not** load with `AutoModelForCausalLM.from_pretrained(...)` / standard Hugging Face `pipeline` yet. |
| > Public files are **weights + tokenizer**; the research runtime loader is required to run generation. |
| |
| ### 1) Download files from this repo |
| |
| | File | Role | |
| |------|------| |
| | `sonamath.pt` | SFT weights package | |
| | `tokenizer.json` | BPE tokenizer (32k) | |
| | `tokenizer_config.json` / `special_tokens_map.json` | Tokenizer metadata | |
| | `config.json` | Public model card config | |
|
|
| ```bash |
| # Option A: huggingface-cli |
| huggingface-cli download HaniAI/SonaMath-0.5B --local-dir ./SonaMath-0.5B |
| |
| # Option B: Python |
| pip install -U "huggingface_hub>=0.23" |
| ``` |
|
|
| ```python |
| from huggingface_hub import snapshot_download |
| |
| path = snapshot_download("HaniAI/SonaMath-0.5B", local_dir="SonaMath-0.5B") |
| print("downloaded to", path) |
| # weights: SonaMath-0.5B/sonamath.pt |
| # tokenizer: SonaMath-0.5B/tokenizer.json |
| ``` |
|
|
| Downloading the repo **does count as usage** and is the supported way to fetch artifacts while the full open runtime is still under development. |
|
|
| ### 2) Prompt format (required) |
|
|
| The model is trained with explicit problem / thinking tags. Build prompts like: |
|
|
| ```text |
| <problem> |
| {your grade-school math word problem here} |
| </problem> |
| <think> |
| ``` |
|
|
| The model continues with reasoning and typically closes with: |
|
|
| ```text |
| ... reasoning ... |
| <answer> |
| {final number} |
| </answer> |
| ``` |
|
|
| ### 3) Recommended decoding (matches public GSM8K number) |
|
|
| | Setting | Value | |
| |---------|-------| |
| | Mode | **greedy** (`temperature = 0`) | |
| | Max new tokens | **1024** | |
| | Stop | `</answer>` or EOS when available in your runtime | |
|
|
| ### 4) Pseudocode (custom research runtime) |
|
|
| ```python |
| # Pseudocode — requires the SonaMath research runtime (not a stock transformers model). |
| # Full open loader will be linked here when released. |
| |
| from sonamath_runtime import load_sonamath, generate # research package (not on PyPI yet) |
| |
| model, tokenizer = load_sonamath( |
| hub_id="HaniAI/SonaMath-0.5B", # or local folder from snapshot_download |
| device="cuda", |
| ) |
| |
| problem = ( |
| "Natalia sold clips to 48 of her friends in April, and then she sold " |
| "half as many clips in May. How many clips did Natalia sell altogether " |
| "in April and May?" |
| ) |
| |
| text = generate( |
| model, |
| tokenizer, |
| problem=problem, |
| max_new_tokens=1024, |
| temperature=0.0, # greedy — same as public eval |
| ) |
| print(text) |
| ``` |
|
|
| ### 5) What works today vs later |
|
|
| | Today (this card) | Coming later | |
| |-------------------|--------------| |
| | Download **weights + tokenizer** from Hub | Public minimal inference package / Space | |
| | Reproduce **eval protocol** (greedy, max_new=1024) | Drop-in scripts without private stack | |
| | Interactive testing in the author’s research stack | Optional HF Space demo | |
| |
| If you only need artifacts for inspection, fine-tuning research, or offline study: |
| |
| ```python |
| from huggingface_hub import hf_hub_download |
|
|
| ckpt = hf_hub_download("HaniAI/SonaMath-0.5B", "sonamath.pt") |
| tok = hf_hub_download("HaniAI/SonaMath-0.5B", "tokenizer.json") |
| print(ckpt, tok) |
| ``` |
| |
| ### 6) Not supported (yet) |
| |
| ```python |
| # ❌ Will NOT work — not a standard transformers architecture export |
| from transformers import AutoModelForCausalLM, AutoTokenizer |
| model = AutoModelForCausalLM.from_pretrained("HaniAI/SonaMath-0.5B") # no |
| ``` |
| |
| Questions / runtime access requests: open a **Discussion** on this model page. |
| |
| ## Two research tracks (keep them separate) |
| |
| ### Track A — Architecture + data-efficient SFT *(this repo)* |
| |
| - Custom architecture specialized for **multi-step math reasoning** in **small LMs** |
| - Pretrain **approx. 3B** tokens (math-focused, English-centric) |
| - **Light SFT** on math Q&A / solution-style data |
| - **Public artifact:** `sonamath.pt` (SFT-stage, weights only) |
| |
| ### Track B — RL theory & methods *(research in progress; not this checkpoint)* |
| |
| - Studying a **new RL formulation** as an **alternative to group-based / GRPO-style** post-training |
| - Status: **theory + early experiments** — **no claim of superiority** on this model card |
| - **Will not** use Track A’s GSM8K score as “proof” of Track B |
| |
| > **SFT checkpoint for architecture research now; RL-vs-GRPO is future work with separate evals.** |
| |
| ## Evaluation (SFT checkpoint only) |
| |
| ### Protocol |
| |
| | Item | Setting | |
| |------|---------| |
| | Benchmark | **GSM8K test (full)** | |
| | Samples | **n = 1319** (entire official test split) | |
| | Decoding | **greedy** (`temperature = 0`) | |
| | Max new tokens | **1024** | |
| | Metric | numeric / exact match on final answer | |
| | Stage | **SFT only** | |
| |
| ### Results |
| |
| | Model | Stage | Params (approx.) | GSM8K full (greedy) | Parse rate | 95% Wilson CI | |
| |------:|-------|-----------------:|--------------------:|-----------:|--------------:| |
| | **SonaMath-0.5B (this repo)** | **SFT** | **0.5B** | **13.8% (182/1319)** | **99.3%** | **[12.0%, 15.8%]** | |
| |
| **Notes** |
| |
| - This is the **authoritative public number** for this release (full test set). |
| - Absolute accuracy is **well below** large or math-specialized models — expected under **approx. 3B** pretrain tokens. |
| - When Track B (RL) is ready, we will report **SFT vs GRPO-style vs our RL** under **matched compute**, on **separate checkpoints**. |
| |
| ## Highlights |
| |
| | | | |
| |---|---| |
| | **Brand** | **SonaMath** | |
| | **Parameters** | approx. 0.5B | |
| | **Released stage** | **SFT preview** | |
| | **Pretrain tokens** | approx. 3B | |
| | **Architecture** | Custom (details partially private while under development) | |
| | **RL (Track B)** | In progress — **not** applied to these weights | |
| | **Language** | English math focus | |
| | **GSM8K (full)** | **13.8% (182/1319)**, greedy, max_new=1024 | |
|
|
| ## Files |
|
|
| | File | Description | |
| |------|-------------| |
| | `sonamath.pt` | **SFT** weights-only package (no optimizer) | |
| | `config.json` | Public metadata | |
| | `tokenizer.json` (+ config maps) | 32k BPE, digit-aware | |
|
|
| ## Intended use |
|
|
| - Research discussion on **small reasoning LMs** and **data-efficient** math training |
| - Download weights/tokenizer for inspection and future runtime integration |
| - Planning compute for larger pretrain, cleaner SFT, and **future** RL ablations |
|
|
| **Not intended for:** production tutoring, grading, high-stakes decisions, or citing this repo as an RL result. |
|
|
| ## Limitations |
|
|
| - Low absolute GSM8K accuracy vs large / math-specialized models |
| - May hallucinate arithmetic and multi-step logic — **verify answers** |
| - Custom architecture: full inference stack not fully open as `transformers` yet |
| - **No public RL checkpoint** in this release |
|
|
| ## Compute context & GPU ask |
|
|
| | Stage | Status / budget | |
| |------|------------------| |
| | Pretrain | approx. 3B tokens (done, limited hardware) | |
| | SFT | Light schedule → **this release** | |
| | RL (vs GRPO-style) | Research track — needs more GPU for fair ablations | |
|
|
| Open a **Discussion** on this page if you can support **A100/H100 hours** for scale-up and matched-compute RL ablations. |
|
|
| ## Safety |
|
|
| Outputs can be confidently wrong. Do not use without human checking for education, finance, or safety-critical settings. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{sonamath2026sft, |
| title = {SonaMath-0.5B: SFT-Stage Research Preview of a Small Math Reasoning Model}, |
| author = {HaniAI}, |
| year = {2026}, |
| note = {SFT weights only; GSM8K full test 13.8\\% (182/1319), greedy, max\\_new=1024; RL research is separate}, |
| url = {https://huggingface.co/HaniAI/SonaMath-0.5B} |
| } |
| ``` |
|
|
| ## License |
|
|
| Apache-2.0 (weights and tokenizer files in this repository). |
|
|