--- license: apache-2.0 language: - en tags: - math - reasoning - text-generation - research-preview - small-language-models - sft - architecture-research - sonamath pipeline_tag: text-generation --- # SonaMath-0.5B > ## Important (please read) > > **The weights in this repository are from the SFT stage only.** > > - They are **not** trained with our experimental RL method. > - They are **not** a drop-in replacement for GRPO / group-based optimizers. > - **GSM8K numbers below measure this SFT checkpoint**, not an RL algorithm. > > We **also** research a new RL formulation (intended as an alternative to GRPO-style group optimization) and a custom reasoning architecture for small LMs. That work is **active / unfinished** and will be reported **separately** when we have matched-compute experiments and RL checkpoints. > > This page is a **transparent research preview** under limited compute — not a claim that “new RL already beats GRPO.” --- **SonaMath** is a **research preview** of a **custom architecture** for **math reasoning in the small-language-model (SLM) regime** (**approx. 0.5B** parameters), trained under a **strict compute budget** (**approx. 3B** pretrain tokens + **light SFT**). > Formerly previewed as *JunMath*; the project brand is now **SonaMath**. ## What this release is / is not | | | |---|---| | **Is** | Early **SFT** weights + tokenizer + public eval protocol | | **Is** | Evidence that a **small custom model** can start learning math format/reasoning under few tokens | | **Is** | A call for compute to scale data, ablations, and (later) RL experiments | | **Is not** | An RL-trained model | | **Is not** | A finished “GRPO killer” or published RL baseline win | | **Is not** | A drop-in `transformers` `AutoModel` checkpoint (custom runtime) | ## How to use (inference) > **Important:** SonaMath uses a **custom architecture**. > It does **not** load with `AutoModelForCausalLM.from_pretrained(...)` / standard Hugging Face `pipeline` yet. > Public files are **weights + tokenizer**; the research runtime loader is required to run generation. ### 1) Download files from this repo | File | Role | |------|------| | `sonamath.pt` | SFT weights package | | `tokenizer.json` | BPE tokenizer (32k) | | `tokenizer_config.json` / `special_tokens_map.json` | Tokenizer metadata | | `config.json` | Public model card config | ```bash # Option A: huggingface-cli huggingface-cli download HaniAI/SonaMath-0.5B --local-dir ./SonaMath-0.5B # Option B: Python pip install -U "huggingface_hub>=0.23" ``` ```python from huggingface_hub import snapshot_download path = snapshot_download("HaniAI/SonaMath-0.5B", local_dir="SonaMath-0.5B") print("downloaded to", path) # weights: SonaMath-0.5B/sonamath.pt # tokenizer: SonaMath-0.5B/tokenizer.json ``` Downloading the repo **does count as usage** and is the supported way to fetch artifacts while the full open runtime is still under development. ### 2) Prompt format (required) The model is trained with explicit problem / thinking tags. Build prompts like: ```text {your grade-school math word problem here} ``` The model continues with reasoning and typically closes with: ```text ... reasoning ... {final number} ``` ### 3) Recommended decoding (matches public GSM8K number) | Setting | Value | |---------|-------| | Mode | **greedy** (`temperature = 0`) | | Max new tokens | **1024** | | Stop | `` or EOS when available in your runtime | ### 4) Pseudocode (custom research runtime) ```python # Pseudocode — requires the SonaMath research runtime (not a stock transformers model). # Full open loader will be linked here when released. from sonamath_runtime import load_sonamath, generate # research package (not on PyPI yet) model, tokenizer = load_sonamath( hub_id="HaniAI/SonaMath-0.5B", # or local folder from snapshot_download device="cuda", ) problem = ( "Natalia sold clips to 48 of her friends in April, and then she sold " "half as many clips in May. How many clips did Natalia sell altogether " "in April and May?" ) text = generate( model, tokenizer, problem=problem, max_new_tokens=1024, temperature=0.0, # greedy — same as public eval ) print(text) ``` ### 5) What works today vs later | Today (this card) | Coming later | |-------------------|--------------| | Download **weights + tokenizer** from Hub | Public minimal inference package / Space | | Reproduce **eval protocol** (greedy, max_new=1024) | Drop-in scripts without private stack | | Interactive testing in the author’s research stack | Optional HF Space demo | If you only need artifacts for inspection, fine-tuning research, or offline study: ```python from huggingface_hub import hf_hub_download ckpt = hf_hub_download("HaniAI/SonaMath-0.5B", "sonamath.pt") tok = hf_hub_download("HaniAI/SonaMath-0.5B", "tokenizer.json") print(ckpt, tok) ``` ### 6) Not supported (yet) ```python # ❌ Will NOT work — not a standard transformers architecture export from transformers import AutoModelForCausalLM, AutoTokenizer model = AutoModelForCausalLM.from_pretrained("HaniAI/SonaMath-0.5B") # no ``` Questions / runtime access requests: open a **Discussion** on this model page. ## Two research tracks (keep them separate) ### Track A — Architecture + data-efficient SFT *(this repo)* - Custom architecture specialized for **multi-step math reasoning** in **small LMs** - Pretrain **approx. 3B** tokens (math-focused, English-centric) - **Light SFT** on math Q&A / solution-style data - **Public artifact:** `sonamath.pt` (SFT-stage, weights only) ### Track B — RL theory & methods *(research in progress; not this checkpoint)* - Studying a **new RL formulation** as an **alternative to group-based / GRPO-style** post-training - Status: **theory + early experiments** — **no claim of superiority** on this model card - **Will not** use Track A’s GSM8K score as “proof” of Track B > **SFT checkpoint for architecture research now; RL-vs-GRPO is future work with separate evals.** ## Evaluation (SFT checkpoint only) ### Protocol | Item | Setting | |------|---------| | Benchmark | **GSM8K test (full)** | | Samples | **n = 1319** (entire official test split) | | Decoding | **greedy** (`temperature = 0`) | | Max new tokens | **1024** | | Metric | numeric / exact match on final answer | | Stage | **SFT only** | ### Results | Model | Stage | Params (approx.) | GSM8K full (greedy) | Parse rate | 95% Wilson CI | |------:|-------|-----------------:|--------------------:|-----------:|--------------:| | **SonaMath-0.5B (this repo)** | **SFT** | **0.5B** | **13.8% (182/1319)** | **99.3%** | **[12.0%, 15.8%]** | **Notes** - This is the **authoritative public number** for this release (full test set). - Absolute accuracy is **well below** large or math-specialized models — expected under **approx. 3B** pretrain tokens. - When Track B (RL) is ready, we will report **SFT vs GRPO-style vs our RL** under **matched compute**, on **separate checkpoints**. ## Highlights | | | |---|---| | **Brand** | **SonaMath** | | **Parameters** | approx. 0.5B | | **Released stage** | **SFT preview** | | **Pretrain tokens** | approx. 3B | | **Architecture** | Custom (details partially private while under development) | | **RL (Track B)** | In progress — **not** applied to these weights | | **Language** | English math focus | | **GSM8K (full)** | **13.8% (182/1319)**, greedy, max_new=1024 | ## Files | File | Description | |------|-------------| | `sonamath.pt` | **SFT** weights-only package (no optimizer) | | `config.json` | Public metadata | | `tokenizer.json` (+ config maps) | 32k BPE, digit-aware | ## Intended use - Research discussion on **small reasoning LMs** and **data-efficient** math training - Download weights/tokenizer for inspection and future runtime integration - Planning compute for larger pretrain, cleaner SFT, and **future** RL ablations **Not intended for:** production tutoring, grading, high-stakes decisions, or citing this repo as an RL result. ## Limitations - Low absolute GSM8K accuracy vs large / math-specialized models - May hallucinate arithmetic and multi-step logic — **verify answers** - Custom architecture: full inference stack not fully open as `transformers` yet - **No public RL checkpoint** in this release ## Compute context & GPU ask | Stage | Status / budget | |------|------------------| | Pretrain | approx. 3B tokens (done, limited hardware) | | SFT | Light schedule → **this release** | | RL (vs GRPO-style) | Research track — needs more GPU for fair ablations | Open a **Discussion** on this page if you can support **A100/H100 hours** for scale-up and matched-compute RL ablations. ## Safety Outputs can be confidently wrong. Do not use without human checking for education, finance, or safety-critical settings. ## Citation ```bibtex @misc{sonamath2026sft, title = {SonaMath-0.5B: SFT-Stage Research Preview of a Small Math Reasoning Model}, author = {HaniAI}, year = {2026}, note = {SFT weights only; GSM8K full test 13.8\\% (182/1319), greedy, max\\_new=1024; RL research is separate}, url = {https://huggingface.co/HaniAI/SonaMath-0.5B} } ``` ## License Apache-2.0 (weights and tokenizer files in this repository).