Argonne-3.0-think / README.md
PursuitOfDataScience's picture
Add companion blog post link to model card
435eb17 verified
|
Raw
History Blame Contribute Delete
37.5 kB
---
license: apache-2.0
language:
- en
library_name: transformers
base_model:
- PursuitOfDataScience/argonne-3.0-base
tags:
- text-generation
- causal-lm
- transformer
- argonne
- reasoning
- chain-of-thought
- think
- sft
- dpo
- model-soup
- tool-use
- distillation
pipeline_tag: text-generation
---
# Argonne 3.0-think
**Argonne 3.0-think** is a 2.88B-parameter **reasoning** model from the Argonne 3.x family. It supports an explicit chain-of-thought "think" mode (`<think> … </think>`) and answers directly in no-think mode. It is built entirely from [Argonne 3.0-base](https://huggingface.co/PursuitOfDataScience/argonne-3.0-base) through a multi-stage pipeline that ends in **training-free weight-soups** β€” one to repair the pretraining base, one to reconcile reasoning with general chat, and further recovery soups after each round of self-improvement.
On a 40-item internal 4-quadrant probe ({math, general} Γ— {no-think, with-CoT}) it scores **33/40**, with strong arithmetic in *both* modes while remaining a loop-free general chat model. (This is a small hand-checked probe β€” for held-out benchmark numbers and how to get the most out of the model, see *Getting more accuracy: test-time compute* below.)
πŸ“ **Companion blog post:** [*How I Taught a Small Language Model to Reason*](https://youzhi.netlify.app/post/2026-07-21-reasoning-model/reasoning-model/) β€” a plain-language walkthrough of how this model was built (calibrated chain-of-thought data + training-free weight-soups), for readers who want the story without the code.
> **Update (2026-07-12) β€” v4.** A modest, **external-teacher-distillation** reasoning update. A stronger open model (**Qwen3-4B**) solved the GSM8K **train** split (contamination-safe, 87.8% correct β†’ 3,714 short worked solutions); those distilled traces β€” plus new **tool-calling** and **coding** data tiers β€” were CoT-SFT'd from `dpo_soup` (β†’ `think_v7`) and cross-souped with v3 (**v4 = `0.3Β·think_v7 + 0.7Β·v3`**). On clean held-out math the teacher distillation raised **greedy ASDiv 25.8β†’30.5** and strengthened native `</think>` termination; SVAMP greedy is ~flat and **self-consistency regressed ~6 pts** (a trade β€” see *Getting more accuracy* and *Limitations*). Honestly, v4 is a **mixed, reasoning-focused update**, not a clean dominance over v3. The new **tool-calling** capability was learned cleanly (the training checkpoint emits 100% valid `<tool_call>`s) but is **not baked into this shipped card** (it dilutes out of the conservative soup and, without a real executor, the model hallucinates tool outputs); it is exposed instead as a **serving-side tool-execution loop** β€” see *Tool use & serving-system capabilities*. Coding did not measurably transfer at this scale (HumanEval β‰ˆ 0, base-capacity-limited). See step 10 in *How it was trained*. v3 (`x_v6v2_040`) is retained upstream for rollback.
> **Update (2026-07-11) β€” v3.** This release targets the model's #1 *deployable* failure: **non-termination** β€” on clean held-out problems (**SVAMP/ASDiv**) ~50–60% of sampled `<think>` traces never close `</think>`, even on one-step problems (the model over-thinks). v3 distills the fix into the **weights**: a **SHORT-trace** CoT-SFT (train *only* on ≀768-token closed-correct traces) so greedy decoding **natively terminates** with the answer β€” turning the decode-time budget-forcing win into a weights-level one β€” plus a contamination-safe grade-school **procedure** tier; the result is then **weight-soup–fused with v2** to keep v2's self-consistency diversity. On clean held-out math (SVAMP / ASDiv, greedy): **18.0β†’22.7 / 22.7β†’27.3**; self-consistency held (**36.3β†’40.3 / 51.0β†’48.0**); pass@32 up (**73.3β†’74.7 / 74.3β†’77.0**). See step 9 in *How it was trained*. *(GSM8K remains contaminated for this model β€” see the note under *Benchmarks* β€” so held-out SVAMP/ASDiv are the honest judge.)*
> **Update (2026-07-08) β€” v2.** The model was improved by a round of **STaR self-improvement** (rejection-sampling fine-tuning on the model's *own* verified-correct GSM8K `<think>` traces) followed by a **third training-free soup** that recovers the no-think arithmetic that step regressed β€” improving held-out math (on the clean **SVAMP/ASDiv** benchmarks, greedy ~9%β†’~18% / ~18%β†’~23%; see *Benchmarks*) while keeping general chat and no-think math intact. *(Earlier versions of this card reported a GSM8K 2%β†’~7.5% gain; GSM8K is now known to be **contaminated** for this model β€” see the note under *Benchmarks* β€” so held-out SVAMP/ASDiv are used instead.)* See step 8 in *How it was trained*. Also: `config.json` now ships an `auto_map`, so `AutoModelForCausalLM.from_pretrained(..., trust_remote_code=True)` loads it directly (no manual `import model` needed).
> **Decoding note (important):** for math/arithmetic, use **greedy decoding** and **do not** apply an aggressive `repetition_penalty` / `no_repeat_ngram_size`. Anti-repeat decoding corrupts multi-digit arithmetic (it blocks the model from re-emitting digits it just used, e.g. turning `80 / 2` into `8 / 2`) and collapses math accuracy. See *Recommended inference settings* below.
## How it was trained
Argonne 3.0-think is the endpoint of the following pipeline. Every stage runs at 13,568-token context with RoPE ΞΈ = 1,000,000 (the base is RoPE-extrapolated from a 1,024-ctx pretraining run).
| # | Stage | Init from | Data | What it does |
|---|-------|-----------|------|--------------|
| 1 | **Pretraining** (from scratch) | β€” | [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb) (~76B tokens, 1,024 ctx) | The 2.88B base, [Argonne 3.0-base](https://huggingface.co/PursuitOfDataScience/argonne-3.0-base). |
| 2 | **Intermix midtraining** | Argonne 3.0-base | FineWeb (crawl `CC-MAIN-2025-21`) + [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) subset `finemath-4plus`, mixed 50:50 β†’ 60:40 by document | Injects grade-school + competition numeracy the pure-FineWeb base lacked. |
| 3 | **Soup base** (training-free) | β€” | *(weight interpolation, no data)* | `0.35 Β· base + 0.65 Β· intermix`. A linear weight-soup of the two same-lineage checkpoints β€” recovers general knowledge the midtraining LR over-wrote while keeping the math. This is the reasoning-ready base. |
| 4 | **SFT** | Soup base | [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) (`train_sft`) | Instruction / multi-turn chat. |
| 5 | **DPO** | SFT model | [KatoHF/chatbot_arena_binarized](https://huggingface.co/datasets/KatoHF/chatbot_arena_binarized) | Preference alignment on real user comparisons. Kept as `dpo_soup` (the general-healthy pre-CoT checkpoint). |
| 6 | **CoT-SFT** | DPO model | `cot_sft_mix_v3` (see table below) | Teaches explicit `<think>` chain-of-thought; brings math to 10/10 but regresses general chat (over-reasons / loops). |
| 7 | **Final weight-soup** (training-free) | β€” | *(weight interpolation, no data)* | `0.15 Β· (DPO model) + 0.85 Β· (CoT-SFT model)`. Fractionally "un-applies" the CoT stage's general-chat regression while keeping its math. The `soup_blend_a085` reasoning model (v1). |
| 8 | **STaR self-improvement + soup-recovery** (v2, 2026-07-08) | step 7 model | Rejection-sampled, **verified-correct** GSM8K `<think>` traces (the model's own, filtered by a `\boxed{}` checker) + a general/no-think anchor; then a training-free soup `0.4 Β· (step 7) + 0.6 Β· (STaR model)` | STaR: fine-tunes on the model's own correct solutions. The third soup recovers the no-think multi-step arithmetic the STaR step regressed (e.g. divisor-counting), keeping general chat intact. |
| 9 | **SHORT-trace termination + procedure distillation β†’ cross-soup** (v3, 2026-07-11) | `dpo_soup` (step-5 pre-CoT) β†’ CoT-SFT β†’ `think_v6`; then cross-soup with v2 | `cot_sft_mix_v6`: **short-only** (≀768-token) closed-correct traces β€” drops the contaminated `easy_gsm8k` and adds a contamination-safe **gsm8k-train-only** short procedure tier (canonicalized `\boxed{}` close). Ship = training-free cross-soup **`0.4 Β· think_v6 + 0.6 Β· v2`** | Short-only SFT makes greedy **natively** close `</think>` with the answer (turns the decode-time budget-force win into a weights-level win) and installs grade-school procedure; the cross-soup with v2 keeps v2's self-consistency diversity. Clean SVAMP/ASDiv greedy +4.7/+4.6, self-cons held. |
| 10 | **External-teacher distillation + tool/coding β†’ cross-soup** (v4, 2026-07-12) | `dpo_soup` β†’ CoT-SFT (`cot_sft_mix_v7`) β†’ `think_v7`; then cross-soup with v3 | `cot_sft_mix_v7` (~35.6k): the v6 short-only backbone + **`teacher_gsm`** (Qwen3-4B's correct gsm8k-**train** solutions, 3,714) + **`code_magicoder`** (decontaminated code) + synthesized **`tool_calc`** (calculator/python tool-use). Ship = training-free cross-soup **`0.3 Β· think_v7 + 0.6·… (0.7 Β· v3)`** | Distills a *stronger* teacher's worked solutions into the model (self-STaR had saturated). Lifts greedy ASDiv (25.8β†’30.5) + native termination; self-consistency regressed ~6 pts (a trade). Tool-calling learned in `think_v7` but not shipped in the soup (see *Tool use*); coding did not transfer at 2.88B. |
**The soups are the key idea.** The CoT-SFT model (step 6) and the DPO model (step 5) live in the same optimization basin (step 6 fine-tunes step 5), so `CoT-SFT = DPO + Ξ”` for a chain-of-thought weight-delta Ξ”. Blending back 15% of the DPO weights (`0.85 Β· CoT-SFT + 0.15 Β· DPO`) scales that delta down just enough to remove the loop/forgetting pathology the CoT stage introduced β€” the grammar loop disappears, lost facts (e.g. "the Red Planet is Mars") return β€” **without hurting the 10/10 math**. Ξ± = 0.85 is a knee: lower values recover even more general chat but start breaking `<think>` trace-closure (the CoT format itself lives in Ξ”). The same recovery trick is reused after each self-improvement round (steps 8 and 9).
### Chain-of-thought SFT data
**v3 (`cot_sft_mix_v6`, ~26.4k examples, all ≀768 tokens).** Deliberately **short-only** β€” training on short, terminating traces is what teaches greedy to natively close `</think>` (the fix for non-termination). The contaminated `easy_gsm8k` tier is dropped and replaced with a contamination-safe **gsm8k-train-only** procedure tier; every long tier from v3-mix is removed.
| Tier | Share | Source |
|------|-----:|--------|
| `direct_tulu` (no-think chat) | 30.3% | [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) |
| `gsm8k_train_short` (NEW, contamination-safe) | 16.4% | [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`) β€” **train split only**, ≀512 tok, verified closed+boxed, canonicalized to the deployed `\boxed{}` close |
| `synth_arith` | 9.5% | Synthetic, correct-by-construction |
| `ms_algebra` / `ms_series` / `ms_geometry` / `ms_divisors` | ~18% | Synthetic multi-step (Python-verified) |
| `med_math` | 7.6% | [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) (levels 1–3) |
| `gen_ultrachat` / `hq_opus` / `med_openmath` / `hard_strict` | small | ultrachat-derived / Opus / OpenMathReasoning / MiniMax-M2.1, all ≀768 tok |
**v1/v2 (`cot_sft_mix_v3`, ~113k examples).** The mix used for pipeline step 6 (still the basis of v1/v2):
| Tier | Rows | Source |
|------|-----:|--------|
| `direct_tulu` (no-think chat) | 34,000 | [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) |
| `synth_arith` | 15,000 | Synthetic, correct-by-construction |
| `gen_ultrachat` (CoT-augmented) | 15,000 | Derived from [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) |
| `hard_strict` | 12,000 | [PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts](https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts) (open-ended, strict-filtered) |
| `easy_gsm8k` *(dropped in v3 β€” contaminated)* | 8,402 | [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`), curated with `<think>` / `\boxed{}` |
| `med_math` | 5,729 | [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) (levels 1–3) |
| `ms_algebra` | 5,000 | Synthetic multi-step (Python-verified) |
| `ms_series` | 5,000 | Synthetic multi-step (Python-verified) |
| `ms_geometry` | 5,000 | Synthetic multi-step (Python-verified) |
| `med_openmath` | 4,620 | Problems from [nvidia/OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) (CoT subset); solutions regenerated |
| `hq_opus` | 2,300 | [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) |
| `ms_divisors` | 1,290 | Synthetic multi-step (Python-verified) |
Every example is capped at 4,000 tokens (768 for v3); all synthetic and `ms_*` traces are re-verified with a `\boxed{}` answer extractor.
### Key hyperparameters
| Stage | LR | Epochs | Effective batch | Ctx | ΞΈ |
|-------|----|-------:|-----------------|----:|---|
| SFT | 2e-5 | 1 | 18 | 13,568 | 1e6 |
| DPO | 1e-6 (Ξ² = 0.03) | 1 | 8 | 13,568 | 1e6 |
| CoT-SFT (v1/v2) | 1e-5 | 1 | 12 (3Γ— H200 DDP) | 13,568 | 1e6 |
| CoT-SFT (v3, short-only) | 1e-5 | 1 | 12 (1Γ— H100) | 13,568 | 1e6 |
## Datasets used (all stages)
1. [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb) β€” pretraining + intermix general half (crawl `CC-MAIN-2025-21`).
2. [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) (`finemath-4plus`) β€” intermix math half.
3. [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) β€” SFT, and the `gen_ultrachat` CoT tier.
4. [KatoHF/chatbot_arena_binarized](https://huggingface.co/datasets/KatoHF/chatbot_arena_binarized) β€” DPO preferences.
5. [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) β€” CoT `direct_tulu` no-think tier.
6. [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`) β€” CoT `easy_gsm8k` tier (v1/v2, contaminated) and the v3 `gsm8k_train_short` tier (train split only).
7. [nvidia/OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) β€” CoT `med_openmath` tier.
8. [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) β€” CoT `med_math` tier.
9. [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) β€” CoT `hq_opus` tier.
10. [PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts](https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts) β€” CoT `hard_strict` tier.
Fully synthetic (no external dataset): `synth_arith`, `ms_algebra`, `ms_series`, `ms_geometry`, `ms_divisors`.
## Model architecture
| Component | Specification |
|-----------|---------------|
| **Parameters** | 2,882,162,688 (~2.88B) |
| **Layers** | 24 transformer blocks |
| **Hidden size** | 3,072 |
| **Attention heads** | 12 query / 4 key-value (GQA) |
| **Head dimension** | 256 |
| **Feed-forward** | SwiGLU MLP, 8,192 intermediate dim |
| **Attention pattern** | Interleaved local/global causal attention |
| **Normalization** | RMSNorm with QK / V / sandwich norms |
| **Position encoding** | RoPE (ΞΈ = 1,000,000) |
| **Logit stabilization** | Final logit softcap = 15.0 |
| **Context length** | 13,568 tokens (RoPE-extrapolated from a 1,024-ctx base) |
| **Vocabulary size** | 151,669 |
| **Tied embeddings** | Yes (input ↔ output) |
| **Precision** | bf16 safetensor shards + `model.safetensors.index.json` |
## Tokenizer
Reuses the Qwen3 tokenizer (vocab 151,669) via the `Qwen2Tokenizer` compatibility class. The chat template (`chat_template.jinja`) supports `enable_thinking` and parses `<think> … </think>`. Tokenizer files are bundled β€” no extra download needed.
## Evaluation
Internal 4-quadrant probe, 10 items each, graded leniently (greedy for no-think, sampled for with-CoT):
| Quadrant | Score |
|----------|:-----:|
| Math, no-think (greedy) | **10 / 10** |
| Math, with-CoT (sampled) | **10 / 10** |
| General, no-think (greedy) | 7 / 10 |
| General, with-CoT (sampled) | 6 / 10 |
| **Total** | **33 / 40** |
The final weight-soup lifts general no-think from 5β†’7 vs. the pre-soup CoT model (the grammar loop is gone; the "Red Planet = Mars" fact is restored) while keeping math perfect. The residual general misses (e.g. naming the third primary color, a taller-than transitivity puzzle) are genuine capability gaps of the 2.88B base β€” present already before CoT training β€” and are not fixable by souping, decoding, or data changes at this scale.
> **Guardrail note (v2 / v3):** this 4-quadrant probe was measured on v1. Each self-improvement round is gated to **preserve** the general/no-think axes (verified head-to-head) while improving the held-out math path; the honest math signal is the clean SVAMP/ASDiv numbers below, not this small hand-checked probe. The v3 no-think guardrail vs v2: **general no-think unchanged** (identical successes and identical residual base-gaps β€” no new loops, no lost facts) and **math no-think 9/10**, with one regression (divisor-counting) documented under *Limitations*.
## Benchmarks (lm-evaluation-harness)
> **⚠ GSM8K contamination (disclosed 2026-07-10).** The CoT-SFT data tier `easy_gsm8k` was built from a curated GSM8K shard that pooled the **train and test** splits with no split filter, so a large fraction of the GSM8K *test* set β€” with worked `<think>`/`\boxed{}` solutions β€” was seen during training; the v2 STaR step also trained on GSM8K problems. **GSM8K is therefore not a valid held-out benchmark for this model**; its rows below are marked ⚠ and should not be read as held-out. (v3 removes `easy_gsm8k` and trains only on the GSM8K **train** split, keeping the SVAMP/ASDiv eval disjoint β€” but GSM8K itself remains off-limits as a held-out metric.) For an honest held-out math signal, see the **SVAMP / ASDiv** numbers under *Getting more accuracy: test-time compute*. The knowledge/commonsense tasks (ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, SciQ, PIQA, …) are **unaffected** β€” standard academic benchmarks not used in any training stage.
Standard academic benchmarks via [EleutherAI lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) v0.4.11, run through its **vLLM backend** (continuous batching) β€” bf16, greedy for generative tasks (no repetition penalty; see the decoding note above), completion-style prompting (no chat template, for comparability).
> **v1 β†’ v2.** Both checkpoints are shown: the previous release (`soup_blend_a085`, v1) and the v2 upload (`blend_star_a06`). The STaR + soup-recovery step (pipeline step 8) lifts the **math** path while every knowledge/commonsense task is unchanged (the STaR delta is math-only). **v3 (step 9) is again expected to be a math-only upgrade** β€” knowledge/commonsense tasks unchanged β€” with the improvement measured on clean SVAMP/ASDiv, not GSM8K.
### Open LLM Leaderboard v1 + GSM8K
| Benchmark | Setup | v1 | v2 |
|---|---|:---:|:---:|
| ARC-Challenge | 25-shot, acc_norm | 34.0 | 34.2 |
| HellaSwag | 10-shot, acc_norm | 58.7 | 58.6 |
| MMLU | 5-shot, acc | 25.0 | 25.0 |
| TruthfulQA (MC2) | 0-shot | 45.1 | 45.4 |
| WinoGrande | 5-shot, acc | 57.9 | 57.8 |
| GSM8K ⚠*contaminated* | 5-shot, exact-match | 6.2 | 7.2 |
| **Average** *(incl. contaminated GSM8K⚠)* | | 37.8 | 38.0 |
*(The lm-eval table below was **not re-run for v3**; the v3 change is a math-focused CoT-SFT + weight-soup, so these knowledge/commonsense rows are expected within noise of v2. The v3 improvement is measured on clean SVAMP/ASDiv β€” see *Getting more accuracy* above.)*
### Additional benchmarks (0-shot)
| Benchmark | Metric | v1 | v2 |
|---|---|:---:|:---:|
| SciQ | acc | 82.9 | 83.2 |
| PIQA | acc | 72.3 | 72.4 |
| BoolQ | acc | 62.3 | 62.4 |
| ARC-Easy | acc | 55.3 | 55.7 |
| LAMBADA (OpenAI) | acc / ppl | 44.6 / 16.7 | 45.3 / 16.8 |
| OpenBookQA | acc_norm | 35.2 | 34.6 |
| CommonsenseQA | acc | 20.1 | 20.1 |
**Reading these numbers (2.88B, from scratch):**
- **Real commonsense/science signal:** SciQ 82.9, PIQA 72.3, HellaSwag 58.7, BoolQ 62.3, ARC-Easy 55.3 β€” the from-scratch base is not hollow.
- **At/near chance where knowledge-dense pretraining is required:** MMLU 25.0 and CommonsenseQA 20.1 (β‰ˆ 5-way random) β€” the base was pretrained on FineWeb + math only.
- **GSM8K is contaminated β€” do not read it as held-out** (see the note above). For honest held-out math see the clean SVAMP/ASDiv numbers under *Getting more accuracy: test-time compute*.
- These are honest reference points for a small from-scratch reasoner, not a leaderboard flex.
*(MathQA and LogiQA were skipped β€” their loaders use legacy dataset scripts unsupported by current `datasets`; GPQA is a gated dataset.)*
## Inference
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
model_id = "PursuitOfDataScience/Argonne-3.0-think"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
dtype=torch.bfloat16,
)
device = "cuda" if torch.cuda.is_available() else "cpu"
model = model.to(device).eval()
# --- Math / factual: GREEDY, no anti-repeat penalty ---
messages = [{"role": "user", "content": "What is 15% of 80?"}]
prompt_ids = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
enable_thinking=True, # set False to answer directly (also 10/10 on math)
)
input_ids = torch.tensor([prompt_ids], dtype=torch.long, device=device)
output_ids = model.generate(
input_ids,
max_length=input_ids.shape[1] + 512,
do_sample=False, # greedy β€” best for arithmetic
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))
```
For open-ended general chat, mild sampling is fine (`do_sample=True, temperature=0.7, top_p=0.9`).
## Recommended inference settings
| Task | do_sample | temperature | top_p | repetition_penalty | notes |
|------|-----------|-------------|-------|--------------------|-------|
| **Math / arithmetic** | `False` (greedy) | β€” | β€” | **1.0 (off)** | Do **not** use `no_repeat_ngram` β€” it corrupts multi-digit numbers. |
| No-think factual Q&A | `False` (greedy) | β€” | β€” | 1.0 | |
| Open-ended chat / brainstorming | `True` | 0.7 | 0.9 | ≀ 1.1 | Keep any repetition penalty mild. |
- Load with `trust_remote_code=True` so the custom `ArgonneModel` / `argonne2` classes (`model.py`) register.
- Use `apply_chat_template(..., enable_thinking=True)` to elicit a `<think> … </think>` trace, or `enable_thinking=False` for a direct answer. Math is strong on the internal probe in both modes; for held-out benchmark numbers and how to get the most out of the model, see **Getting more accuracy: test-time compute** below.
- The custom `generate` uses `max_length` (total length), not `max_new_tokens`.
- Weights are bf16 safetensor shards with a `model.safetensors.index.json` weight map for sharded loading.
## Getting more accuracy: test-time compute
The **33/40** internal probe (and "10/10 math") is a small, hand-checked sanity probe (10 items per quadrant): it shows the *capability is present*, not competition-grade math. For an **honest held-out** signal we use **SVAMP** and **ASDiv** β€” elementary math-word-problem benchmarks that appear in **none** of the training stages (unlike GSM8K, which is contaminated for this model β€” see the note under *Benchmarks*). The model holds substantial **latent** capability that cheap test-time compute unlocks:
| decoding (v2, with-think) | SVAMP (n=300) | ASDiv (n=300) |
|---|:---:|:---:|
| greedy, single sample | 18.0 | 22.7 |
| **+ budget-forcing** (force-close `</think>` past a think-token budget) | 20.7 | 29.3 |
| **self-consistency** (sample K=32, majority-vote the `\boxed{}` answer) | **36.3** | **51.0** |
| pass@32 (a correct answer is *somewhere* in 32 samples) | **73.3** | **74.3** |
> **v3 (2026-07-11, shipped):** the goal β€” raise the *greedy* row toward the budget-force row (native termination) while holding self-consistency β€” is met:
>
> | decoding (v3, with-think) | SVAMP (n=300) | ASDiv (n=300) |
> |---|:---:|:---:|
> | greedy, single sample | **22.7** (v2 18.0) | **27.3** (v2 22.7) |
> | + budget-forcing | 24.3 | 32.3 |
> | self-consistency (K=32) | 40.3 (v2 36.3) | 48.0 (v2 51.0) |
> | pass@32 | 74.7 (v2 73.3) | 77.0 (v2 74.3) |
>
> Greedy is now β‰ˆ its own budget-forced number (the model **terminates natively** instead of rambling to max length), self-consistency is held, and pass@32 is up β€” a clean single-pass upgrade. Cost: one fragile no-think probe (divisor-counting) regressed β€” see *Limitations*.
> **v4 (2026-07-12, shipped) β€” broad held-out gate (n=400/set, +MAWPS & GSM-Plus, Wilson 95% CIs):** v4 vs v3, both re-measured at n=400. The external-teacher distillation lifts greedy on ASDiv and holds pass@K, but self-consistency regressed on MAWPS β€” a **trade, not a clean win**:
>
> | set β€” greedy / self-cons(K=32) / pass@32 | v3 | v4 |
> |---|:---:|:---:|
> | SVAMP (clean) | 23.5 / 39.5 / 75.5 | 20.5 / 38.0 / 76.3 |
> | ASDiv (clean) | 25.8 / 47.3 / 76.0 | **30.5** / 47.8 / 78.0 |
> | MAWPS (clean) | 21.3 / 45.3 / 78.8 | 21.5 / 40.0 / 78.0 |
> | GSM-Plus (semi-clean, adversarial) | 9.3 / 12.8 / 46.8 | 8.3 / 12.8 / 47.0 |
>
> Greedy ASDiv clearly up (teacher distillation landed), other greedy β‰ˆ flat; self-consistency down on MAWPS (the code/tool/teacher diet cut sample diversity); pass@K held. This is an honest mixed result β€” v3 is retained upstream for rollback if the self-consistency path matters more for your use.
**Read:** on clean, never-trained problems the model is a genuine small reasoner β€” **pass@32 of 73–74%** shows the capability is *real, not memorized*. (A permutation null-control confirms this: against a magnitude-matched chance floor, the pass@32 and self-consistency scores clear the null by tens of points β€” the ceiling is genuine, not lucky small-integer collisions.) The dominant single-shot failure is **non-termination** (~half of greedy traces never close `</think>`), which is exactly why **budget-forcing** (force a stop) and **self-consistency** (vote over K samples) are the biggest free wins β€” and exactly what the v3 training round targets at the weights level.
**Recommended recipe (single model):** sample **K β‰ˆ 16–32** traces (`temperature 0.8, top_p 0.95`), **force-close** the `<think>` block if it runs past a token budget, then **majority-vote** the extracted `\boxed{}` answers β€” roughly **2Γ— greedy at zero training cost**. Among *single-model* pickers this is the best: neither a learned same-base verifier nor confidence-weighted / budget-forced voting beats plain majority vote (reliably judging a solution needs capability this 2.88B base lacks). **For the highest accuracy, add a stronger external reranker** β€” it nearly doubles the vote again (SVAMP/ASDiv ~75, β‰ˆ the pass@32 ceiling); see *Tool use & serving-system capabilities* below.
### Fast inference with vLLM
The custom `argonne2` architecture (a Gemma2-style sandwich-norm layer + Qwen3 qk-norm + an extra value-norm + final logit soft-cap, full-causal, tied embeddings) ports cleanly to **vLLM** β€” validated numerically exact (token-for-token greedy) against the reference `model.py`. vLLM's continuous batching makes the large-K sampling above cheap (~65Γ— the naive per-token HF decode loop, and it actually fills the GPU). The custom-model class is in the training repo.
## Tool use & serving-system capabilities
The v4 round added **tool-calling** and **coding** training data, and studied the two ways to push this
small model further at *inference* (both are serving-system components in the training repo, not baked into
the single shipped card):
- **Tool-calling (calculator / python).** The v4 training checkpoint (`think_v7`) learned the format
cleanly β€” on held-out arithmetic it emits a **100%-valid `<tool_call>`** with the correct expression. But
(a) that behavior **dilutes out** of the conservative shipped soup (v4 emits ~0% tool calls), and (b) if
the model *generates its own* `<tool_response>` it hallucinates the result (~53% correct). The fix is a
**real tool-execution loop** β€” [`reasoning/tool_decode.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_decode.py): stop generation at `</tool_call>`, actually execute the call,
inject the **real** `<tool_response>`, and resume. With real execution, `think_v7` goes **53% β†’ 100%** on
held-out single-operation arithmetic (+46.7 pts). The chat template already renders the standard
`tokenizer.apply_chat_template(..., tools=[…])` format, so the emitted calls are API-compatible. *(This is a
proof on tool-friendly single-op problems β€” it shows the arithmetic-offload mechanism works, not that
multi-step word problems become 100%.)*
- **External-verifier best-of-N β€” the highest-accuracy recipe.** For the pass@K β†’ pass@1 **selection gap**
(the right answer is present in ~77–79% of 32 samples but the same-base majority vote only *picks* ~36–49%),
no *same-base* picker helps (majority vote / confidence-weighting / a same-base learned verifier all saturate
β€” the base can't reliably judge its own solutions). A **stronger external reranker does**: sample K=32 with
the shipped card, then have **Qwen3-4B** *reason through each distinct candidate answer and keep the one it
verifies* (`ext_verify.py`, "reasoned" lens). Validated on the shipped weights (v4), n=500, McNemar vs the
same-base vote (all p<0.001):
| clean set | same-base self-consistency (vote) | **+ Qwen3-4B reranker** | pass@32 (ceiling) |
|---|:---:|:---:|:---:|
| SVAMP | 36.4 | **74.8** (+38.4) | 77.0 |
| ASDiv | 49.0 | **76.0** (+27.0) | 79.2 |
| MAWPS | 38.4 | **58.0** (+19.6) | 77.2 |
The reranker recovers **~97% of the pass@32 ceiling** on SVAMP/ASDiv and lifts the deployable metric **~41 β†’
~70 on average (+28 pts)** β€” including the MAWPS self-consistency that the v4 weights had traded away. Note
the judge must *reason*: a 1-token "Yes/No" verifier scores *below* the vote. This is a **2-model serving
win** β€” the accuracy is the external reasoner's competence applied to this model's candidate set, not the
single 2.88B card unlocked. Tooling: [`reasoning/ext_verify.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/ext_verify.py) (two-phase generate β†’ rerank).
Both point to the same lesson: on this 2.88B base, the remaining gains are **serving-system** moves (a tool
executor, an external reranker) or a **stronger base** β€” not further single-checkpoint weight edits. A broad
n=1000 held-out sweep (SVAMP/ASDiv/MAWPS/GSM-Plus) confirms this at the weight level: v3, v4, and every
weight-soup between them sit on **one flat Pareto line** β€” identical pass@32 (~78%, the capability ceiling is
fixed), trading only greedy ↔ self-consistency. No single-checkpoint edit improves both axes; the reranker
above is the way to cash the fixed ceiling.
## Limitations
- **2.88B parameters** β€” far smaller than frontier models; expect weaker performance on hard reasoning, long-form knowledge, and code.
- **Base-capability gaps persist:** some simple general questions (e.g. listing all three primary colors, multi-hop comparison puzzles) are answered incorrectly regardless of think mode β€” these are limits of the underlying pretraining, not of the reasoning recipe.
- **v3 divisor-counting regression:** v3's short-trace diet erodes one fragile synthetic capability β€” counting the number of divisors of an integer via prime factorization (e.g. "how many divisors does 12 have?") is answered correctly by v2 but garbled by v3. This is a single narrow templated pattern; broad held-out math (SVAMP/ASDiv) improved. Noted for transparency.
- **v4 self-consistency trade:** v4's external-teacher distillation + tool/coding diet improved greedy ASDiv (+4.7) and native termination but cut sample diversity, **regressing self-consistency ~6 pts** (mainly MAWPS) vs v3 β€” an honest mixed update, not a clean win. If your use relies on K-sample majority voting, v3 may serve you better; v3 (`x_v6v2_040`) is retained upstream for rollback. Tool-calling and coding data were added but are **not** in this shipped card (see *Tool use & serving-system capabilities*).
- **Anti-repeat decoding breaks math** (see notes above) β€” a property of this small model's arithmetic being token-repetition-heavy.
- **The selection gap is base-limited:** the correct answer is *present* in ~73% of 32 samples but only *picked* ~36–51% of the time; no cheap picker (majority vote, confidence-weighting, a same-base learned verifier) closes this gap β€” it caps at the base's ability to *judge* a solution.
- Context extended via RoPE extrapolation; very-long-context retrieval may degrade.
- English-centric SFT/DPO data; limited multilingual ability. No safety filtering or content moderation has been applied.
## Source code
All training code is on the `main` branch of the Argonne repo:
https://github.com/PursuitOfDataScience/ArgonneAI
The custom architecture (`ArgonneModel`, `model_type = argonne2`) is bundled with this
model as `model.py`. The exact scripts that produced this checkpoint, by pipeline stage:
| Stage | Script(s) |
|-------|-----------|
| Architecture | [`model.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/model.py) |
| Pretraining (base) | [`pretrain.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/pretrain.py) |
| Intermix midtraining | [`preprocess_finemath.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/preprocess_finemath.py) β†’ [`reasoning/build_intermix.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_intermix.py) β†’ [`midtraining.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/midtraining.py) |
| Soup base (training-free) | [`reasoning/build_soup_base.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_soup_base.py) |
| SFT | [`sft.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/sft.py) |
| DPO | [`dpo.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/dpo.py) |
| CoT-SFT data (v1/v2) | [`reasoning/build_sft_mix.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_sft_mix.py), [`reasoning/build_mix_v3.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v3.py) |
| CoT-SFT data (v3, short-only) | [`reasoning/build_mix_v6.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v6.py) |
| CoT-SFT data (v4: teacher + tool + coding) | [`reasoning/gen_teacher.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/gen_teacher.py) (Qwen3-4B teacher), [`reasoning/build_mix_v7.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v7.py) |
| CoT-SFT training | [`reasoning/cot-sft.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/cot-sft.py) |
| Weight-soups (training-free) | [`reasoning/build_ckpt_soup.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_ckpt_soup.py) |
| Deployment | [`reasoning/deploy_hf.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/deploy_hf.py) |
| Serving-system (research) | [`reasoning/ext_verify.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/ext_verify.py) (external-verifier best-of-N), [`reasoning/tool_decode.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_decode.py) (tool-execution loop), [`reasoning/vllm_argonne.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/vllm_argonne.py) (vLLM port) |
| Evaluation | [`reasoning/clean_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/clean_eval.py) (clean SVAMP/ASDiv/MAWPS + GSM-Plus, Wilson CIs β€” the honest judge), [`reasoning/code_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/code_eval.py) (HumanEval), [`reasoning/tool_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_eval.py) (tool-call format), [`reasoning/eval_numeracy.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/eval_numeracy.py) |
| Full writeup / recipe | [`reasoning/thinking_training.md`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/thinking_training.md) |
The SLURM launcher scripts that wire these together with per-stage hyperparameters are
intentionally untracked in the repo (they carry cluster-specific paths); every stage's
hyperparameters are recorded in [`reasoning/thinking_training.md`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/thinking_training.md) and summarized in the
tables above.
## Citation
```bibtex
@misc{argonne30think,
author = {PursuitOfDataScience},
title = {Argonne 3.0-think},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/PursuitOfDataScience/Argonne-3.0-think}
}
```