Text Generation
Transformers
Safetensors
English
argonne2
feature-extraction
causal-lm
transformer
argonne
reasoning
chain-of-thought
think
sft
dpo
model-soup
tool-use
distillation
conversational
custom_code
Instructions to use PursuitOfDataScience/Argonne-3.0-think with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PursuitOfDataScience/Argonne-3.0-think with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PursuitOfDataScience/Argonne-3.0-think", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("PursuitOfDataScience/Argonne-3.0-think", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PursuitOfDataScience/Argonne-3.0-think with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PursuitOfDataScience/Argonne-3.0-think" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PursuitOfDataScience/Argonne-3.0-think", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PursuitOfDataScience/Argonne-3.0-think
- SGLang
How to use PursuitOfDataScience/Argonne-3.0-think with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PursuitOfDataScience/Argonne-3.0-think" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PursuitOfDataScience/Argonne-3.0-think", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PursuitOfDataScience/Argonne-3.0-think" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PursuitOfDataScience/Argonne-3.0-think", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PursuitOfDataScience/Argonne-3.0-think with Docker Model Runner:
docker model run hf.co/PursuitOfDataScience/Argonne-3.0-think
| license: apache-2.0 | |
| language: | |
| - en | |
| library_name: transformers | |
| base_model: | |
| - PursuitOfDataScience/argonne-3.0-base | |
| tags: | |
| - text-generation | |
| - causal-lm | |
| - transformer | |
| - argonne | |
| - reasoning | |
| - chain-of-thought | |
| - think | |
| - sft | |
| - dpo | |
| - model-soup | |
| - tool-use | |
| - distillation | |
| pipeline_tag: text-generation | |
| # Argonne 3.0-think | |
| **Argonne 3.0-think** is a 2.88B-parameter **reasoning** model from the Argonne 3.x family. It supports an explicit chain-of-thought "think" mode (`<think> β¦ </think>`) and answers directly in no-think mode. It is built entirely from [Argonne 3.0-base](https://huggingface.co/PursuitOfDataScience/argonne-3.0-base) through a multi-stage pipeline that ends in **training-free weight-soups** β one to repair the pretraining base, one to reconcile reasoning with general chat, and further recovery soups after each round of self-improvement. | |
| On a 40-item internal 4-quadrant probe ({math, general} Γ {no-think, with-CoT}) it scores **33/40**, with strong arithmetic in *both* modes while remaining a loop-free general chat model. (This is a small hand-checked probe β for held-out benchmark numbers and how to get the most out of the model, see *Getting more accuracy: test-time compute* below.) | |
| π **Companion blog post:** [*How I Taught a Small Language Model to Reason*](https://youzhi.netlify.app/post/2026-07-21-reasoning-model/reasoning-model/) β a plain-language walkthrough of how this model was built (calibrated chain-of-thought data + training-free weight-soups), for readers who want the story without the code. | |
| > **Update (2026-07-12) β v4.** A modest, **external-teacher-distillation** reasoning update. A stronger open model (**Qwen3-4B**) solved the GSM8K **train** split (contamination-safe, 87.8% correct β 3,714 short worked solutions); those distilled traces β plus new **tool-calling** and **coding** data tiers β were CoT-SFT'd from `dpo_soup` (β `think_v7`) and cross-souped with v3 (**v4 = `0.3Β·think_v7 + 0.7Β·v3`**). On clean held-out math the teacher distillation raised **greedy ASDiv 25.8β30.5** and strengthened native `</think>` termination; SVAMP greedy is ~flat and **self-consistency regressed ~6 pts** (a trade β see *Getting more accuracy* and *Limitations*). Honestly, v4 is a **mixed, reasoning-focused update**, not a clean dominance over v3. The new **tool-calling** capability was learned cleanly (the training checkpoint emits 100% valid `<tool_call>`s) but is **not baked into this shipped card** (it dilutes out of the conservative soup and, without a real executor, the model hallucinates tool outputs); it is exposed instead as a **serving-side tool-execution loop** β see *Tool use & serving-system capabilities*. Coding did not measurably transfer at this scale (HumanEval β 0, base-capacity-limited). See step 10 in *How it was trained*. v3 (`x_v6v2_040`) is retained upstream for rollback. | |
| > **Update (2026-07-11) β v3.** This release targets the model's #1 *deployable* failure: **non-termination** β on clean held-out problems (**SVAMP/ASDiv**) ~50β60% of sampled `<think>` traces never close `</think>`, even on one-step problems (the model over-thinks). v3 distills the fix into the **weights**: a **SHORT-trace** CoT-SFT (train *only* on β€768-token closed-correct traces) so greedy decoding **natively terminates** with the answer β turning the decode-time budget-forcing win into a weights-level one β plus a contamination-safe grade-school **procedure** tier; the result is then **weight-soupβfused with v2** to keep v2's self-consistency diversity. On clean held-out math (SVAMP / ASDiv, greedy): **18.0β22.7 / 22.7β27.3**; self-consistency held (**36.3β40.3 / 51.0β48.0**); pass@32 up (**73.3β74.7 / 74.3β77.0**). See step 9 in *How it was trained*. *(GSM8K remains contaminated for this model β see the note under *Benchmarks* β so held-out SVAMP/ASDiv are the honest judge.)* | |
| > **Update (2026-07-08) β v2.** The model was improved by a round of **STaR self-improvement** (rejection-sampling fine-tuning on the model's *own* verified-correct GSM8K `<think>` traces) followed by a **third training-free soup** that recovers the no-think arithmetic that step regressed β improving held-out math (on the clean **SVAMP/ASDiv** benchmarks, greedy ~9%β~18% / ~18%β~23%; see *Benchmarks*) while keeping general chat and no-think math intact. *(Earlier versions of this card reported a GSM8K 2%β~7.5% gain; GSM8K is now known to be **contaminated** for this model β see the note under *Benchmarks* β so held-out SVAMP/ASDiv are used instead.)* See step 8 in *How it was trained*. Also: `config.json` now ships an `auto_map`, so `AutoModelForCausalLM.from_pretrained(..., trust_remote_code=True)` loads it directly (no manual `import model` needed). | |
| > **Decoding note (important):** for math/arithmetic, use **greedy decoding** and **do not** apply an aggressive `repetition_penalty` / `no_repeat_ngram_size`. Anti-repeat decoding corrupts multi-digit arithmetic (it blocks the model from re-emitting digits it just used, e.g. turning `80 / 2` into `8 / 2`) and collapses math accuracy. See *Recommended inference settings* below. | |
| ## How it was trained | |
| Argonne 3.0-think is the endpoint of the following pipeline. Every stage runs at 13,568-token context with RoPE ΞΈ = 1,000,000 (the base is RoPE-extrapolated from a 1,024-ctx pretraining run). | |
| | # | Stage | Init from | Data | What it does | | |
| |---|-------|-----------|------|--------------| | |
| | 1 | **Pretraining** (from scratch) | β | [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb) (~76B tokens, 1,024 ctx) | The 2.88B base, [Argonne 3.0-base](https://huggingface.co/PursuitOfDataScience/argonne-3.0-base). | | |
| | 2 | **Intermix midtraining** | Argonne 3.0-base | FineWeb (crawl `CC-MAIN-2025-21`) + [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) subset `finemath-4plus`, mixed 50:50 β 60:40 by document | Injects grade-school + competition numeracy the pure-FineWeb base lacked. | | |
| | 3 | **Soup base** (training-free) | β | *(weight interpolation, no data)* | `0.35 Β· base + 0.65 Β· intermix`. A linear weight-soup of the two same-lineage checkpoints β recovers general knowledge the midtraining LR over-wrote while keeping the math. This is the reasoning-ready base. | | |
| | 4 | **SFT** | Soup base | [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) (`train_sft`) | Instruction / multi-turn chat. | | |
| | 5 | **DPO** | SFT model | [KatoHF/chatbot_arena_binarized](https://huggingface.co/datasets/KatoHF/chatbot_arena_binarized) | Preference alignment on real user comparisons. Kept as `dpo_soup` (the general-healthy pre-CoT checkpoint). | | |
| | 6 | **CoT-SFT** | DPO model | `cot_sft_mix_v3` (see table below) | Teaches explicit `<think>` chain-of-thought; brings math to 10/10 but regresses general chat (over-reasons / loops). | | |
| | 7 | **Final weight-soup** (training-free) | β | *(weight interpolation, no data)* | `0.15 Β· (DPO model) + 0.85 Β· (CoT-SFT model)`. Fractionally "un-applies" the CoT stage's general-chat regression while keeping its math. The `soup_blend_a085` reasoning model (v1). | | |
| | 8 | **STaR self-improvement + soup-recovery** (v2, 2026-07-08) | step 7 model | Rejection-sampled, **verified-correct** GSM8K `<think>` traces (the model's own, filtered by a `\boxed{}` checker) + a general/no-think anchor; then a training-free soup `0.4 Β· (step 7) + 0.6 Β· (STaR model)` | STaR: fine-tunes on the model's own correct solutions. The third soup recovers the no-think multi-step arithmetic the STaR step regressed (e.g. divisor-counting), keeping general chat intact. | | |
| | 9 | **SHORT-trace termination + procedure distillation β cross-soup** (v3, 2026-07-11) | `dpo_soup` (step-5 pre-CoT) β CoT-SFT β `think_v6`; then cross-soup with v2 | `cot_sft_mix_v6`: **short-only** (β€768-token) closed-correct traces β drops the contaminated `easy_gsm8k` and adds a contamination-safe **gsm8k-train-only** short procedure tier (canonicalized `\boxed{}` close). Ship = training-free cross-soup **`0.4 Β· think_v6 + 0.6 Β· v2`** | Short-only SFT makes greedy **natively** close `</think>` with the answer (turns the decode-time budget-force win into a weights-level win) and installs grade-school procedure; the cross-soup with v2 keeps v2's self-consistency diversity. Clean SVAMP/ASDiv greedy +4.7/+4.6, self-cons held. | | |
| | 10 | **External-teacher distillation + tool/coding β cross-soup** (v4, 2026-07-12) | `dpo_soup` β CoT-SFT (`cot_sft_mix_v7`) β `think_v7`; then cross-soup with v3 | `cot_sft_mix_v7` (~35.6k): the v6 short-only backbone + **`teacher_gsm`** (Qwen3-4B's correct gsm8k-**train** solutions, 3,714) + **`code_magicoder`** (decontaminated code) + synthesized **`tool_calc`** (calculator/python tool-use). Ship = training-free cross-soup **`0.3 Β· think_v7 + 0.6Β·β¦ (0.7 Β· v3)`** | Distills a *stronger* teacher's worked solutions into the model (self-STaR had saturated). Lifts greedy ASDiv (25.8β30.5) + native termination; self-consistency regressed ~6 pts (a trade). Tool-calling learned in `think_v7` but not shipped in the soup (see *Tool use*); coding did not transfer at 2.88B. | | |
| **The soups are the key idea.** The CoT-SFT model (step 6) and the DPO model (step 5) live in the same optimization basin (step 6 fine-tunes step 5), so `CoT-SFT = DPO + Ξ` for a chain-of-thought weight-delta Ξ. Blending back 15% of the DPO weights (`0.85 Β· CoT-SFT + 0.15 Β· DPO`) scales that delta down just enough to remove the loop/forgetting pathology the CoT stage introduced β the grammar loop disappears, lost facts (e.g. "the Red Planet is Mars") return β **without hurting the 10/10 math**. Ξ± = 0.85 is a knee: lower values recover even more general chat but start breaking `<think>` trace-closure (the CoT format itself lives in Ξ). The same recovery trick is reused after each self-improvement round (steps 8 and 9). | |
| ### Chain-of-thought SFT data | |
| **v3 (`cot_sft_mix_v6`, ~26.4k examples, all β€768 tokens).** Deliberately **short-only** β training on short, terminating traces is what teaches greedy to natively close `</think>` (the fix for non-termination). The contaminated `easy_gsm8k` tier is dropped and replaced with a contamination-safe **gsm8k-train-only** procedure tier; every long tier from v3-mix is removed. | |
| | Tier | Share | Source | | |
| |------|-----:|--------| | |
| | `direct_tulu` (no-think chat) | 30.3% | [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) | | |
| | `gsm8k_train_short` (NEW, contamination-safe) | 16.4% | [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`) β **train split only**, β€512 tok, verified closed+boxed, canonicalized to the deployed `\boxed{}` close | | |
| | `synth_arith` | 9.5% | Synthetic, correct-by-construction | | |
| | `ms_algebra` / `ms_series` / `ms_geometry` / `ms_divisors` | ~18% | Synthetic multi-step (Python-verified) | | |
| | `med_math` | 7.6% | [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) (levels 1β3) | | |
| | `gen_ultrachat` / `hq_opus` / `med_openmath` / `hard_strict` | small | ultrachat-derived / Opus / OpenMathReasoning / MiniMax-M2.1, all β€768 tok | | |
| **v1/v2 (`cot_sft_mix_v3`, ~113k examples).** The mix used for pipeline step 6 (still the basis of v1/v2): | |
| | Tier | Rows | Source | | |
| |------|-----:|--------| | |
| | `direct_tulu` (no-think chat) | 34,000 | [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) | | |
| | `synth_arith` | 15,000 | Synthetic, correct-by-construction | | |
| | `gen_ultrachat` (CoT-augmented) | 15,000 | Derived from [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) | | |
| | `hard_strict` | 12,000 | [PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts](https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts) (open-ended, strict-filtered) | | |
| | `easy_gsm8k` *(dropped in v3 β contaminated)* | 8,402 | [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`), curated with `<think>` / `\boxed{}` | | |
| | `med_math` | 5,729 | [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) (levels 1β3) | | |
| | `ms_algebra` | 5,000 | Synthetic multi-step (Python-verified) | | |
| | `ms_series` | 5,000 | Synthetic multi-step (Python-verified) | | |
| | `ms_geometry` | 5,000 | Synthetic multi-step (Python-verified) | | |
| | `med_openmath` | 4,620 | Problems from [nvidia/OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) (CoT subset); solutions regenerated | | |
| | `hq_opus` | 2,300 | [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) | | |
| | `ms_divisors` | 1,290 | Synthetic multi-step (Python-verified) | | |
| Every example is capped at 4,000 tokens (768 for v3); all synthetic and `ms_*` traces are re-verified with a `\boxed{}` answer extractor. | |
| ### Key hyperparameters | |
| | Stage | LR | Epochs | Effective batch | Ctx | ΞΈ | | |
| |-------|----|-------:|-----------------|----:|---| | |
| | SFT | 2e-5 | 1 | 18 | 13,568 | 1e6 | | |
| | DPO | 1e-6 (Ξ² = 0.03) | 1 | 8 | 13,568 | 1e6 | | |
| | CoT-SFT (v1/v2) | 1e-5 | 1 | 12 (3Γ H200 DDP) | 13,568 | 1e6 | | |
| | CoT-SFT (v3, short-only) | 1e-5 | 1 | 12 (1Γ H100) | 13,568 | 1e6 | | |
| ## Datasets used (all stages) | |
| 1. [HuggingFaceFW/fineweb](https://huggingface.co/datasets/HuggingFaceFW/fineweb) β pretraining + intermix general half (crawl `CC-MAIN-2025-21`). | |
| 2. [HuggingFaceTB/finemath](https://huggingface.co/datasets/HuggingFaceTB/finemath) (`finemath-4plus`) β intermix math half. | |
| 3. [HuggingFaceH4/ultrachat_200k](https://huggingface.co/datasets/HuggingFaceH4/ultrachat_200k) β SFT, and the `gen_ultrachat` CoT tier. | |
| 4. [KatoHF/chatbot_arena_binarized](https://huggingface.co/datasets/KatoHF/chatbot_arena_binarized) β DPO preferences. | |
| 5. [allenai/tulu-3-sft-mixture](https://huggingface.co/datasets/allenai/tulu-3-sft-mixture) β CoT `direct_tulu` no-think tier. | |
| 6. [openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k) (`main`) β CoT `easy_gsm8k` tier (v1/v2, contaminated) and the v3 `gsm8k_train_short` tier (train split only). | |
| 7. [nvidia/OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) β CoT `med_openmath` tier. | |
| 8. [nlile/hendrycks-MATH-benchmark](https://huggingface.co/datasets/nlile/hendrycks-MATH-benchmark) β CoT `med_math` tier. | |
| 9. [nohurry/Opus-4.6-Reasoning-3000x-filtered](https://huggingface.co/datasets/nohurry/Opus-4.6-Reasoning-3000x-filtered) β CoT `hq_opus` tier. | |
| 10. [PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts](https://huggingface.co/datasets/PursuitOfDataScience/MiniMax-M2.1-Mixture-of-Thoughts) β CoT `hard_strict` tier. | |
| Fully synthetic (no external dataset): `synth_arith`, `ms_algebra`, `ms_series`, `ms_geometry`, `ms_divisors`. | |
| ## Model architecture | |
| | Component | Specification | | |
| |-----------|---------------| | |
| | **Parameters** | 2,882,162,688 (~2.88B) | | |
| | **Layers** | 24 transformer blocks | | |
| | **Hidden size** | 3,072 | | |
| | **Attention heads** | 12 query / 4 key-value (GQA) | | |
| | **Head dimension** | 256 | | |
| | **Feed-forward** | SwiGLU MLP, 8,192 intermediate dim | | |
| | **Attention pattern** | Interleaved local/global causal attention | | |
| | **Normalization** | RMSNorm with QK / V / sandwich norms | | |
| | **Position encoding** | RoPE (ΞΈ = 1,000,000) | | |
| | **Logit stabilization** | Final logit softcap = 15.0 | | |
| | **Context length** | 13,568 tokens (RoPE-extrapolated from a 1,024-ctx base) | | |
| | **Vocabulary size** | 151,669 | | |
| | **Tied embeddings** | Yes (input β output) | | |
| | **Precision** | bf16 safetensor shards + `model.safetensors.index.json` | | |
| ## Tokenizer | |
| Reuses the Qwen3 tokenizer (vocab 151,669) via the `Qwen2Tokenizer` compatibility class. The chat template (`chat_template.jinja`) supports `enable_thinking` and parses `<think> β¦ </think>`. Tokenizer files are bundled β no extra download needed. | |
| ## Evaluation | |
| Internal 4-quadrant probe, 10 items each, graded leniently (greedy for no-think, sampled for with-CoT): | |
| | Quadrant | Score | | |
| |----------|:-----:| | |
| | Math, no-think (greedy) | **10 / 10** | | |
| | Math, with-CoT (sampled) | **10 / 10** | | |
| | General, no-think (greedy) | 7 / 10 | | |
| | General, with-CoT (sampled) | 6 / 10 | | |
| | **Total** | **33 / 40** | | |
| The final weight-soup lifts general no-think from 5β7 vs. the pre-soup CoT model (the grammar loop is gone; the "Red Planet = Mars" fact is restored) while keeping math perfect. The residual general misses (e.g. naming the third primary color, a taller-than transitivity puzzle) are genuine capability gaps of the 2.88B base β present already before CoT training β and are not fixable by souping, decoding, or data changes at this scale. | |
| > **Guardrail note (v2 / v3):** this 4-quadrant probe was measured on v1. Each self-improvement round is gated to **preserve** the general/no-think axes (verified head-to-head) while improving the held-out math path; the honest math signal is the clean SVAMP/ASDiv numbers below, not this small hand-checked probe. The v3 no-think guardrail vs v2: **general no-think unchanged** (identical successes and identical residual base-gaps β no new loops, no lost facts) and **math no-think 9/10**, with one regression (divisor-counting) documented under *Limitations*. | |
| ## Benchmarks (lm-evaluation-harness) | |
| > **β GSM8K contamination (disclosed 2026-07-10).** The CoT-SFT data tier `easy_gsm8k` was built from a curated GSM8K shard that pooled the **train and test** splits with no split filter, so a large fraction of the GSM8K *test* set β with worked `<think>`/`\boxed{}` solutions β was seen during training; the v2 STaR step also trained on GSM8K problems. **GSM8K is therefore not a valid held-out benchmark for this model**; its rows below are marked β and should not be read as held-out. (v3 removes `easy_gsm8k` and trains only on the GSM8K **train** split, keeping the SVAMP/ASDiv eval disjoint β but GSM8K itself remains off-limits as a held-out metric.) For an honest held-out math signal, see the **SVAMP / ASDiv** numbers under *Getting more accuracy: test-time compute*. The knowledge/commonsense tasks (ARC, HellaSwag, MMLU, TruthfulQA, WinoGrande, SciQ, PIQA, β¦) are **unaffected** β standard academic benchmarks not used in any training stage. | |
| Standard academic benchmarks via [EleutherAI lm-evaluation-harness](https://github.com/EleutherAI/lm-evaluation-harness) v0.4.11, run through its **vLLM backend** (continuous batching) β bf16, greedy for generative tasks (no repetition penalty; see the decoding note above), completion-style prompting (no chat template, for comparability). | |
| > **v1 β v2.** Both checkpoints are shown: the previous release (`soup_blend_a085`, v1) and the v2 upload (`blend_star_a06`). The STaR + soup-recovery step (pipeline step 8) lifts the **math** path while every knowledge/commonsense task is unchanged (the STaR delta is math-only). **v3 (step 9) is again expected to be a math-only upgrade** β knowledge/commonsense tasks unchanged β with the improvement measured on clean SVAMP/ASDiv, not GSM8K. | |
| ### Open LLM Leaderboard v1 + GSM8K | |
| | Benchmark | Setup | v1 | v2 | | |
| |---|---|:---:|:---:| | |
| | ARC-Challenge | 25-shot, acc_norm | 34.0 | 34.2 | | |
| | HellaSwag | 10-shot, acc_norm | 58.7 | 58.6 | | |
| | MMLU | 5-shot, acc | 25.0 | 25.0 | | |
| | TruthfulQA (MC2) | 0-shot | 45.1 | 45.4 | | |
| | WinoGrande | 5-shot, acc | 57.9 | 57.8 | | |
| | GSM8K β *contaminated* | 5-shot, exact-match | 6.2 | 7.2 | | |
| | **Average** *(incl. contaminated GSM8Kβ )* | | 37.8 | 38.0 | | |
| *(The lm-eval table below was **not re-run for v3**; the v3 change is a math-focused CoT-SFT + weight-soup, so these knowledge/commonsense rows are expected within noise of v2. The v3 improvement is measured on clean SVAMP/ASDiv β see *Getting more accuracy* above.)* | |
| ### Additional benchmarks (0-shot) | |
| | Benchmark | Metric | v1 | v2 | | |
| |---|---|:---:|:---:| | |
| | SciQ | acc | 82.9 | 83.2 | | |
| | PIQA | acc | 72.3 | 72.4 | | |
| | BoolQ | acc | 62.3 | 62.4 | | |
| | ARC-Easy | acc | 55.3 | 55.7 | | |
| | LAMBADA (OpenAI) | acc / ppl | 44.6 / 16.7 | 45.3 / 16.8 | | |
| | OpenBookQA | acc_norm | 35.2 | 34.6 | | |
| | CommonsenseQA | acc | 20.1 | 20.1 | | |
| **Reading these numbers (2.88B, from scratch):** | |
| - **Real commonsense/science signal:** SciQ 82.9, PIQA 72.3, HellaSwag 58.7, BoolQ 62.3, ARC-Easy 55.3 β the from-scratch base is not hollow. | |
| - **At/near chance where knowledge-dense pretraining is required:** MMLU 25.0 and CommonsenseQA 20.1 (β 5-way random) β the base was pretrained on FineWeb + math only. | |
| - **GSM8K is contaminated β do not read it as held-out** (see the note above). For honest held-out math see the clean SVAMP/ASDiv numbers under *Getting more accuracy: test-time compute*. | |
| - These are honest reference points for a small from-scratch reasoner, not a leaderboard flex. | |
| *(MathQA and LogiQA were skipped β their loaders use legacy dataset scripts unsupported by current `datasets`; GPQA is a gated dataset.)* | |
| ## Inference | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| import torch | |
| model_id = "PursuitOfDataScience/Argonne-3.0-think" | |
| tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True) | |
| model = AutoModelForCausalLM.from_pretrained( | |
| model_id, | |
| trust_remote_code=True, | |
| dtype=torch.bfloat16, | |
| ) | |
| device = "cuda" if torch.cuda.is_available() else "cpu" | |
| model = model.to(device).eval() | |
| # --- Math / factual: GREEDY, no anti-repeat penalty --- | |
| messages = [{"role": "user", "content": "What is 15% of 80?"}] | |
| prompt_ids = tokenizer.apply_chat_template( | |
| messages, | |
| tokenize=True, | |
| add_generation_prompt=True, | |
| enable_thinking=True, # set False to answer directly (also 10/10 on math) | |
| ) | |
| input_ids = torch.tensor([prompt_ids], dtype=torch.long, device=device) | |
| output_ids = model.generate( | |
| input_ids, | |
| max_length=input_ids.shape[1] + 512, | |
| do_sample=False, # greedy β best for arithmetic | |
| ) | |
| print(tokenizer.decode(output_ids[0], skip_special_tokens=True)) | |
| ``` | |
| For open-ended general chat, mild sampling is fine (`do_sample=True, temperature=0.7, top_p=0.9`). | |
| ## Recommended inference settings | |
| | Task | do_sample | temperature | top_p | repetition_penalty | notes | | |
| |------|-----------|-------------|-------|--------------------|-------| | |
| | **Math / arithmetic** | `False` (greedy) | β | β | **1.0 (off)** | Do **not** use `no_repeat_ngram` β it corrupts multi-digit numbers. | | |
| | No-think factual Q&A | `False` (greedy) | β | β | 1.0 | | | |
| | Open-ended chat / brainstorming | `True` | 0.7 | 0.9 | β€ 1.1 | Keep any repetition penalty mild. | | |
| - Load with `trust_remote_code=True` so the custom `ArgonneModel` / `argonne2` classes (`model.py`) register. | |
| - Use `apply_chat_template(..., enable_thinking=True)` to elicit a `<think> β¦ </think>` trace, or `enable_thinking=False` for a direct answer. Math is strong on the internal probe in both modes; for held-out benchmark numbers and how to get the most out of the model, see **Getting more accuracy: test-time compute** below. | |
| - The custom `generate` uses `max_length` (total length), not `max_new_tokens`. | |
| - Weights are bf16 safetensor shards with a `model.safetensors.index.json` weight map for sharded loading. | |
| ## Getting more accuracy: test-time compute | |
| The **33/40** internal probe (and "10/10 math") is a small, hand-checked sanity probe (10 items per quadrant): it shows the *capability is present*, not competition-grade math. For an **honest held-out** signal we use **SVAMP** and **ASDiv** β elementary math-word-problem benchmarks that appear in **none** of the training stages (unlike GSM8K, which is contaminated for this model β see the note under *Benchmarks*). The model holds substantial **latent** capability that cheap test-time compute unlocks: | |
| | decoding (v2, with-think) | SVAMP (n=300) | ASDiv (n=300) | | |
| |---|:---:|:---:| | |
| | greedy, single sample | 18.0 | 22.7 | | |
| | **+ budget-forcing** (force-close `</think>` past a think-token budget) | 20.7 | 29.3 | | |
| | **self-consistency** (sample K=32, majority-vote the `\boxed{}` answer) | **36.3** | **51.0** | | |
| | pass@32 (a correct answer is *somewhere* in 32 samples) | **73.3** | **74.3** | | |
| > **v3 (2026-07-11, shipped):** the goal β raise the *greedy* row toward the budget-force row (native termination) while holding self-consistency β is met: | |
| > | |
| > | decoding (v3, with-think) | SVAMP (n=300) | ASDiv (n=300) | | |
| > |---|:---:|:---:| | |
| > | greedy, single sample | **22.7** (v2 18.0) | **27.3** (v2 22.7) | | |
| > | + budget-forcing | 24.3 | 32.3 | | |
| > | self-consistency (K=32) | 40.3 (v2 36.3) | 48.0 (v2 51.0) | | |
| > | pass@32 | 74.7 (v2 73.3) | 77.0 (v2 74.3) | | |
| > | |
| > Greedy is now β its own budget-forced number (the model **terminates natively** instead of rambling to max length), self-consistency is held, and pass@32 is up β a clean single-pass upgrade. Cost: one fragile no-think probe (divisor-counting) regressed β see *Limitations*. | |
| > **v4 (2026-07-12, shipped) β broad held-out gate (n=400/set, +MAWPS & GSM-Plus, Wilson 95% CIs):** v4 vs v3, both re-measured at n=400. The external-teacher distillation lifts greedy on ASDiv and holds pass@K, but self-consistency regressed on MAWPS β a **trade, not a clean win**: | |
| > | |
| > | set β greedy / self-cons(K=32) / pass@32 | v3 | v4 | | |
| > |---|:---:|:---:| | |
| > | SVAMP (clean) | 23.5 / 39.5 / 75.5 | 20.5 / 38.0 / 76.3 | | |
| > | ASDiv (clean) | 25.8 / 47.3 / 76.0 | **30.5** / 47.8 / 78.0 | | |
| > | MAWPS (clean) | 21.3 / 45.3 / 78.8 | 21.5 / 40.0 / 78.0 | | |
| > | GSM-Plus (semi-clean, adversarial) | 9.3 / 12.8 / 46.8 | 8.3 / 12.8 / 47.0 | | |
| > | |
| > Greedy ASDiv clearly up (teacher distillation landed), other greedy β flat; self-consistency down on MAWPS (the code/tool/teacher diet cut sample diversity); pass@K held. This is an honest mixed result β v3 is retained upstream for rollback if the self-consistency path matters more for your use. | |
| **Read:** on clean, never-trained problems the model is a genuine small reasoner β **pass@32 of 73β74%** shows the capability is *real, not memorized*. (A permutation null-control confirms this: against a magnitude-matched chance floor, the pass@32 and self-consistency scores clear the null by tens of points β the ceiling is genuine, not lucky small-integer collisions.) The dominant single-shot failure is **non-termination** (~half of greedy traces never close `</think>`), which is exactly why **budget-forcing** (force a stop) and **self-consistency** (vote over K samples) are the biggest free wins β and exactly what the v3 training round targets at the weights level. | |
| **Recommended recipe (single model):** sample **K β 16β32** traces (`temperature 0.8, top_p 0.95`), **force-close** the `<think>` block if it runs past a token budget, then **majority-vote** the extracted `\boxed{}` answers β roughly **2Γ greedy at zero training cost**. Among *single-model* pickers this is the best: neither a learned same-base verifier nor confidence-weighted / budget-forced voting beats plain majority vote (reliably judging a solution needs capability this 2.88B base lacks). **For the highest accuracy, add a stronger external reranker** β it nearly doubles the vote again (SVAMP/ASDiv ~75, β the pass@32 ceiling); see *Tool use & serving-system capabilities* below. | |
| ### Fast inference with vLLM | |
| The custom `argonne2` architecture (a Gemma2-style sandwich-norm layer + Qwen3 qk-norm + an extra value-norm + final logit soft-cap, full-causal, tied embeddings) ports cleanly to **vLLM** β validated numerically exact (token-for-token greedy) against the reference `model.py`. vLLM's continuous batching makes the large-K sampling above cheap (~65Γ the naive per-token HF decode loop, and it actually fills the GPU). The custom-model class is in the training repo. | |
| ## Tool use & serving-system capabilities | |
| The v4 round added **tool-calling** and **coding** training data, and studied the two ways to push this | |
| small model further at *inference* (both are serving-system components in the training repo, not baked into | |
| the single shipped card): | |
| - **Tool-calling (calculator / python).** The v4 training checkpoint (`think_v7`) learned the format | |
| cleanly β on held-out arithmetic it emits a **100%-valid `<tool_call>`** with the correct expression. But | |
| (a) that behavior **dilutes out** of the conservative shipped soup (v4 emits ~0% tool calls), and (b) if | |
| the model *generates its own* `<tool_response>` it hallucinates the result (~53% correct). The fix is a | |
| **real tool-execution loop** β [`reasoning/tool_decode.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_decode.py): stop generation at `</tool_call>`, actually execute the call, | |
| inject the **real** `<tool_response>`, and resume. With real execution, `think_v7` goes **53% β 100%** on | |
| held-out single-operation arithmetic (+46.7 pts). The chat template already renders the standard | |
| `tokenizer.apply_chat_template(..., tools=[β¦])` format, so the emitted calls are API-compatible. *(This is a | |
| proof on tool-friendly single-op problems β it shows the arithmetic-offload mechanism works, not that | |
| multi-step word problems become 100%.)* | |
| - **External-verifier best-of-N β the highest-accuracy recipe.** For the pass@K β pass@1 **selection gap** | |
| (the right answer is present in ~77β79% of 32 samples but the same-base majority vote only *picks* ~36β49%), | |
| no *same-base* picker helps (majority vote / confidence-weighting / a same-base learned verifier all saturate | |
| β the base can't reliably judge its own solutions). A **stronger external reranker does**: sample K=32 with | |
| the shipped card, then have **Qwen3-4B** *reason through each distinct candidate answer and keep the one it | |
| verifies* (`ext_verify.py`, "reasoned" lens). Validated on the shipped weights (v4), n=500, McNemar vs the | |
| same-base vote (all p<0.001): | |
| | clean set | same-base self-consistency (vote) | **+ Qwen3-4B reranker** | pass@32 (ceiling) | | |
| |---|:---:|:---:|:---:| | |
| | SVAMP | 36.4 | **74.8** (+38.4) | 77.0 | | |
| | ASDiv | 49.0 | **76.0** (+27.0) | 79.2 | | |
| | MAWPS | 38.4 | **58.0** (+19.6) | 77.2 | | |
| The reranker recovers **~97% of the pass@32 ceiling** on SVAMP/ASDiv and lifts the deployable metric **~41 β | |
| ~70 on average (+28 pts)** β including the MAWPS self-consistency that the v4 weights had traded away. Note | |
| the judge must *reason*: a 1-token "Yes/No" verifier scores *below* the vote. This is a **2-model serving | |
| win** β the accuracy is the external reasoner's competence applied to this model's candidate set, not the | |
| single 2.88B card unlocked. Tooling: [`reasoning/ext_verify.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/ext_verify.py) (two-phase generate β rerank). | |
| Both point to the same lesson: on this 2.88B base, the remaining gains are **serving-system** moves (a tool | |
| executor, an external reranker) or a **stronger base** β not further single-checkpoint weight edits. A broad | |
| n=1000 held-out sweep (SVAMP/ASDiv/MAWPS/GSM-Plus) confirms this at the weight level: v3, v4, and every | |
| weight-soup between them sit on **one flat Pareto line** β identical pass@32 (~78%, the capability ceiling is | |
| fixed), trading only greedy β self-consistency. No single-checkpoint edit improves both axes; the reranker | |
| above is the way to cash the fixed ceiling. | |
| ## Limitations | |
| - **2.88B parameters** β far smaller than frontier models; expect weaker performance on hard reasoning, long-form knowledge, and code. | |
| - **Base-capability gaps persist:** some simple general questions (e.g. listing all three primary colors, multi-hop comparison puzzles) are answered incorrectly regardless of think mode β these are limits of the underlying pretraining, not of the reasoning recipe. | |
| - **v3 divisor-counting regression:** v3's short-trace diet erodes one fragile synthetic capability β counting the number of divisors of an integer via prime factorization (e.g. "how many divisors does 12 have?") is answered correctly by v2 but garbled by v3. This is a single narrow templated pattern; broad held-out math (SVAMP/ASDiv) improved. Noted for transparency. | |
| - **v4 self-consistency trade:** v4's external-teacher distillation + tool/coding diet improved greedy ASDiv (+4.7) and native termination but cut sample diversity, **regressing self-consistency ~6 pts** (mainly MAWPS) vs v3 β an honest mixed update, not a clean win. If your use relies on K-sample majority voting, v3 may serve you better; v3 (`x_v6v2_040`) is retained upstream for rollback. Tool-calling and coding data were added but are **not** in this shipped card (see *Tool use & serving-system capabilities*). | |
| - **Anti-repeat decoding breaks math** (see notes above) β a property of this small model's arithmetic being token-repetition-heavy. | |
| - **The selection gap is base-limited:** the correct answer is *present* in ~73% of 32 samples but only *picked* ~36β51% of the time; no cheap picker (majority vote, confidence-weighting, a same-base learned verifier) closes this gap β it caps at the base's ability to *judge* a solution. | |
| - Context extended via RoPE extrapolation; very-long-context retrieval may degrade. | |
| - English-centric SFT/DPO data; limited multilingual ability. No safety filtering or content moderation has been applied. | |
| ## Source code | |
| All training code is on the `main` branch of the Argonne repo: | |
| https://github.com/PursuitOfDataScience/ArgonneAI | |
| The custom architecture (`ArgonneModel`, `model_type = argonne2`) is bundled with this | |
| model as `model.py`. The exact scripts that produced this checkpoint, by pipeline stage: | |
| | Stage | Script(s) | | |
| |-------|-----------| | |
| | Architecture | [`model.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/model.py) | | |
| | Pretraining (base) | [`pretrain.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/pretrain.py) | | |
| | Intermix midtraining | [`preprocess_finemath.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/preprocess_finemath.py) β [`reasoning/build_intermix.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_intermix.py) β [`midtraining.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/midtraining.py) | | |
| | Soup base (training-free) | [`reasoning/build_soup_base.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_soup_base.py) | | |
| | SFT | [`sft.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/sft.py) | | |
| | DPO | [`dpo.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/dpo.py) | | |
| | CoT-SFT data (v1/v2) | [`reasoning/build_sft_mix.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_sft_mix.py), [`reasoning/build_mix_v3.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v3.py) | | |
| | CoT-SFT data (v3, short-only) | [`reasoning/build_mix_v6.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v6.py) | | |
| | CoT-SFT data (v4: teacher + tool + coding) | [`reasoning/gen_teacher.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/gen_teacher.py) (Qwen3-4B teacher), [`reasoning/build_mix_v7.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_mix_v7.py) | | |
| | CoT-SFT training | [`reasoning/cot-sft.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/cot-sft.py) | | |
| | Weight-soups (training-free) | [`reasoning/build_ckpt_soup.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/build_ckpt_soup.py) | | |
| | Deployment | [`reasoning/deploy_hf.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/deploy_hf.py) | | |
| | Serving-system (research) | [`reasoning/ext_verify.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/ext_verify.py) (external-verifier best-of-N), [`reasoning/tool_decode.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_decode.py) (tool-execution loop), [`reasoning/vllm_argonne.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/vllm_argonne.py) (vLLM port) | | |
| | Evaluation | [`reasoning/clean_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/clean_eval.py) (clean SVAMP/ASDiv/MAWPS + GSM-Plus, Wilson CIs β the honest judge), [`reasoning/code_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/code_eval.py) (HumanEval), [`reasoning/tool_eval.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/tool_eval.py) (tool-call format), [`reasoning/eval_numeracy.py`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/eval_numeracy.py) | | |
| | Full writeup / recipe | [`reasoning/thinking_training.md`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/thinking_training.md) | | |
| The SLURM launcher scripts that wire these together with per-stage hyperparameters are | |
| intentionally untracked in the repo (they carry cluster-specific paths); every stage's | |
| hyperparameters are recorded in [`reasoning/thinking_training.md`](https://github.com/PursuitOfDataScience/ArgonneAI/blob/main/reasoning/thinking_training.md) and summarized in the | |
| tables above. | |
| ## Citation | |
| ```bibtex | |
| @misc{argonne30think, | |
| author = {PursuitOfDataScience}, | |
| title = {Argonne 3.0-think}, | |
| year = {2026}, | |
| publisher = {Hugging Face}, | |
| url = {https://huggingface.co/PursuitOfDataScience/Argonne-3.0-think} | |
| } | |
| ``` | |