nkthebass's picture
Card: document GGUF add_space_prefix=false + leading-space template fix (#23840)
bffc2a8 verified
|
Raw
History Blame Contribute Delete
11.3 kB
---
license: apache-2.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
base_model: nkthebass/tinybrainbot-320mV2-base
tags:
- tinybrainbot
- small-language-model
- from-scratch
- gqa
- instruct
- chat
---
# TinyBrainBot 320M V2 β€” Instruct
A **~326M-parameter** decoder-only language model, trained from scratch on **~10B tokens** and then **supervised-fine-tuned** for chat / instruction following. This is the **instruct** model.
- Base model: **`tinybrainbot-320mV2-base`** (full pretraining details there).
- Successor to the 303M V2 instruct.
**TL;DR:** A compact general-purpose assistant trained from scratch on ~10B tokens, with an added **conversational + in-context-recall** fine-tune (robust multi-turn chat, diverse instruction phrasings, remembers facts stated earlier in the conversation). It matches/beats **Pythia-410M** on general benchmarks (on far fewer training tokens) and sweeps **GPT-2-124M**, follows instructions, and has **2–3-digit arithmetic far stronger than its GSM8K reasoning score would suggest** (94–99% on 2–3-digit addition) β€” though it hits a clear length-generalization wall beyond 3 digits. It is **not** math-specialized.
---
## Model details
| | |
|---|---|
| Parameters | **325,899,264** (~326M) |
| Architecture | Decoder-only transformer, pre-norm, RMSNorm, SwiGLU MLP, RoPE |
| Hidden size | 1024 |
| Layers | 26 |
| Attention heads | 16 (query) / **4 KV heads** (grouped-query attention) |
| FFN size | 2816 |
| Context length | 1024 |
| Vocabulary | 32,000 |
| Tokenizer | `tbb-32k-v2` β€” 32k BPE (67% English / 20% code / 13% math), with reserved `<think>`/`</think>` special tokens |
| Precision | trained in fp16 with an fp32 master copy (autocast) |
## Usage
Prompt with the chat format:
```
<|user|>
{user message}
<|end|>
<|assistant|>
{assistant reply}
<|end|>
```
Example (greedy):
```
<|user|>
What is the capital of France?
<|end|>
<|assistant|>
Paris.
<|end|>
```
The model gives concise direct answers and shows worked steps for arithmetic.
### Conversational + recall update
This release adds a **conversational + in-context-recall** fine-tuning pass on top of the base instruct SFT. It:
- answers open-ended imperatives robustly β€” `list all the planets in the solar system` β†’ *Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune.*;
- handles **multi-turn chat and in-context recall** (state a fact early, ask about it later);
- stays consistent across the F16 GGUF (LM Studio / Ollama / llama.cpp) and the fp16 `transformers` weights.
*(An earlier build gave terse/empty answers to some imperative phrasings through the GGUF export. That was a real tokenization mismatch in the export β€” not a capacity limit. This release fixes it two ways: the conversational pass makes the model robust to it, **and** the F16 GGUF now sets `tokenizer.ggml.add_space_prefix=false` + ships a leading-space chat template so llama.cpp tokenizes the chat format token-for-token identically to the native tokenizer, per llama.cpp [#23840](https://github.com/ggml-org/llama.cpp/issues/23840).)*
---
## Training procedure
![training loss curve](loss-curve-instruct.png)
*Pretraining (blue, ~10B tokens) β†’ instruct SFT (orange, from step 51k) β†’ conversational + recall SFT (green, steps 52k–55k, final loss ~0.5).*
### Pretraining (base)
Pretrained from scratch on **10.03B tokens** (51,000 steps, WSD schedule, peak LR 6e-4, two-phase broad β†’ quality-anneal, final loss **1.436**) across **2Γ— Tesla V100-PCIE-16GB** with PyTorch DDP (gloo), fp16 autocast + fused AdamW. Full pretraining details and the 13-source data mix are documented on the **base** model card.
### Supervised fine-tuning (this model)
| Hyperparameter | Value |
|---|---|
| Steps | ~1,433 (resumed from base @ 51,000) |
| Tokens | ~94M |
| LR | 1.5e-5, constant |
| Warmup | 100 steps |
| Global batch | 4 micro Γ— 8 grad-accum Γ— 2 GPUs Γ— 1024 seq |
| Loss masking | assistant-only |
| Best checkpoint | step 52,000 (selected on eval, not last) |
**SFT data mix** (sampling weights; sum = 9.5; math β‰ˆ 15.8% of the mixture):
| Source | Weight |
|---|---:|
| longdef-sft (~16k long / multi-step answers) | 2.0 |
| smoltalk | 2.0 |
| **math-v2** (~70k verified worked examples) | 1.5 |
| qa-distill | 1.0 |
| lamini-instructions | 1.0 |
| soda-dialogues | 1.0 |
| greetings | 0.5 |
| reasoning-distill | 0.5 |
*`math-v2` is ~70k programmatically-generated, tolerance-verified arithmetic worked examples (multiplication / division / decimals, with the final answer checked). It appears in both pretraining and SFT β€” see the note below.*
> **On where the arithmetic comes from:** the base and instruct models score almost identically on the GPT-3 Arithmetic suite (**aggregate 31.4% vs 31.2%**). That base/instruct parity suggests **most arithmetic computation was acquired during pretraining**, while SFT mainly shaped instruction-following and response format (the one sub-task where SFT clearly helps is composite / order-of-operations, 7.0% vs 3.7%). Pretraining also delivered far more math *tokens* in absolute terms β€” on the order of ~215M vs ~15M β€” despite SFT's higher math *proportion*.
### Conversational + recall fine-tuning (final stage β€” this release)
A further SFT pass continued from the instruct checkpoint (**step 52,000 β†’ 55,000**) to make the model a robust multi-turn conversationalist, teach in-context recall, and fix brittle behavior on open-ended instruction phrasings.
| Hyperparameter | Value |
|---|---|
| Steps | 3,000 (resumed from instruct @ 52,000) |
| LR | 1.5e-5, constant (WSD, decay-fraction 0) |
| Global batch | 2 micro Γ— 16 grad-accum Γ— 2 GPUs Γ— 1024 seq |
| Loss masking | assistant-only |
| Final loss | ~0.5 |
**Data mix** (sampling weights):
| Source | Weight |
|---|---:|
| **convo-recall** (6k multi-turn in-context-recall dialogues) | 2.5 |
| **convo-core** (4k instruction-phrasing + small-talk) | 2.5 |
| smoltalk | 1.5 |
| math-v3 | 0.75 |
| longdef-sft | 0.75 |
| greetings | 0.5 |
| qa-distill | 0.5 |
`convo-recall` and `convo-core` are programmatically generated: `convo-recall` teaches **in-context memory** (the user states facts early β€” name / city / pet / counts β€” then asks about them later), and `convo-core` covers **diverse instruction verbs** ("list all / name / give me / what are the …") over closed sets with correct complete answers, plus natural small-talk. Alongside the conversational gains, this pass makes the model **robust to the SPM normalization that the GGUF/HF exports drop** β€” so it now behaves consistently in llama.cpp / LM Studio / Ollama *and* in `transformers`, instead of degrading on some phrasings through the export.
---
## Evaluation
Measured on our own log-likelihood MC harness (lm-eval style, fixed seed). Headline metric = `acc_norm` for HellaSwag/ARC/OpenBookQA, `acc` for WinoGrande/MMLU. Reference values are published lm-eval approximations β€” treat gaps under ~Β±2 points as ties.
### vs the previous 303M instruct (full test sets)
| Benchmark | n | **320M V2 instruct** | 303M instruct |
|---|:--:|:--:|:--:|
| HellaSwag | 10042 | **34.5** | 30.7 |
| ARC-Easy (acc_norm) | 2376 | **49.3** | 47.6 |
| ARC-Easy (raw acc) | 2376 | **57.0** | 51.0 |
| ARC-Challenge | 1172 | 27.6 | 27.6 |
| OpenBookQA | 500 | **31.8** | 29.0 |
| WinoGrande | 1267 | **53.3** | 52.2 |
| MMLU | 14042 | **28.0** | 27.1 |
β†’ **5 wins, 1 tie, 0 losses** over the previous generation.
### vs reference models (headline metric)
| Benchmark | **320M V2 instruct** | GPT-2-124M | Pythia-410M | SmolLM-360M |
|---|:--:|:--:|:--:|:--:|
| HellaSwag | 34.5 | 31 | 34 | 54 |
| ARC-Easy | 49.3 (57 raw) | 44 | 52 | 70 |
| ARC-Challenge | 27.6 | 22 | 24 | 37 |
| OpenBookQA | 31.8 | 29 | 30 | 42 |
| WinoGrande | 53.3 | 52 | 53 | 57 |
| MMLU | 28.0 | 26 | 25 | 34 |
β†’ Sweeps **GPT-2-124M**; **~5 wins + 1 draw vs Pythia-410M**. SmolLM-360M (trained on ~600B aggressively-filtered tokens) remains the frontier for this size.
**Training efficiency.** These results come from **~10B pretraining tokens** β€” roughly an order of magnitude fewer than the Pythia suite's ~300B. The Pythia-410M parity is therefore best read as a **token-efficiency** result (curated data + quality anneal) rather than a scale win.
### Math β€” computation vs reasoning
The model was trained on arithmetic **computation**, not word-problem **reasoning** β€” the two benchmarks below show that split clearly.
**GSM8K** (grade-school word problems, full 1319-problem test, zero-shot chain-of-thought):
| Model | GSM8K |
|---|:--:|
| GPT-2-124M | ~0% |
| **320M V2 instruct** | **0.53%** |
| Pythia-410M | ~1–2% |
| SmolLM2-360M-Instruct | ~3–5% |
β†’ At the floor **for a general-purpose model of this size and training mix**; stronger sub-1B *math-specialized* models can score substantially higher. GSM8K rewards multi-step semantic reasoning, which this recipe did not target.
**GPT-3 Arithmetic** (Brown et al. 2020 protocol, exact-match, n=300/sub-task):
| Sub-task | Accuracy |
|---|:--:|
| 2-digit addition | 99.0% |
| 3-digit addition | 94.0% |
| 2-digit subtraction | 49.3% |
| 3-digit subtraction | 42.7% |
| 4-digit addition / subtraction | 0.3% / 0.3% |
| 5-digit addition / subtraction | 0.0% / 0.0% |
| 2-digit multiplication | 21.7% |
| single-digit composite (order of ops) | 7.0% |
| **Aggregate (all 10 sub-tasks)** | **31.4%** |
β†’ **Strong through 3 digits, then a hard wall.** 2–3-digit addition is near-solved (94–99%, with correct carrying), but 4+-digit accuracy collapses to ~0%: the model executes a fixed **~3-column** addition routine and silently drops the higher place values β€” a **length-generalization limit** tied to the training distribution (`math-v2` operands are ≀3 digits), *not* truncation (generations complete normally and end with a stated answer). Subtraction sits ~42–50% β€” it handles `aβˆ’b` when `a>b` but drops the sign on **negative results**. 2-digit multiplication ~22%; single-digit composite (order of operations) ~7%.
> **How to read this:** the GPT-3 Arithmetic suite mainly probes **exact symbolic computation** and short-range algorithmic generalization; it should **not** be interpreted as evidence of strong mathematical *reasoning* (see GSM8K above). The two results together are the point: strong at computing, weak at reasoning.
---
## Intended use & limitations
**Intended use:** a capable general chat assistant at ~326M scale, on-device / low-resource deployment, research on small-model SFT, and arithmetic computation.
**Limitations:**
- **Math reasoning** (word problems, GSM8K/MATH) is at the floor β€” the model computes but does not reason through multi-step problems.
- **Negative-result subtraction** is unreliable (drops the sign).
- **WinoGrande and MMLU** sit near the random floor β€” consistent with the capacity and data limits of a ~326M model under this training recipe.
- Trained predominantly on English; 1024-token context; **no RLHF/safety tuning** β€” outputs may be incorrect or inappropriate and should not be relied upon unchecked.
## Hardware & framework
2Γ— NVIDIA Tesla V100-PCIE-16GB Β· Windows Β· PyTorch DDP (gloo) Β· fp16 autocast (fp32 master) Β· fused AdamW Β· custom TinyBrainBot trainer.