TinyBalls-v1-Base

The base (non-instruct) model: 113.3M parameters, trained from scratch on 2.01B tokens of English text. For the chat and tool-use version see igidn/TinyBalls-110M-V1.

Architecture

params 113.3M (151.0M counting the tied head twice)
layers 12
hidden size 768
attention 12 query heads / 4 KV heads (GQA), head dim 64
FFN SwiGLU, intermediate 2048, pre-activations clamped at 10.0
vocab 49,154 (SmolLM2 byte-level BPE, plus ChatML and tool specials reserved but untrained)
context 2,048 (the length it was trained at)
tied embeddings yes
RoPE theta 10,000

Weights are float32 and load as a standard transformers LlamaForCausalLM.

Usage

This is a completion model. There is no chat template, and the ChatML tokens in its vocabulary are untrained, do not prompt it with <|im_start|>.

from transformers import AutoModelForCausalLM, AutoTokenizer

tok = AutoTokenizer.from_pretrained("igidn/TinyBalls-v1-Base")
model = AutoModelForCausalLM.from_pretrained(
    "igidn/TinyBalls-v1-Base", torch_dtype="float32")

ids = tok("The capital of France is", return_tensors="pt").input_ids
out = model.generate(ids, max_new_tokens=48, do_sample=True,
                     temperature=0.8, top_p=0.95)
print(tok.decode(out[0][ids.shape[1]:]))

Documents were packed with <|endoftext|> (id 0) as the separator, which is also the EOS and pad token here. The generation_config.json defaults (temperature=0.8, top_p=0.95, sampling on) are what you want: greedy decoding collapses into repetition on this checkpoint, under greedy sampling the 4-gram repeat rate is 0.94 and distinct-token rate 0.04, versus 0.05–0.09 and 0.38–0.43 under top-p sampling.

Training

Two phases on a 2.01B-token English-only mix (1.5B broad, then a 0.5B quality anneal over the final 25% of steps), one LR schedule spanning both:

domain tokens
web (FineWeb-Edu) 670M
books (Gutenberg, pre-1929, DOAB, LoC) 150M
Wikipedia (FineWiki) 130M
code (UltraData-Code, codeparrot) 120M
math (FineMath, OpenWebMath, AutoMathText) 110M
synthetic prose (Cosmopedia) 80M
papers (peS2o, arXiv, Proof-pile-2) 80M
QA and forums (Stack Exchange, Ubuntu IRC) 80M
micro-domains (patents, legal, regulations, ...) 80M

Files

file
model.safetensors fp32 LlamaForCausalLM weights
config.json architecture, rope_theta 10000, 2048 context
tokenizer.json the exact tokenizer these weights were trained with
tokenizer_config.json specials; no chat template
generation_config.json sampling defaults, EOS = `<
Modelfile minimal Ollama serving file

License

MIT.

-igidn

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for igidn/TinyBalls-v1-Base

Finetunes
1 model
Quantizations
1 model

Datasets used to train igidn/TinyBalls-v1-Base