Qielbas Tiny 9M M2
A 9,557,716-parameter English causal language model.
Architecture
Same architecture and tokenizer as M1: four GatedDeltaNet2 blocks and one causal attention block, width 320, SwiGLU hidden size 1048, RMSNorm, RoPE, tied embeddings, context 256 and a 4096-token BPE vocabulary.
Changes from M1
Both releases start from the same Muon pretrained checkpoint average. M2 adds a 300M-token continuation with 15% WikiText-103 raw training text and 85% of the original data mixture. WikiText-2 and WikiText-103 validation/test articles were excluded by title and exact text before training.
The continuation uses fresh Muon + AdamW, peak LR 0.0001, and seed 7. ARC adaptation uses fresh AdamW, peak LR 0.00005, five epochs and seed 123. Its loss is 0.15 choice CE + 0.85 text replay CE; M1 used 1.0 + 0.25. Replay keeps the 15% WikiText share. Epoch 5 was selected by ARC-Easy held-out validation accuracy; M1 used mean Easy/Challenge accuracy with a language-loss guard. Benchmark test data was not used for training or epoch selection.
Data
The original 6.7B tokens comprise FineWeb-Edu (2.7B), Cosmopedia v2 (0.3B), FinePhrase FAQ and Tutorial (0.5B each), DCLM-Edu (1.0B), FinePDFs-Edu English (1.2B), and FineWiki English (0.5B). The additional 300M tokens contain 45M WikiText exposures and 255M exposures from that mixture, bringing pretraining to 7.0B tokens. The audited WikiText pool contains 176M unique tokens.
ARC adaptation uses 3345 Easy/Challenge training questions and 867 held-out validation questions, with general-text replay.
Benchmarks
| Metric | M1 | M2 |
|---|---|---|
| Tiny-ML Efficiency | 80.1426 | 81.4170 |
| Raw Overall | 71.8696 | 73.0124 |
| ARC-Easy accuracy | 48.2744% | 47.4327% |
| BLiMP accuracy | 75.3224% | 75.6836% |
| WikiText-2 byte perplexity | 2.90791 | 2.33674 |
M2 gains 1.2744 Efficiency points (+1.59%).
Same frozen Glint-1.3 protocol as M1: 256-token contexts, no BOS, unnormalized answer likelihood, 2376 ARC-Easy questions and 67000 BLiMP pairs. Efficiency uses the retained Tiny-ML normalization. Results are self-reported; the separate official-harness M1 evaluation is not mixed into this table. Reports and predictions are included.
Usage
Native PyTorch loader; requires a BF16-capable CUDA GPU and the pinned dependencies.
pip install -r requirements.txt
import torch
from load_model import load_model
model, tokenizer = load_model(".")
inputs = torch.tensor([tokenizer.encode("The Earth revolves around").ids], device="cuda")
with torch.inference_mode(), torch.autocast("cuda", dtype=torch.bfloat16):
next_token = model(inputs)[0, -1].argmax().item()
print(tokenizer.decode([next_token]))
- Downloads last month
- 20