Text Generation
Transformers
Safetensors
English
llama
gpt-u
tiny-lm
pretrained-from-scratch
text-generation-inference

image

GPT-U-20M

A 20.5M-parameter Llama-architecture language model pretrained from scratch on 2.6B tokens of web text, educational text and code.

Architecture

Type decoder-only transformer, LlamaForCausalLM: RoPE, SwiGLU, RMSNorm, pre-norm, no biases
Parameters 20,453,760 (6,291,456 tied embedding + 8 × 1,770,240 per block + 384 final norm)
Layers 8
Hidden size 384
Attention 6 heads × 64 dims, full multi-head (no GQA), SDPA
MLP SwiGLU, intermediate size 1,024
Vocabulary 16,384 (byte-level BPE, trained on the same mix)
Context length 1,024 tokens
RoPE θ 10,000
Input/output embeddings tied

The 16,384-token vocabulary keeps the embedding at 31% of the parameters (a 50k vocabulary would need 19.3M parameters at width 384) and lets tokens be stored as uint16.

Tokenizer. Byte-level BPE with a cl100k-style split that keeps whitespace runs whole (Python indentation), one token per digit, and <|endoftext|> (id 0) as BOS/EOS/padding. Round-trip decode(encode(x)) == x holds on 1,000 held-out samples (tabs, CRLF, emoji, CJK, RTL, control characters).

Held-out domain Tokens / word Bytes / token
DCLM 1.48 4.07
FineWeb-Edu 1.45 4.24
Code (27% Python) 5.34 2.01
Python only 3.78 3.21

Training data

One epoch over 2.6B tokens, no repetition:

Dataset Target share Realized share Tokens seen (≈)
mlfoundations/dclm-baseline-1.0 45% 45.04% 1.17B
HuggingFaceFW/fineweb-edu (sample-10BT) 35% 35.03% 0.91B
HuggingFaceCode/stack-v3-train 20% 19.93% 0.52B
  • Each source is streamed from the Hub and tokenized up to its token budget; the documents are then interleaved per document in a seeded random order, i.e. with fixed proportions for the whole run (no domain curriculum), so the realized shares match the targets exactly. (Holding the four streams open at once was avoided: it needed 7–8 GB of RAM and destabilized the training machine.)
  • The Stack v3 is grouped by repository; repositories were split into individual files, vendored files dropped, and Python held at 26.6% of the code tokens (it is ~4% of Stack v3 naturally).
  • Documents are separated by <|endoftext|> and packed into contiguous 1,025-token blocks (1,024 inputs + 1 shifted target), no padding. Batches are sampled from a seeded permutation of all blocks.
  • Validation holdouts (2M tokens per domain) come from source files that never fed training or the tokenizer.

Training

Hyperparameter Value
Tokens per step 131,072 (2 GPUs × 32 sequences × 1,024 tokens × 2 accumulation steps)
Steps 19,840
Optimizer AdamW, β = (0.9, 0.95), ε = 1e-8
Peak learning rate 1.8e-3
Schedule WSD: 500 warmup steps → constant to step 15,872 → linear decay to 1.8e-4 at step 19,840
Weight decay 0.1 on 2D matrices only (none on norms and embeddings)
Gradient clipping 1.0
z-loss 1e-4
Dropout 0
Initialization N(0, 0.02); o_proj and down_proj scaled by 1/√(2·8)
Precision bf16 autocast, fp32 master weights

Trained on 2× RTX 3060 12GB (no NVLink) under WSL2 with PyTorch DDP and torch.compile (mode: compile): 4.6 h of active training time, 167,893 tokens/s on average (MFU ≈ 52% of the GPUs' measured bf16 matmul throughput), max GPU temperature 76°C, 3 automatic resume(s) from checkpoints after the machine was shut down, 0 loss-spike rollback(s).

Results

Validation loss on the 2M-token holdouts (1,600 fixed 1,024-token windows per domain) and bits per byte, bpb = loss_nats × tokens / (bytes × ln 2), which is comparable across tokenizers:

Domain Val loss (nats/token) Bits per byte
DCLM-baseline 3.5044 1.2377
FineWeb-Edu 3.2586 1.1070
The Stack v3 (code) 1.8075 1.3031
Aggregate 2.8568 1.1967

The aggregate is the equal-weight mean of the three holdouts (same number of tokens each). Code tokens are far easier to predict than prose with this tokenizer, which pulls the aggregate down; the prose domains (3.26–3.50 nats/token) are the fairer comparison point for web-text models, and bits per byte is the fairer one across tokenizers.

Zero-shot benchmarks (lm-evaluation-harness). <|endoftext|> is prepended as BOS. hellaswag and winogrande sit at chance, as expected at 20M parameters; arc_easy and piqa land clearly above it, which reflects surface lexical knowledge more than reasoning. They are reported for transparency, not as a quality signal.

Task Metric GPT-U-20M Chance
arc_easy acc 40.03 25
arc_easy acc_norm 37.42 25
piqa acc 58.60 50
piqa acc_norm 58.54 50
hellaswag acc 26.64 25
hellaswag acc_norm 27.85 25
lambada_openai acc 21.79 0
lambada_openai perplexity 138.45 —
winogrande acc 51.78 50

Sample generations (temperature 0.8, top-p 0.95) are in out/samples.md.

Loss curve

Training and validation loss

Usage

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("DedeProGames/GPT-U-20M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/GPT-U-20M")

inputs = tokenizer("The history of the Roman Empire", return_tensors="pt")  # prepends <|endoftext|> as BOS
output = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.8, top_p=0.95)
print(tokenizer.decode(output[0], skip_special_tokens=True))

Checkpoints

  • main: final weights (step 19,840, end of the decay phase).
  • stable-15872: weights at step 15,872, the end of the constant-learning-rate phase, meant for continued pretraining (e.g. extending to 4–5B tokens before decaying again). Load with revision="stable-15872".

Limitations

  • It is tiny. 20M parameters is a research and teaching scale; the model has shallow world knowledge and weak reasoning.
  • Heavily over-trained for its size by design: 127 tokens per parameter, about 6× the Chinchilla-optimal ratio (~20).
  • Not for production use. It hallucinates freely and produces fluent but often false statements.
  • Generated code is not functional. It imitates the surface form of code; do not run it.
  • English only (web and educational text); no instruction tuning, no safety tuning. Web data can carry biases and offensive content.
  • Context is limited to 1,024 tokens.

Data provenance and licenses

Source License Notes
DCLM-baseline 1.0 CC-BY-4.0 Common Crawl web text filtered by DataComp-LM; see the dataset card.
FineWeb-Edu (sample-10BT) ODC-By 1.0 Common Crawl pages filtered for educational value; use is also subject to Common Crawl's terms of use.
The Stack v3 (train) ODC-By GitHub code as of August 2025; files are permissively licensed or carry no detected license (non-permissive files excluded upstream), PII-redacted, with an opt-out process. Individual files keep their original licenses.

The model weights are released under Apache-2.0. Tokenized data is not redistributed.

Reproducibility

Everything used to build this model is in the repository: scripts/ (tokenizer training, data preparation, training, evaluation, publishing), configs/model.json, logs/train.jsonl (every 10 steps: loss, lr, grad norm, throughput, MFU, GPU temperatures, memory), logs/tokenizer_report.json and out/data_index.json (token counts per shard and domain, holdout bytes). Raw evaluation outputs: out/eval.json and out/lm_eval_results.json.

Downloads last month
607
Safetensors
Model size
20.5M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train DedeProGames/GPT-U-20M

Space using DedeProGames/GPT-U-20M 1

Collection including DedeProGames/GPT-U-20M