Instructions to use DedeProGames/GPT-U-20M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DedeProGames/GPT-U-20M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="DedeProGames/GPT-U-20M")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("DedeProGames/GPT-U-20M") model = AutoModelForCausalLM.from_pretrained("DedeProGames/GPT-U-20M", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use DedeProGames/GPT-U-20M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DedeProGames/GPT-U-20M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/GPT-U-20M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/DedeProGames/GPT-U-20M
- SGLang
How to use DedeProGames/GPT-U-20M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "DedeProGames/GPT-U-20M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/GPT-U-20M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "DedeProGames/GPT-U-20M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/GPT-U-20M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use DedeProGames/GPT-U-20M with Docker Model Runner:
docker model run hf.co/DedeProGames/GPT-U-20M
GPT-U-20M
A 20.5M-parameter Llama-architecture language model pretrained from scratch on 2.6B tokens of web text, educational text and code.
Architecture
| Type | decoder-only transformer, LlamaForCausalLM: RoPE, SwiGLU, RMSNorm, pre-norm, no biases |
| Parameters | 20,453,760 (6,291,456 tied embedding + 8 × 1,770,240 per block + 384 final norm) |
| Layers | 8 |
| Hidden size | 384 |
| Attention | 6 heads × 64 dims, full multi-head (no GQA), SDPA |
| MLP | SwiGLU, intermediate size 1,024 |
| Vocabulary | 16,384 (byte-level BPE, trained on the same mix) |
| Context length | 1,024 tokens |
| RoPE θ | 10,000 |
| Input/output embeddings | tied |
The 16,384-token vocabulary keeps the embedding at 31% of the parameters (a 50k vocabulary would need 19.3M parameters at width 384) and lets tokens be stored as uint16.
Tokenizer. Byte-level BPE with a cl100k-style split that keeps whitespace runs whole (Python indentation), one token per digit, and <|endoftext|> (id 0) as BOS/EOS/padding. Round-trip decode(encode(x)) == x holds on 1,000 held-out samples (tabs, CRLF, emoji, CJK, RTL, control characters).
| Held-out domain | Tokens / word | Bytes / token |
|---|---|---|
| DCLM | 1.48 | 4.07 |
| FineWeb-Edu | 1.45 | 4.24 |
| Code (27% Python) | 5.34 | 2.01 |
| Python only | 3.78 | 3.21 |
Training data
One epoch over 2.6B tokens, no repetition:
| Dataset | Target share | Realized share | Tokens seen (≈) |
|---|---|---|---|
| mlfoundations/dclm-baseline-1.0 | 45% | 45.04% | 1.17B |
HuggingFaceFW/fineweb-edu (sample-10BT) |
35% | 35.03% | 0.91B |
| HuggingFaceCode/stack-v3-train | 20% | 19.93% | 0.52B |
- Each source is streamed from the Hub and tokenized up to its token budget; the documents are then interleaved per document in a seeded random order, i.e. with fixed proportions for the whole run (no domain curriculum), so the realized shares match the targets exactly. (Holding the four streams open at once was avoided: it needed 7–8 GB of RAM and destabilized the training machine.)
- The Stack v3 is grouped by repository; repositories were split into individual files, vendored files dropped, and Python held at 26.6% of the code tokens (it is ~4% of Stack v3 naturally).
- Documents are separated by
<|endoftext|>and packed into contiguous 1,025-token blocks (1,024 inputs + 1 shifted target), no padding. Batches are sampled from a seeded permutation of all blocks. - Validation holdouts (2M tokens per domain) come from source files that never fed training or the tokenizer.
Training
| Hyperparameter | Value |
|---|---|
| Tokens per step | 131,072 (2 GPUs × 32 sequences × 1,024 tokens × 2 accumulation steps) |
| Steps | 19,840 |
| Optimizer | AdamW, β = (0.9, 0.95), ε = 1e-8 |
| Peak learning rate | 1.8e-3 |
| Schedule | WSD: 500 warmup steps → constant to step 15,872 → linear decay to 1.8e-4 at step 19,840 |
| Weight decay | 0.1 on 2D matrices only (none on norms and embeddings) |
| Gradient clipping | 1.0 |
| z-loss | 1e-4 |
| Dropout | 0 |
| Initialization | N(0, 0.02); o_proj and down_proj scaled by 1/√(2·8) |
| Precision | bf16 autocast, fp32 master weights |
Trained on 2× RTX 3060 12GB (no NVLink) under WSL2 with PyTorch DDP and torch.compile (mode: compile): 4.6 h of active training time, 167,893 tokens/s on average (MFU ≈ 52% of the GPUs' measured bf16 matmul throughput), max GPU temperature 76°C, 3 automatic resume(s) from checkpoints after the machine was shut down, 0 loss-spike rollback(s).
Results
Validation loss on the 2M-token holdouts (1,600 fixed 1,024-token windows per domain) and bits per byte, bpb = loss_nats × tokens / (bytes × ln 2), which is comparable across tokenizers:
| Domain | Val loss (nats/token) | Bits per byte |
|---|---|---|
| DCLM-baseline | 3.5044 | 1.2377 |
| FineWeb-Edu | 3.2586 | 1.1070 |
| The Stack v3 (code) | 1.8075 | 1.3031 |
| Aggregate | 2.8568 | 1.1967 |
The aggregate is the equal-weight mean of the three holdouts (same number of tokens each). Code tokens are far easier to predict than prose with this tokenizer, which pulls the aggregate down; the prose domains (3.26–3.50 nats/token) are the fairer comparison point for web-text models, and bits per byte is the fairer one across tokenizers.
Zero-shot benchmarks (lm-evaluation-harness). <|endoftext|> is prepended as BOS. hellaswag and winogrande sit at chance, as expected at 20M parameters; arc_easy and piqa land clearly above it, which reflects surface lexical knowledge more than reasoning. They are reported for transparency, not as a quality signal.
| Task | Metric | GPT-U-20M | Chance |
|---|---|---|---|
| arc_easy | acc | 40.03 | 25 |
| arc_easy | acc_norm | 37.42 | 25 |
| piqa | acc | 58.60 | 50 |
| piqa | acc_norm | 58.54 | 50 |
| hellaswag | acc | 26.64 | 25 |
| hellaswag | acc_norm | 27.85 | 25 |
| lambada_openai | acc | 21.79 | 0 |
| lambada_openai | perplexity | 138.45 | — |
| winogrande | acc | 51.78 | 50 |
Sample generations (temperature 0.8, top-p 0.95) are in out/samples.md.
Loss curve
Usage
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("DedeProGames/GPT-U-20M")
model = AutoModelForCausalLM.from_pretrained("DedeProGames/GPT-U-20M")
inputs = tokenizer("The history of the Roman Empire", return_tensors="pt") # prepends <|endoftext|> as BOS
output = model.generate(**inputs, max_new_tokens=100, do_sample=True, temperature=0.8, top_p=0.95)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Checkpoints
main: final weights (step 19,840, end of the decay phase).stable-15872: weights at step 15,872, the end of the constant-learning-rate phase, meant for continued pretraining (e.g. extending to 4–5B tokens before decaying again). Load withrevision="stable-15872".
Limitations
- It is tiny. 20M parameters is a research and teaching scale; the model has shallow world knowledge and weak reasoning.
- Heavily over-trained for its size by design: 127 tokens per parameter, about 6× the Chinchilla-optimal ratio (~20).
- Not for production use. It hallucinates freely and produces fluent but often false statements.
- Generated code is not functional. It imitates the surface form of code; do not run it.
- English only (web and educational text); no instruction tuning, no safety tuning. Web data can carry biases and offensive content.
- Context is limited to 1,024 tokens.
Data provenance and licenses
| Source | License | Notes |
|---|---|---|
| DCLM-baseline 1.0 | CC-BY-4.0 | Common Crawl web text filtered by DataComp-LM; see the dataset card. |
FineWeb-Edu (sample-10BT) |
ODC-By 1.0 | Common Crawl pages filtered for educational value; use is also subject to Common Crawl's terms of use. |
| The Stack v3 (train) | ODC-By | GitHub code as of August 2025; files are permissively licensed or carry no detected license (non-permissive files excluded upstream), PII-redacted, with an opt-out process. Individual files keep their original licenses. |
The model weights are released under Apache-2.0. Tokenized data is not redistributed.
Reproducibility
Everything used to build this model is in the repository: scripts/ (tokenizer training, data preparation, training, evaluation, publishing), configs/model.json, logs/train.jsonl (every 10 steps: loss, lr, grad norm, throughput, MFU, GPU temperatures, memory), logs/tokenizer_report.json and out/data_index.json (token counts per shard and domain, holdout bytes). Raw evaluation outputs: out/eval.json and out/lm_eval_results.json.
- Downloads last month
- 607

