metadata
license: apache-2.0
datasets:
- HuggingFaceFW/fineweb
language:
- en
tags:
- width-vs-depth
- slms
- custom_tokenizer
- llama
- llama-architecture
- tlms
- ablation
Width Vs Depth
We tested two different architecture configurations on 500M tokens of FineWeb to investigate whether depth or width is more effective for tiny language models (TLMs).
Model Architectures
Config A: Deep-Narrow (code name: depth_311)
- Base Architecture:
LlamaForCausalLM - Tokenizer:
Harley-ml/Dillionv2-1.3M - Transformers Version:
5.13.1 - Hidden Size:
128 - Vocab Size: 2564
- Number of Layers:
21 - Number of Heads:
4 - Number of KV Heads:
2 - Intermediate Size:
313 - Head Dim:
32 - Max Position Embeddings:
256 - RoPe Theta:
2125.0 - Tie Word Embeddings:
true - Hidden Activation:
silu - MLP Bias:
false - Initializer Range:
0.2 - RMS Norm Eps:
1e-06 - Pretraining Tp:
1 - Use Cache:
false - Hidden Size/Layers:
6.10 - Total Parameters:
3,889,920
Config B: Shallow-Wide (code name: width_311)
- Base Architecture:
LlamaForCausalLM - Tokenizer:
Harley-ml/Dillionv2-1.3M - Transformers Version:
5.13.1 - Hidden Size:
192 - Vocab Size: 2564
- Number of Layers:
9 - Number of Heads:
6 - Number of KV Heads:
2 - Intermediate Size:
469 - Head Dim:
32 - Max Position Embeddings:
256 - RoPe Theta:
2125.0 - Tie Word Embeddings:
true - Hidden Activation:
silu - MLP Bias:
false - Initializer Range:
0.2 - RMS Norm Eps:
1e-06 - Pretraining Tp:
1 - Use Cache:
false - Hidden Size/Layers:
21.33 - Total Parameters:
3,816,576
Training Setup
- Epochs:
1 - Max Steps:
-1.0 - Batch Size:
256 - Sequence Length:
256 - Gradient Accumulation:
2 - Gradient Clipping:
1.0 - Gradient Checkpointing:
true - Learning Rate:
3e-3 - Eval Split:
0.001 - Weight Decay:
0.01 - Optimizer:
AdamW - AdamW Betas:
(0.9, 0.95) - AdamW Eps:
1e-8 - Scheduler:
WSD - WSD Warmup Ratio:
0.015 - WSD Stable Ratio:
0.78 - WSD Decay Ratio:
0.20 - WSD Minium LR Ratio:
0.0 - WSD Number of Cycles:
0.5 - DType:
float16 - Torch.Compile:
false - DataLoader Workers:
4 - Seed:
311
Results
Accuracy is normalized by length and shown as a percentage.
| Config | Final Val Loss ↓ | Arc Easy ↑ | Arc Challenge ↑ | HellaSwag ↑ | PiQA ↑ | Swag ↑ | Blimp ↑ | Avg ↑ |
|---|---|---|---|---|---|---|---|---|
| Config A | 3.13697 |
29.17 | 21.67 | 27.01 | 54.03 | 32.65 | 67.66 | 38.70 |
| Config B | 3.14935 |
28.91 | 20.73 | 26.93 | 53.65 | 32.24 | 68.18 | 38.44 |
Config A scores higher than Config B on nearly every task, demonstrating that even at minuscule scales, greater depth can outperform greater width.
Model Checkpoints
The two models are stored separately in different folders in this repository. To load them, use:
from transformers import AutoModelForCausalLM, AutoTokenizer
config_a = AutoModelForCausalLM.from_pretrained(
"fromziro/Width-Vs-Depth",
subfolder="config_a",
)
# load config b instead:
# config_b = AutoModelForCausalLM.from_pretrained(
# "fromziro/Width-Vs-Depth",
# subfolder="config_b",
# )
tokenizer = AutoTokenizer.from_pretrained("fromziro/Width-Vs-Depth")
License
Apache 2.0.
Citation
@misc{width-vs-depth,
title = {Width-vs-Depth at Small Scales},
organization = {FromZero},
authors = {Paul Courneya, Jonathon LY},
year = {2026},
url = {https://huggingface.co/fromziro/Width-Vs-Depth]
}