Width-Vs-Depth / README.md
Harley-ml's picture
Update README.md
06697ba verified
|
Raw
History Blame Contribute Delete
3.84 kB
---
license: apache-2.0
datasets:
- HuggingFaceFW/fineweb
language:
- en
tags:
- width-vs-depth
- slms
- custom_tokenizer
- llama
- llama-architecture
- tlms
- ablation
---
# Width Vs Depth
We tested two different architecture configurations on 500M tokens of FineWeb to investigate whether depth or width is more effective for tiny language models (TLMs).
## Model Architectures
### Config A: Deep-Narrow (code name: `depth_311`)
- Base Architecture: `LlamaForCausalLM`
- Tokenizer: [`Harley-ml/Dillionv2-1.3M`](https://huggingface.co/Harley-ml/Dillionv2-1.3M)
- Transformers Version: `5.13.1`
- Hidden Size: `128`
- Vocab Size: 2564
- Number of Layers: `21`
- Number of Heads: `4`
- Number of KV Heads: `2`
- Intermediate Size: `313`
- Head Dim: `32`
- Max Position Embeddings: `256`
- RoPe Theta: `2125.0`
- Tie Word Embeddings: `true`
- Hidden Activation: `silu`
- MLP Bias: `false`
- Initializer Range: `0.2`
- RMS Norm Eps: `1e-06`
- Pretraining Tp: `1`
- Use Cache: `false`
- **Hidden Size/Layers**: `6.10`
- **Total Parameters**: `3,889,920`
### Config B: Shallow-Wide (code name: `width_311`)
- Base Architecture: `LlamaForCausalLM`
- Tokenizer: [`Harley-ml/Dillionv2-1.3M`](https://huggingface.co/Harley-ml/Dillionv2-1.3M)
- Transformers Version: `5.13.1`
- Hidden Size: `192`
- Vocab Size: 2564
- Number of Layers: `9`
- Number of Heads: `6`
- Number of KV Heads: `2`
- Intermediate Size: `469`
- Head Dim: `32`
- Max Position Embeddings: `256`
- RoPe Theta: `2125.0`
- Tie Word Embeddings: `true`
- Hidden Activation: `silu`
- MLP Bias: `false`
- Initializer Range: `0.2`
- RMS Norm Eps: `1e-06`
- Pretraining Tp: `1`
- Use Cache: `false`
- **Hidden Size/Layers**: `21.33`
- **Total Parameters**: `3,816,576`
## Training Setup
- Epochs: `1`
- Max Steps: `-1.0`
- Batch Size: `256`
- Sequence Length: `256`
- Gradient Accumulation: `2`
- Gradient Clipping: `1.0`
- Gradient Checkpointing: `true`
- Learning Rate: `3e-3`
- Eval Split: `0.001`
- Weight Decay: `0.01`
- Optimizer: `AdamW`
- AdamW Betas: `(0.9, 0.95)`
- AdamW Eps: `1e-8`
- Scheduler: `WSD`
- WSD Warmup Ratio: `0.015`
- WSD Stable Ratio: `0.78`
- WSD Decay Ratio: `0.20`
- WSD Minium LR Ratio: `0.0`
- WSD Number of Cycles: `0.5`
- DType: `float16`
- Torch.Compile: `false`
- DataLoader Workers: `4`
- Seed: `311`
## Results
Accuracy is normalized by length and shown as a percentage.
| Config | Final Val Loss ↓ | Arc Easy ↑ | Arc Challenge ↑ | HellaSwag ↑ | PiQA ↑ | Swag ↑ | Blimp ↑ | Avg ↑ |
| ------ | ---------------: | ---------: | --------------: | ----------: | --------: | --------: | --------: | --------: |
| Config A | `3.13697` | **29.17** | **21.67** | **27.01** | **54.03** | **32.65** | 67.66 | **38.70** |
| Config B | `3.14935` | 28.91 | 20.73 | 26.93 | 53.65 | 32.24 | **68.18** | 38.44 |
Config A scores higher than Config B on nearly every task, demonstrating that even at minuscule scales, greater depth can outperform greater width.
## Model Checkpoints
The two models are stored separately in different folders in this repository. To load them, use:
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
config_a = AutoModelForCausalLM.from_pretrained(
"fromziro/Width-Vs-Depth",
subfolder="config_a",
)
# load config b instead:
# config_b = AutoModelForCausalLM.from_pretrained(
# "fromziro/Width-Vs-Depth",
# subfolder="config_b",
# )
tokenizer = AutoTokenizer.from_pretrained("fromziro/Width-Vs-Depth")
```
## License
Apache 2.0.
## Citation
```
@misc{width-vs-depth,
title = {Width-vs-Depth at Small Scales},
organization = {FromZero},
authors = {Paul Courneya, Jonathon LY},
year = {2026},
url = {https://huggingface.co/fromziro/Width-Vs-Depth]
}
```