metadata
license: apache-2.0
language:
- en
tags:
- ablations
- datasets
- 1m
- llama
- custom_tokenizer
DataBench
Rather than training one model, we trained multiple models, each on a different dataset. Everything else was kept the same, so that the only changing variable would be the dataset itself.
Model Architecture
- Base Architecture:
LlamaForCasualLM - Tokenizer:
Harley-ml/Dillionv2-1.3M - Transformers Version:
5.13.1 - Hidden Size:
128 - Vocab Size: 2564
- Number of Layers:
8 - Number of Heads:
4 - Number of KV Heads:
2 - Intermediate Size:
344 - Head Dim:
32 - Max Position Embeddings:
256 - RoPe Theta:
2500.0 - Tie Word Embeddings:
true - Hidden Activation:
silu - MLP Bias:
false - Initializer Range:
0.2 - RMS Norm Eps:
1e-06 - Pretraining Tp:
1 - Use Cache:
false - Total Parameters:
1,780,352
Training Setup
- Epochs:
1 - Max Steps:
-1.0 - Batch Size:
400 - Sequence Length:
256 - Gradient Accumulation:
2 - Gradient Clipping:
1.0 - Gradient Checkpointing:
true - Learning Rate:
2.5e-3 - Eval Split:
0.00165 - Weight Decay:
0.01 - Optimizer:
AdamW - AdamW Betas:
(0.9, 0.95) - AdamW Eps:
1e-8 - Scheduler:
WSD - WSD Warmup Ratio:
0.015 - WSD Stable Ratio:
0.78 - WSD Decay Ratio:
0.20 - WSD Minium LR Ratio:
0.0 - WSD Number of Cycles:
0.5 - DType:
float16 - Torch.Compile:
true - DataLoader Workers:
2 - Seed:
311
Results
Accuracy is normalized by length and shown as a percentage.
| Dataset | ARC-Easy | HellaSwag | PIQA | Avg ↑ |
|---|---|---|---|---|
| FineWeb | 27.82% | 26.89% | 52.83% | 35.85% |
| DCLM-1.0-Baseline | 29.08% | 27.06% | 52.23% | 36.12% |
| DOAB | 31.14% | 28.00% | 52.18% | 37.11% |
| Project Gutenberg | 25.84% | 24.63% | 50.21% | 33.56% |
| EOT-2004-Raw | 29.08% | 27.54% | 51.31% | 35.98% |
| Wikipedia | 28.28% | 27.61% | 51.14% | 35.68% |
(Will add ArithMark-3.0 later)
Notice
This is a work in progress and is currently not completed. By the end of this project, we aim to have tested over 50 datasets.
License
Apache 2.0.