--- license: apache-2.0 language: - en tags: - ablations - datasets - 1m - llama - custom_tokenizer --- # DataBench Rather than training one model, we trained multiple models, each on a different dataset. Everything else was kept the same, so that the only changing variable would be the dataset itself. ## Model Architecture - Base Architecture: `LlamaForCasualLM` - Tokenizer: [`Harley-ml/Dillionv2-1.3M`](https://huggingface.co/Harley-ml/Dillionv2-1.3M) - Transformers Version: `5.13.1` - Hidden Size: `128` - Vocab Size: 2564 - Number of Layers: `8` - Number of Heads: `4` - Number of KV Heads: `2` - Intermediate Size: `344` - Head Dim: `32` - Max Position Embeddings: `256` - RoPe Theta: `2500.0` - Tie Word Embeddings: `true` - Hidden Activation: `silu` - MLP Bias: `false` - Initializer Range: `0.2` - RMS Norm Eps: `1e-06` - Pretraining Tp: `1` - Use Cache: `false` - **Total Parameters**: `1,780,352` ## Training Setup - Epochs: `1` - Max Steps: `-1.0` - Batch Size: `400` - Sequence Length: `256` - Gradient Accumulation: `2` - Gradient Clipping: `1.0` - Gradient Checkpointing: `true` - Learning Rate: `2.5e-3` - Eval Split: `0.00165` - Weight Decay: `0.01` - Optimizer: `AdamW` - AdamW Betas: `(0.9, 0.95)` - AdamW Eps: `1e-8` - Scheduler: `WSD` - WSD Warmup Ratio: `0.015` - WSD Stable Ratio: `0.78` - WSD Decay Ratio: `0.20` - WSD Minium LR Ratio: `0.0` - WSD Number of Cycles: `0.5` - DType: `float16` - Torch.Compile: `true` - DataLoader Workers: `2` - Seed: `311` # Results Accuracy is normalized by length and shown as a percentage. | Dataset | ARC-Easy | HellaSwag | PIQA | **Avg ↑** | | :---------------- | ---------: | ---------: | ---------: | ---------: | | FineWeb | 27.82% | 26.89% | 52.83% | **35.85%** | | DCLM-1.0-Baseline | 29.08% | 27.06% | 52.23% | **36.12%** | | DOAB | 31.14% | 28.00% | 52.18% | **37.11%** | | Project Gutenberg | 25.84% | 24.63% | 50.21% | **33.56%** | | EOT-2004-Raw | 29.08% | 27.54% | 51.31% | **35.98%** | | Wikipedia | 28.28% | 27.61% | 51.14% | **35.68%** | (Will add ArithMark-3.0 later) ## Notice This is a work in progress and is currently not completed. By the end of this project, we aim to have tested over 50 datasets. ## License Apache 2.0.