File size: 1,627 Bytes
a17482f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 | ---
license: apache-2.0
language: en
tags:
- tiny
- slm
- small-language-model
- from-scratch
- gqa
- swiglu
- rope
- training-script
pipeline_tag: text-generation
metrics:
- accuracy
- perplexity
---
# ram-18m Training Script
Training script for **ram-18m**: an 18,290,304-param LLaMA-style language model.
## Architecture
| Parameter | Value |
|-----------|-------|
| d_model | 384 |
| n_heads | 6 |
| n_kv_heads | 2 (GQA) |
| head_dim | 64 |
| n_layers | 7 |
| FFN | SwiGLU, 4x (1536) |
| Norm | RMSNorm |
| Positional | RoPE (θ=10000) |
| Vocab | 8192 (BPE) |
| Tied embed/head | yes |
| **Total params** | **18,290,304** |
## Default Training Config
- Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
- Optimizer: AdamW (β=0.9/0.95, wd=0)
- LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
- Batch: 32, seq_len 512
- Steps: 12,207 (~2B tokens)
- Grad clip: 1.0
## Usage
```bash
# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare
# 2. Train
python3 train_ram_18m.py --stage train
# 3. Or do both
python3 train_ram_18m.py --stage all
# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt
```
## Requirements
```
pip install torch transformers datasets numpy tokenizers
```
## Notes
- Requested by GGUFGuy in [model-requests #30](https://huggingface.co/spaces/Compactbot/model-requests/discussions/30)
- The model is NOT trained yet — this is the script only.
- GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
- Checkpoints saved every 500 steps to the script directory. |