Compactbot's picture
Add README with architecture details and usage (#2)
a17482f
|
Raw History Blame Contribute Delete
1.63 kB
---
license: apache-2.0
language: en
tags:
- tiny
- slm
- small-language-model
- from-scratch
- gqa
- swiglu
- rope
- training-script
pipeline_tag: text-generation
metrics:
- accuracy
- perplexity
---
# ram-18m Training Script
Training script for **ram-18m**: an 18,290,304-param LLaMA-style language model.
## Architecture
| Parameter | Value |
|-----------|-------|
| d_model | 384 |
| n_heads | 6 |
| n_kv_heads | 2 (GQA) |
| head_dim | 64 |
| n_layers | 7 |
| FFN | SwiGLU, 4x (1536) |
| Norm | RMSNorm |
| Positional | RoPE (θ=10000) |
| Vocab | 8192 (BPE) |
| Tied embed/head | yes |
| **Total params** | **18,290,304** |
## Default Training Config
- Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
- Optimizer: AdamW (β=0.9/0.95, wd=0)
- LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
- Batch: 32, seq_len 512
- Steps: 12,207 (~2B tokens)
- Grad clip: 1.0
## Usage
```bash
# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare
# 2. Train
python3 train_ram_18m.py --stage train
# 3. Or do both
python3 train_ram_18m.py --stage all
# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt
```
## Requirements
```
pip install torch transformers datasets numpy tokenizers
```
## Notes
- Requested by GGUFGuy in [model-requests #30](https://huggingface.co/spaces/Compactbot/model-requests/discussions/30)
- The model is NOT trained yet — this is the script only.
- GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
- Checkpoints saved every 500 steps to the script directory.