Compactbot's picture
Add README with architecture details and usage (#2)
a17482f
|
Raw History Blame Contribute Delete
1.63 kB
metadata
license: apache-2.0
language: en
tags:
  - tiny
  - slm
  - small-language-model
  - from-scratch
  - gqa
  - swiglu
  - rope
  - training-script
pipeline_tag: text-generation
metrics:
  - accuracy
  - perplexity

ram-18m Training Script

Training script for ram-18m: an 18,290,304-param LLaMA-style language model.

Architecture

Parameter Value
d_model 384
n_heads 6
n_kv_heads 2 (GQA)
head_dim 64
n_layers 7
FFN SwiGLU, 4x (1536)
Norm RMSNorm
Positional RoPE (θ=10000)
Vocab 8192 (BPE)
Tied embed/head yes
Total params 18,290,304

Default Training Config

  • Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
  • Optimizer: AdamW (β=0.9/0.95, wd=0)
  • LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
  • Batch: 32, seq_len 512
  • Steps: 12,207 (~2B tokens)
  • Grad clip: 1.0

Usage

# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare

# 2. Train
python3 train_ram_18m.py --stage train

# 3. Or do both
python3 train_ram_18m.py --stage all

# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt

Requirements

pip install torch transformers datasets numpy tokenizers

Notes

  • Requested by GGUFGuy in model-requests #30
  • The model is NOT trained yet — this is the script only.
  • GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
  • Checkpoints saved every 500 steps to the script directory.