--- license: apache-2.0 language: en tags: - tiny - slm - small-language-model - from-scratch - gqa - swiglu - rope - training-script pipeline_tag: text-generation metrics: - accuracy - perplexity --- # ram-18m Training Script Training script for **ram-18m**: an 18,290,304-param LLaMA-style language model. ## Architecture | Parameter | Value | |-----------|-------| | d_model | 384 | | n_heads | 6 | | n_kv_heads | 2 (GQA) | | head_dim | 64 | | n_layers | 7 | | FFN | SwiGLU, 4x (1536) | | Norm | RMSNorm | | Positional | RoPE (θ=10000) | | Vocab | 8192 (BPE) | | Tied embed/head | yes | | **Total params** | **18,290,304** | ## Default Training Config - Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens - Optimizer: AdamW (β=0.9/0.95, wd=0) - LR: 2e-4, cosine decay to 2e-5, warmup 200 steps - Batch: 32, seq_len 512 - Steps: 12,207 (~2B tokens) - Grad clip: 1.0 ## Usage ```bash # 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens) python3 train_ram_18m.py --stage prepare # 2. Train python3 train_ram_18m.py --stage train # 3. Or do both python3 train_ram_18m.py --stage all # 4. Eval a checkpoint python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt ``` ## Requirements ``` pip install torch transformers datasets numpy tokenizers ``` ## Notes - Requested by GGUFGuy in [model-requests #30](https://huggingface.co/spaces/Compactbot/model-requests/discussions/30) - The model is NOT trained yet — this is the script only. - GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower. - Checkpoints saved every 500 steps to the script directory.