File size: 1,627 Bytes
a17482f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
---
license: apache-2.0
language: en
tags:
  - tiny
  - slm
  - small-language-model
  - from-scratch
  - gqa
  - swiglu
  - rope
  - training-script
pipeline_tag: text-generation
metrics:
  - accuracy
  - perplexity
---

# ram-18m Training Script

Training script for **ram-18m**: an 18,290,304-param LLaMA-style language model.

## Architecture

| Parameter | Value |
|-----------|-------|
| d_model | 384 |
| n_heads | 6 |
| n_kv_heads | 2 (GQA) |
| head_dim | 64 |
| n_layers | 7 |
| FFN | SwiGLU, 4x (1536) |
| Norm | RMSNorm |
| Positional | RoPE (θ=10000) |
| Vocab | 8192 (BPE) |
| Tied embed/head | yes |
| **Total params** | **18,290,304** |

## Default Training Config

- Data: FineWeb-Edu L3 (sample-100BT), ~2B tokens
- Optimizer: AdamW (β=0.9/0.95, wd=0)
- LR: 2e-4, cosine decay to 2e-5, warmup 200 steps
- Batch: 32, seq_len 512
- Steps: 12,207 (~2B tokens)
- Grad clip: 1.0

## Usage

```bash
# 1. Prepare data (trains BPE tokenizer + tokenizes 2B tokens)
python3 train_ram_18m.py --stage prepare

# 2. Train
python3 train_ram_18m.py --stage train

# 3. Or do both
python3 train_ram_18m.py --stage all

# 4. Eval a checkpoint
python3 train_ram_18m.py --stage eval --ckpt ckpt_step12207.pt
```

## Requirements

```
pip install torch transformers datasets numpy tokenizers
```

## Notes

- Requested by GGUFGuy in [model-requests #30](https://huggingface.co/spaces/Compactbot/model-requests/discussions/30)
- The model is NOT trained yet — this is the script only.
- GPU recommended (RTX 3090+ for reasonable speed); CPU works but is ~50x slower.
- Checkpoints saved every 500 steps to the script directory.