File size: 6,192 Bytes
d7562c8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
---
license: cc-by-sa-4.0
language:
- en
library_name: transformers
pipeline_tag: text-generation
tags:
- from-scratch
- mixture-of-experts
- custom-gpu-stack
- research
datasets:
- ajaxdavis/alpha-er-corpus
---

# alpha-er

**alpha-er** (二, *èr* — "two") is a 100M-parameter language model trained end to end on a
**from-scratch GPU stack**: our own ioctl driver, memory manager, command channels, sm_86 SASS
assembler and kernel IR. **No CUDA, no cuBLAS, no vendor runtime** was involved in training —
every matrix multiply ran on hand-written machine code.

It was trained on **one RTX 3070** at **~96,000 tokens/second**, for 1.97B tokens in 5.7 hours.

## What it is, and what it is not

alpha-er writes fluent, grammatical English in the correct register for a prompt. **It is not
factually reliable and it does not answer questions correctly.** Validation perplexity is ~88;
this is a small model trained for one afternoon, published as a research artifact of the GPU
stack rather than as a useful assistant.

Real, unedited samples from the released weights (temperature 0.8, top-k 40):

| Prompt | Output |
|---|---|
| `<\|user\|>What is the capital of France?<\|assistant\|>` | *"To create a new and innovative approach, you can crafting a rich culture, and interests, such a rich, uniquely and diverse interests…"* |
| `The history of the Roman Empire` | *". In addition to the ancient Egyptian and ancient Egypt, but on the East Vietria. So, both are the elegance of the past, are a Germanician Greek civilization, the Greeks (1714)…"* |

Note what this does and does not show. The model has learned **which words belong together in a
history text** — Egypt, Greek civilization, antiquity — without learning any history. Syntax is
solid; semantics are weak. That is the honest reading of a 100M model at perplexity 88.

## Architecture

Three features make this **not a Llama**, and loading it as one would silently produce a
different model.

**Conditional MLP.** The feed-forward block is split into **G = 64 experts of width 320**. Each
token is routed to exactly one, so the model *stores* a 20,480-wide FFN but any token *pays for*
320. This is the identity the whole design rests on: FLOPs/token = 6 × **active** parameters, not
6 × total.

**Positional routing.** `expert(t) = floor(t · G / T)` — a token's expert depends only on its
position within its own sequence. An earlier version routed on the index in the *flattened batch*,
which made each sequence reach only 4 of 64 experts and made the weights meaningful only at the
exact batch shape they were trained at. Routing on `t` makes a checkpoint portable: the same
sequence gives identical logits at any batch width.

**Factored projections.** QKV, the attention output and the LM head are each a rank-128
bottleneck with a LayerNorm on the bottleneck. The norm is load-bearing — without it the factored
form diverged (grad_norm 51 against a dense baseline's 1.25).

| | |
|---|---|
| Parameters | 100,281,600 |
| Layers / d_model / heads | 2 / 1024 / 8 |
| FFN | 20,480 total, 320 active per token (G=64) |
| Context | 512 |
| Vocabulary | 12,288 (byte-level BPE) |
| Position encoding | learned |
| Attention | causal, **logit soft-cap 30** (`30·tanh(s/30)`) |
| Activation | GELU (tanh approximation) |

## Sequence length is part of the architecture

Expert boundaries fall at multiples of `T/G`, so the model only reproduces its training behaviour
at its trained context length. **Pad the prompt to 512 and read the logits at the last real
position.** This is exact, not an approximation: attention is causal, so padding after the prompt
cannot influence it, and each token's expert depends only on its own position. `generate()` in
`modeling_alpha.py` does this for you.

## Usage

```python
import torch
from modeling_alpha import AlphaErConfig, AlphaErForCausalLM
from tokenization_alpha import AlphaErTokenizer
from safetensors.torch import load_file
import json

cfg_d = json.load(open("config.json"))
cfg = AlphaErConfig(**{k: v for k, v in cfg_d.items()
                       if k in AlphaErConfig.__init__.__code__.co_varnames})
model = AlphaErForCausalLM(cfg)
model.load_state_dict(load_file("model.safetensors"), strict=False)
model.eval()

tok = AlphaErTokenizer.from_file("tokenizer_artifacts.json")
ids = tok.encode("<|user|>Hello!<|assistant|>")
out = model.generate(torch.tensor([ids]), max_new_tokens=60)[0].tolist()
print(tok.decode(out[len(ids):]))
```

`modeling_alpha.py` is a **PyTorch re-expression** of the trainer's forward pass, not the trainer
itself. It is checked elementwise against the real model at one position inside every one of the
64 expert windows: **max |Δlogit| = 6.8e-05, relative 3.7e-06** — float32 round-off.

## Training

| | |
|---|---|
| Tokens | 1.97B (20,000 steps × 98,304) |
| Batch | 16 × 512, gradient accumulation 12 |
| Optimizer | AdamW, lr 3e-4 cosine, warmup 500, weight decay 0.1, grad clip 1.0 |
| Loss | cross-entropy with **sampled softmax** (512 shared negatives) during training; full softmax for evaluation |
| Final val loss | 4.4803 (best 4.4119) — perplexity ~88 vs 12,288 for uniform |
| Throughput | ~96,000 tok/s on one RTX 3070 |

The validation curve flattened after roughly step 8,000, moving only 4.70 → ~4.45 over the second
half. The likely cause is **active capacity**: each token passes through a single 320-wide expert.
More steps would not fix that; fewer and wider experts would.

## Data

Trained on [`ajaxdavis/alpha-er-corpus`](https://huggingface.co/datasets/ajaxdavis/alpha-er-corpus)
— FineWeb-Edu/DCLM/FinePDFs, Concordance-EN, and SmolTalk. Licensed **CC-BY-SA-4.0**, inherited
from Concordance-EN's share-alike terms.

## Limitations

- **Not factually reliable.** It will confidently produce false statements.
- **No alignment, no safety tuning, no RLHF.** Trained on web text; it can reproduce the biases
  and content of that text.
- **512-token context**, and generation must pad to it.
- Repetition loops are common at low temperature.

Published as a research artifact demonstrating that a hand-built, fully-understood GPU stack can
train a real language model. Do not deploy it.