File size: 5,177 Bytes
f1f3289
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
library_name: pytorch
pipeline_tag: text-generation
datasets:
  - roneneldan/TinyStories
language:
  - en
tags:
  - linear-attention
  - mixture-of-experts
  - delta-rule
  - muon
  - from-scratch
---

# miniKimiK3 β€” 50M

A 50M parameter language model written from scratch in plain PyTorch, combining **Kimi Delta Attention**, **Multi-head Latent Attention**, a **latent Mixture-of-Experts** with auxiliary-loss-free load balancing, **hyper-connection block residuals**, and a **Muon + AdamW hybrid optimizer**. Trained on TinyStories on a single Apple Silicon GPU.

No Triton, no CUDA kernels, no `transformers` β€” every component is implemented directly so the math is readable end to end.

**Code, full architecture notes and derivations:** https://github.com/Shiveshrane/MiniKimiK3

## Results

| | |
|---|---|
| parameters | 50.17M total, 36.71M active per token |
| validation loss | **1.6429** (perplexity β‰ˆ 5.17) |
| trained on | 82M tokens (10,000 steps Γ— 8,192 tokens), 1.22 epochs of 67M |
| hardware | one Apple Silicon GPU via MPS, 16.1 hours |
| throughput | ~2,300 tokens/s training, ~13 tokens/s single-stream decode |

## Samples

Prompt: `Once upon a time there was a little girl who`, temperature 0.8, top-k 50.

> Once upon a time there was a little girl who loved to go to the park. One day she was walking with her mom and she saw a big tree. She wanted to climb the tree, but her mom said no.
>
> The little girl was sad, but she kept walking. Suddenly, she saw a man walking towards the tree. He was very big and he had a bag in his hand.
>
> The little girl was scared, so she ran away. But then she heard a voice in the tree. It was the voice of someone yelling.

> Once upon a time there was a little girl who liked to take a bath. Her mommy would take her in the bathtub and filled it with warm water. The little girl liked to splash around in the warm water.
>
> One day, the little girl was feeling very miserable. She had no water, so she lay down in the bathtub and cried out loud.
>
> The mommy said, "What's wrong?"
>
> The little girl said, "I'm so miserable. I don't want to get out of the bathtub."

## Architecture

```
13 layers, 4n+1 rule    9 KDA + 4 MLA (layer 12 forced to MLA)
hidden 512              8 heads Γ— 64
KV / MoE latent 128     4Γ— compression
MoE per layer           16 routed (top-4) + 2 shared, StiGLU experts
vocab 8192              byte-level BPE, tied input/output embedding
```

- **Kimi Delta Attention** β€” linear attention whose state update is one step of gradient descent on $\lVert S^\top k - v\rVert^2$, giving an error-correcting write instead of blind accumulation, with per-channel gated decay. Trains through a chunked parallel form; decodes as a recurrence with a fixed $64\times64$ state per head, so cost per token is constant in sequence length.
- **Multi-head Latent Attention** every fourth layer, for exact long-range lookups a fixed-size state cannot hold.
- **Latent MoE** β€” routed experts operate inside a 128-wide latent rather than the 512-wide model dimension, which is what makes 16 experts per layer affordable at this scale.
- **Quantile balancing** β€” per-expert routing bias solved as a fixed point, so expert load is equalized with no auxiliary loss term and no loss weight to tune.
- **Muon** for the 2-D hidden matrices (Newton-Schulz orthogonalized momentum, per-head on q/k/v projections), AdamW for embeddings, norms and biases.

## Files

| file | |
|---|---|
| `best.pt` | checkpoint: weights, both optimizers' state, step, val loss |
| `tokenizer.json` | the byte-level BPE this checkpoint was trained with β€” ids are meaningless without it |

## Usage

```bash
git clone https://github.com/Shiveshrane/MiniKimiK3
cd MiniKimiK3

hf download ItsProtesilaus/MiniKimiK3 best.pt --local-dir checkpoints
hf download ItsProtesilaus/MiniKimiK3 tokenizer.json --local-dir data

python3 tests/generate.py "Once upon a time" --ckpt checkpoints/best.pt
```

The architecture is reconstructed from the checkpoint's stored config, so no flags need to match by hand.

## Limitations

- Trained only on TinyStories β€” simple-vocabulary children's stories. It has no world knowledge, no instruction following, no chat behaviour, and will not perform on any general benchmark.
- 82M training tokens for 50M parameters is well under a compute-optimal budget (~20 tokens/parameter would be 1B), so it is undertrained.
- No positional encoding of any kind: order information reaches the model only through KDA's decay and short convolutions plus the causal mask.
- `generate` is single-sequence; batched sampling needs a rework.
- Research and educational use. Not evaluated for safety, bias, or factuality.

## Training details

| | |
|---|---|
| corpus | 300k TinyStories β†’ 66.99M train / 0.38M val tokens, uint16 |
| batch | 8 Γ— 4 accumulation Γ— 256 tokens = 8,192 tokens/step |
| optimizer | Muon (lr 0.02) + AdamW (lr 3e-3), weight decay 0.1, grad clip 1.0 |
| schedule | 300 warmup steps, cosine to 10% over 10,000 steps |
| init | $\mathcal{N}(0, 0.02)$ β€” required with tied embeddings, which would otherwise start at loss β‰ˆ 494 instead of $\ln 8192 \approx 9.0$ |