File size: 5,530 Bytes
5e6d9f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3bdad68
5e6d9f5
3bdad68
 
 
 
 
 
 
 
 
 
5e6d9f5
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
---
license: mit
tags:
  - mla
  - deepseek-moe
  - mtp
  - custom-code
  - tinystories
  - from-scratch
language:
  - en
library_name: transformers
model-index:
  - name: DeepSeek-Flash-Mini (nano)
    results:
      - task:
          type: language-modeling
          name: TinyStories perplexity
        dataset:
          type: roneneldan/TinyStories
          name: TinyStories
        metrics:
          - type: perplexity
            value: 7.67
            name: validation perplexity
---

# DeepSeek-Flash-Mini

A **from-scratch, lightweight Mixture-of-Experts language model** that reproduces the core
recipe of DeepSeek-V3 / V4-style efficient transformers at toy scale. It is fully trainable and
runnable on a laptop (Apple M-series MPS supported).

> ⚠️ This is a **15M-parameter research toy** trained on ~5.3M tokens of TinyStories. It tells
> small coherent stories. It is **not** a general assistant and will score near random on
> knowledge/QA benchmarks (see MMLU note below). The value here is the **architecture**, not the
> raw score.

## Architecture

| Component | What it does |
|---|---|
| **MLA** (Multi-head Latent Attention) | Low-rank KV compression β€” only a `kv_lora_rank`-dim latent is cached, so KV memory is ~6.4Γ— smaller than a same-size MHA. Two mathematically-equivalent paths: `naive` (explicit K/V restore, for training) and `absorb` (compute attention in latent space, for long-context decode). |
| **DeepSeekMoE** | Fine-grained routed experts + shared experts. Every token also passes a shared expert; the first `n_dense_layers` stay dense FFN. |
| **Aux-loss-free load balancing** | Each expert carries a gradient-free bias `b_i`; routing uses `argtopk(s_i + b_i)` while aggregation still uses the true `s_i`. Biases self-adjust every step toward balanced load β€” no auxiliary loss term needed. |
| **MTP** (Multi-Token Prediction) | A lightweight head predicts `t_{i+2}` sharing the embedding/output. At inference it acts as a **draft model for self-speculative decoding** (distribution-identical to autoregressive decoding), giving ~1.37Γ— speedup. |

## Training

- **Config**: `nano` β€” dim 256, 6 layers, 8 routed experts (2 active) + 1 shared, `kv_lora_rank` 64.
  - Total params **14.92M**, activated **6.90M** (46%).
- **Data**: TinyStories (22.5M chars β†’ 5.3M tokens), custom BPE vocab 8192.
- **Recipe**: 3500 steps, cosine schedule + warmup, bf16, AdamW, gradient clip, MPS.
- **Time**: ~32 minutes on Apple M-series (MPS).

## Evaluation β€” where it actually ranks

The meaningful benchmark for this model is **TinyStories validation perplexity** (lower = better),
compared with other small models trained on the same corpus:

| Model | Params | Train tokens | Val PPL (TinyStories) |
|---|---:|---:|---:|
| karpathy/stories15M | 15.2M | ~1B | **2.92** |
| child-12m | 12.3M | ~1B | 3.45 |
| pluto-15M | 15M | ~1B | 3.64 |
| TinyStories-28M | 28M | ~1B | ~3.0–3.5 |
| **DeepSeek-Flash-Mini (nano)** | **14.9M** | **5.3M** | **7.67** |
| microgpt | ~? | 327M | 9.49 |

**Reading the table**: our 7.67 sits in the middle. The gap to the top models is dominated by
**training-token count (5.3M vs ~1B)**, not architecture β€” MLA/MoE/MTP here are an engineering
demonstration. (Cross-tokenizer perplexities are not strictly comparable; we use a custom BPE, so
a fully fair comparison would report bits-per-byte. The table is for rough orientation.)

### Efficiency metrics (measured)

| Metric | Value |
|---|---|
| KV cache vs same-size MHA | **6.4Γ— smaller** (960 B vs 6144 B per token) |
| MTP speculative decoding | **1.37Γ— speedup** (92% draft acceptance) |
| MPS decode throughput | ~40 tok/s (naive) |

### MMLU (subset, measured)

Loglikelihood multiple-choice (mean-NLL argmin), CPU, **10 subjects / 1,432 questions**:

- **Overall accuracy: 23.9%** (random-chance baseline 25%)
- Per-subject range: 17% (computer_security) – 31% (high_school_mathematics, abstract_algebra)

**Interpretation**: essentially at the random floor. Expected β€” the model was trained on 5.3M
tokens of TinyStories and holds no world knowledge. The small deviations (e.g. math 31%) are
statistical noise, not competence. Reported for honesty, **not** as a ranking claim. See
`mmlu_result.json` for the per-subject breakdown. The honest "ranking" for this model is the
TinyStories perplexity table above.

## How to use

```bash
pip install torch safetensors
# (this repo already bundles config.py / model/ / dataio/ / generate.py)

from load_and_generate import load_model
model, cfg = load_model(".")          # needs config.json + model.safetensors in repo dir

# or CLI
python load_and_generate.py --prompt "Once upon a time" --max-new-tokens 80
python load_and_generate.py --prompt "Once upon a time" --spec   # MTP speculative decoding
```

Loads the safetensors weights into the bundled model code and generates TinyStories-style text.

## Files

- `model.safetensors` β€” nano weights (68 MB)
- `config.json` β€” architecture hyperparameters (`ModelConfig` schema)
- `tokenizer.json` β€” custom BPE tokenizer (vocab 8192)
- `config.py`, `model/`, `dataio/`, `generate.py` β€” self-contained inference code
- `load_and_generate.py` β€” convenience loader + CLI

## Limitations

- Tiny context (512 tokens), English TinyStories only, no instruction-tuning.
- Not a chat/QA model; do not expect factual answers.
- The three presets (`nano`/`small`/`base`) are defined in `config.py`; only `nano` is trained and shipped here.

## License

MIT β€” Β© 2026 nowordsxiaomu.