Transformers
Safetensors
English
mla
deepseek-moe
mtp
custom-code
tinystories
from-scratch
Eval Results (legacy)
Instructions to use nowordsxiaomu/DeepSeek-Flash-Mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nowordsxiaomu/DeepSeek-Flash-Mini with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nowordsxiaomu/DeepSeek-Flash-Mini", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 5,530 Bytes
5e6d9f5 3bdad68 5e6d9f5 3bdad68 5e6d9f5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 | ---
license: mit
tags:
- mla
- deepseek-moe
- mtp
- custom-code
- tinystories
- from-scratch
language:
- en
library_name: transformers
model-index:
- name: DeepSeek-Flash-Mini (nano)
results:
- task:
type: language-modeling
name: TinyStories perplexity
dataset:
type: roneneldan/TinyStories
name: TinyStories
metrics:
- type: perplexity
value: 7.67
name: validation perplexity
---
# DeepSeek-Flash-Mini
A **from-scratch, lightweight Mixture-of-Experts language model** that reproduces the core
recipe of DeepSeek-V3 / V4-style efficient transformers at toy scale. It is fully trainable and
runnable on a laptop (Apple M-series MPS supported).
> β οΈ This is a **15M-parameter research toy** trained on ~5.3M tokens of TinyStories. It tells
> small coherent stories. It is **not** a general assistant and will score near random on
> knowledge/QA benchmarks (see MMLU note below). The value here is the **architecture**, not the
> raw score.
## Architecture
| Component | What it does |
|---|---|
| **MLA** (Multi-head Latent Attention) | Low-rank KV compression β only a `kv_lora_rank`-dim latent is cached, so KV memory is ~6.4Γ smaller than a same-size MHA. Two mathematically-equivalent paths: `naive` (explicit K/V restore, for training) and `absorb` (compute attention in latent space, for long-context decode). |
| **DeepSeekMoE** | Fine-grained routed experts + shared experts. Every token also passes a shared expert; the first `n_dense_layers` stay dense FFN. |
| **Aux-loss-free load balancing** | Each expert carries a gradient-free bias `b_i`; routing uses `argtopk(s_i + b_i)` while aggregation still uses the true `s_i`. Biases self-adjust every step toward balanced load β no auxiliary loss term needed. |
| **MTP** (Multi-Token Prediction) | A lightweight head predicts `t_{i+2}` sharing the embedding/output. At inference it acts as a **draft model for self-speculative decoding** (distribution-identical to autoregressive decoding), giving ~1.37Γ speedup. |
## Training
- **Config**: `nano` β dim 256, 6 layers, 8 routed experts (2 active) + 1 shared, `kv_lora_rank` 64.
- Total params **14.92M**, activated **6.90M** (46%).
- **Data**: TinyStories (22.5M chars β 5.3M tokens), custom BPE vocab 8192.
- **Recipe**: 3500 steps, cosine schedule + warmup, bf16, AdamW, gradient clip, MPS.
- **Time**: ~32 minutes on Apple M-series (MPS).
## Evaluation β where it actually ranks
The meaningful benchmark for this model is **TinyStories validation perplexity** (lower = better),
compared with other small models trained on the same corpus:
| Model | Params | Train tokens | Val PPL (TinyStories) |
|---|---:|---:|---:|
| karpathy/stories15M | 15.2M | ~1B | **2.92** |
| child-12m | 12.3M | ~1B | 3.45 |
| pluto-15M | 15M | ~1B | 3.64 |
| TinyStories-28M | 28M | ~1B | ~3.0β3.5 |
| **DeepSeek-Flash-Mini (nano)** | **14.9M** | **5.3M** | **7.67** |
| microgpt | ~? | 327M | 9.49 |
**Reading the table**: our 7.67 sits in the middle. The gap to the top models is dominated by
**training-token count (5.3M vs ~1B)**, not architecture β MLA/MoE/MTP here are an engineering
demonstration. (Cross-tokenizer perplexities are not strictly comparable; we use a custom BPE, so
a fully fair comparison would report bits-per-byte. The table is for rough orientation.)
### Efficiency metrics (measured)
| Metric | Value |
|---|---|
| KV cache vs same-size MHA | **6.4Γ smaller** (960 B vs 6144 B per token) |
| MTP speculative decoding | **1.37Γ speedup** (92% draft acceptance) |
| MPS decode throughput | ~40 tok/s (naive) |
### MMLU (subset, measured)
Loglikelihood multiple-choice (mean-NLL argmin), CPU, **10 subjects / 1,432 questions**:
- **Overall accuracy: 23.9%** (random-chance baseline 25%)
- Per-subject range: 17% (computer_security) β 31% (high_school_mathematics, abstract_algebra)
**Interpretation**: essentially at the random floor. Expected β the model was trained on 5.3M
tokens of TinyStories and holds no world knowledge. The small deviations (e.g. math 31%) are
statistical noise, not competence. Reported for honesty, **not** as a ranking claim. See
`mmlu_result.json` for the per-subject breakdown. The honest "ranking" for this model is the
TinyStories perplexity table above.
## How to use
```bash
pip install torch safetensors
# (this repo already bundles config.py / model/ / dataio/ / generate.py)
from load_and_generate import load_model
model, cfg = load_model(".") # needs config.json + model.safetensors in repo dir
# or CLI
python load_and_generate.py --prompt "Once upon a time" --max-new-tokens 80
python load_and_generate.py --prompt "Once upon a time" --spec # MTP speculative decoding
```
Loads the safetensors weights into the bundled model code and generates TinyStories-style text.
## Files
- `model.safetensors` β nano weights (68 MB)
- `config.json` β architecture hyperparameters (`ModelConfig` schema)
- `tokenizer.json` β custom BPE tokenizer (vocab 8192)
- `config.py`, `model/`, `dataio/`, `generate.py` β self-contained inference code
- `load_and_generate.py` β convenience loader + CLI
## Limitations
- Tiny context (512 tokens), English TinyStories only, no instruction-tuning.
- Not a chat/QA model; do not expect factual answers.
- The three presets (`nano`/`small`/`base`) are defined in `config.py`; only `nano` is trained and shipped here.
## License
MIT β Β© 2026 nowordsxiaomu.
|