hyperdex-trainer / README.md
GGUFGuy's picture
Duplicate from GGUFGuy/hyperdex-trainer
2a49d9a
|
Raw History Blame Contribute Delete
4.48 kB
---
title: NanoDex
emoji: 🧬
colorFrom: green
colorTo: blue
sdk: docker
app_port: 7860
pinned: true
license: apache-2.0
short_description: Train a decoder-only LM from scratch on fineweb-edu
hf_oauth: true
hf_oauth_scopes:
- read-repos
- write-repos
- manage-repos
hf_oauth_expiration_minutes: 480
startup_duration_timeout: 45m
datasets:
- HuggingFaceFW/fineweb-edu
---
# 🧬 NanoDex
**Train your first decoder-only language model from scratch, in the cloud.**
Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss
curve you watch fall in real time.
Pick a size (**492k / 1.06M / 5.17M / 8.06M** parameters), pick a token budget
(**200M – 1.5B** tokens of fineweb-edu), and queue it. A worker picks it up, you
watch the loss fall live, and when it's done the model is pushed to *your*
Hugging Face account with a generated model card.
## Pages
| Route | What it is |
|---|---|
| `/` | Landing page when signed out; your dashboard when signed in |
| `/train` | Four-step run wizard — size, tokens, name, review |
| `/models` | Everything you've trained: training, queued (with position), ready |
| `/models/{id}` | One run in detail — live loss chart, metrics, logs, publish, playground |
| `/queue` | Global queue and worker status |
| `/profile` | Account, totals, permissions |
| `/about` | The full recipe |
## Architecture
A standard modern decoder-only transformer (`LlamaForCausalLM`): SiLU MLP,
RMSNorm (ε=1e-5), rotary position embeddings (θ=10000), grouped-query attention,
tied input/output embeddings, no biases. The four tiers are that same recipe
scaled down in width and depth.
| Tier | Parameters | Layers | Hidden | Heads (KV) | FFN | Tokens / step |
|---|---:|---:|---:|---|---:|---:|
| NanoDex-500k | 492,192 | 3 | 96 | 6 (2) | 256 | 131,072 |
| NanoDex-1M | 1,062,272 | 5 | 128 | 8 (4) | 288 | 262,144 |
| NanoDex-5M | 5,172,384 | 9 | 224 | 8 (2) | 592 | 393,216 |
| NanoDex-8M | 8,060,256 | 9 | 288 | 9 (3) | 704 | 524,288 |
A typical small-LM vocabulary (~49k tokens) would be 28M embedding parameters on
its own — more than three times the biggest model here. So NanoDex ships its own
**2,048-token byte-level BPE** trained on fineweb-edu (~2.7 chars/token). The
parameter counts above are real totals, embeddings included.
## Data
[`HuggingFaceFW/fineweb-edu`](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu),
config `sample-10BT`. A 1.505B-token slice is tokenized **once** into a flat
`uint16` file (~2.8 GB) that every run reads from, so workers are never
data-bound. Each run starts at its own random offset. Building the cache is a
one-time cost per container; a run asking for more tokens than are cached yet
waits for the writer and reports it in the run log.
## Training
AdamW (β=0.9/0.95, wd=0.1, grad-clip 1.0) · 2% warmup then cosine decay to 10%
of peak · bf16 autocast · 512-token context · 128–256 sequences per forward
pass · 131k–524k tokens per optimizer step, with automatic micro-batch backoff
if CUDA runs out of memory. One worker thread per GPU claims the oldest queued run
atomically, so a 4-GPU Space trains four models at once.
## Layout
```
main.py FastAPI app — pages, JSON API, Hugging Face OAuth
Dockerfile
web/templates/ landing, home, train wizard, models, detail, queue,
profile, about, error
web/static/ app.css, app.js (dependency-free charts + polling)
nanodex/config.py tier definitions + exact parameter counting
nanodex/data.py shared fineweb-edu token cache
nanodex/trainer.py the training loop
nanodex/worker.py one worker per GPU
nanodex/hub.py push a finished run to the user's account
nanodex/estimate.py ETA, self-calibrating from completed runs
nanodex/db.py SQLite job store
tokenizer/ the 2,048-token BPE
```
## A note on scale
1.5 billion tokens into an 8M-parameter model is ~9× past Chinchilla-optimal,
and into a 500k-parameter model over 150× past it — on purpose. Tiny models keep
improving long past the compute-optimal point, and the goal here isn't FLOP
efficiency — it's watching cross-entropy fall from 7.6
(uniform over 2,048 tokens) toward 4, and seeing a network that started as pure
noise begin to emit English words.
Nothing trained here is a useful assistant. That was never the point.
Built with 🤗 by [hugging-science](https://huggingface.co/hugging-science).