|
Download README.md from SLM-Archive/hyperdex-trainer: direct link, hf CLI and curl.
- Browser
- Download file 4.48 kB
-
https://huggingface.co/SLM-Archive/hyperdex-trainer/resolve/main/README.md
- Command line
-
hf download hf://SLM-Archive/hyperdex-trainer/README.md
-
curl -L -o README.md https://huggingface.co/SLM-Archive/hyperdex-trainer/resolve/main/README.md
4.48 kB
| title: NanoDex | |
| emoji: 🧬 | |
| colorFrom: green | |
| colorTo: blue | |
| sdk: docker | |
| app_port: 7860 | |
| pinned: true | |
| license: apache-2.0 | |
| short_description: Train a decoder-only LM from scratch on fineweb-edu | |
| hf_oauth: true | |
| hf_oauth_scopes: | |
| - read-repos | |
| - write-repos | |
| - manage-repos | |
| hf_oauth_expiration_minutes: 480 | |
| startup_duration_timeout: 45m | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| # 🧬 NanoDex | |
| **Train your first decoder-only language model from scratch, in the cloud.** | |
| Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss | |
| curve you watch fall in real time. | |
| Pick a size (**492k / 1.06M / 5.17M / 8.06M** parameters), pick a token budget | |
| (**200M – 1.5B** tokens of fineweb-edu), and queue it. A worker picks it up, you | |
| watch the loss fall live, and when it's done the model is pushed to *your* | |
| Hugging Face account with a generated model card. | |
| ## Pages | |
| | Route | What it is | | |
| |---|---| | |
| | `/` | Landing page when signed out; your dashboard when signed in | | |
| | `/train` | Four-step run wizard — size, tokens, name, review | | |
| | `/models` | Everything you've trained: training, queued (with position), ready | | |
| | `/models/{id}` | One run in detail — live loss chart, metrics, logs, publish, playground | | |
| | `/queue` | Global queue and worker status | | |
| | `/profile` | Account, totals, permissions | | |
| | `/about` | The full recipe | | |
| ## Architecture | |
| A standard modern decoder-only transformer (`LlamaForCausalLM`): SiLU MLP, | |
| RMSNorm (ε=1e-5), rotary position embeddings (θ=10000), grouped-query attention, | |
| tied input/output embeddings, no biases. The four tiers are that same recipe | |
| scaled down in width and depth. | |
| | Tier | Parameters | Layers | Hidden | Heads (KV) | FFN | Tokens / step | | |
| |---|---:|---:|---:|---|---:|---:| | |
| | NanoDex-500k | 492,192 | 3 | 96 | 6 (2) | 256 | 131,072 | | |
| | NanoDex-1M | 1,062,272 | 5 | 128 | 8 (4) | 288 | 262,144 | | |
| | NanoDex-5M | 5,172,384 | 9 | 224 | 8 (2) | 592 | 393,216 | | |
| | NanoDex-8M | 8,060,256 | 9 | 288 | 9 (3) | 704 | 524,288 | | |
| A typical small-LM vocabulary (~49k tokens) would be 28M embedding parameters on | |
| its own — more than three times the biggest model here. So NanoDex ships its own | |
| **2,048-token byte-level BPE** trained on fineweb-edu (~2.7 chars/token). The | |
| parameter counts above are real totals, embeddings included. | |
| ## Data | |
| [`HuggingFaceFW/fineweb-edu`](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), | |
| config `sample-10BT`. A 1.505B-token slice is tokenized **once** into a flat | |
| `uint16` file (~2.8 GB) that every run reads from, so workers are never | |
| data-bound. Each run starts at its own random offset. Building the cache is a | |
| one-time cost per container; a run asking for more tokens than are cached yet | |
| waits for the writer and reports it in the run log. | |
| ## Training | |
| AdamW (β=0.9/0.95, wd=0.1, grad-clip 1.0) · 2% warmup then cosine decay to 10% | |
| of peak · bf16 autocast · 512-token context · 128–256 sequences per forward | |
| pass · 131k–524k tokens per optimizer step, with automatic micro-batch backoff | |
| if CUDA runs out of memory. One worker thread per GPU claims the oldest queued run | |
| atomically, so a 4-GPU Space trains four models at once. | |
| ## Layout | |
| ``` | |
| main.py FastAPI app — pages, JSON API, Hugging Face OAuth | |
| Dockerfile | |
| web/templates/ landing, home, train wizard, models, detail, queue, | |
| profile, about, error | |
| web/static/ app.css, app.js (dependency-free charts + polling) | |
| nanodex/config.py tier definitions + exact parameter counting | |
| nanodex/data.py shared fineweb-edu token cache | |
| nanodex/trainer.py the training loop | |
| nanodex/worker.py one worker per GPU | |
| nanodex/hub.py push a finished run to the user's account | |
| nanodex/estimate.py ETA, self-calibrating from completed runs | |
| nanodex/db.py SQLite job store | |
| tokenizer/ the 2,048-token BPE | |
| ``` | |
| ## A note on scale | |
| 1.5 billion tokens into an 8M-parameter model is ~9× past Chinchilla-optimal, | |
| and into a 500k-parameter model over 150× past it — on purpose. Tiny models keep | |
| improving long past the compute-optimal point, and the goal here isn't FLOP | |
| efficiency — it's watching cross-entropy fall from 7.6 | |
| (uniform over 2,048 tokens) toward 4, and seeing a network that started as pure | |
| noise begin to emit English words. | |
| Nothing trained here is a useful assistant. That was never the point. | |
| Built with 🤗 by [hugging-science](https://huggingface.co/hugging-science). | |