--- title: NanoDex emoji: 🧬 colorFrom: green colorTo: blue sdk: docker app_port: 7860 pinned: true license: apache-2.0 short_description: Train a decoder-only LM from scratch on fineweb-edu hf_oauth: true hf_oauth_scopes: - read-repos - write-repos - manage-repos hf_oauth_expiration_minutes: 480 startup_duration_timeout: 45m datasets: - HuggingFaceFW/fineweb-edu --- # 🧬 NanoDex **Train your first decoder-only language model from scratch, in the cloud.** Not a fine-tune. Not an adapter. Random weights, a real corpus, and a real loss curve you watch fall in real time. Pick a size (**492k / 1.06M / 5.17M / 8.06M** parameters), pick a token budget (**200M – 1.5B** tokens of fineweb-edu), and queue it. A worker picks it up, you watch the loss fall live, and when it's done the model is pushed to *your* Hugging Face account with a generated model card. ## Pages | Route | What it is | |---|---| | `/` | Landing page when signed out; your dashboard when signed in | | `/train` | Four-step run wizard — size, tokens, name, review | | `/models` | Everything you've trained: training, queued (with position), ready | | `/models/{id}` | One run in detail — live loss chart, metrics, logs, publish, playground | | `/queue` | Global queue and worker status | | `/profile` | Account, totals, permissions | | `/about` | The full recipe | ## Architecture A standard modern decoder-only transformer (`LlamaForCausalLM`): SiLU MLP, RMSNorm (ε=1e-5), rotary position embeddings (θ=10000), grouped-query attention, tied input/output embeddings, no biases. The four tiers are that same recipe scaled down in width and depth. | Tier | Parameters | Layers | Hidden | Heads (KV) | FFN | Tokens / step | |---|---:|---:|---:|---|---:|---:| | NanoDex-500k | 492,192 | 3 | 96 | 6 (2) | 256 | 131,072 | | NanoDex-1M | 1,062,272 | 5 | 128 | 8 (4) | 288 | 262,144 | | NanoDex-5M | 5,172,384 | 9 | 224 | 8 (2) | 592 | 393,216 | | NanoDex-8M | 8,060,256 | 9 | 288 | 9 (3) | 704 | 524,288 | A typical small-LM vocabulary (~49k tokens) would be 28M embedding parameters on its own — more than three times the biggest model here. So NanoDex ships its own **2,048-token byte-level BPE** trained on fineweb-edu (~2.7 chars/token). The parameter counts above are real totals, embeddings included. ## Data [`HuggingFaceFW/fineweb-edu`](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu), config `sample-10BT`. A 1.505B-token slice is tokenized **once** into a flat `uint16` file (~2.8 GB) that every run reads from, so workers are never data-bound. Each run starts at its own random offset. Building the cache is a one-time cost per container; a run asking for more tokens than are cached yet waits for the writer and reports it in the run log. ## Training AdamW (β=0.9/0.95, wd=0.1, grad-clip 1.0) · 2% warmup then cosine decay to 10% of peak · bf16 autocast · 512-token context · 128–256 sequences per forward pass · 131k–524k tokens per optimizer step, with automatic micro-batch backoff if CUDA runs out of memory. One worker thread per GPU claims the oldest queued run atomically, so a 4-GPU Space trains four models at once. ## Layout ``` main.py FastAPI app — pages, JSON API, Hugging Face OAuth Dockerfile web/templates/ landing, home, train wizard, models, detail, queue, profile, about, error web/static/ app.css, app.js (dependency-free charts + polling) nanodex/config.py tier definitions + exact parameter counting nanodex/data.py shared fineweb-edu token cache nanodex/trainer.py the training loop nanodex/worker.py one worker per GPU nanodex/hub.py push a finished run to the user's account nanodex/estimate.py ETA, self-calibrating from completed runs nanodex/db.py SQLite job store tokenizer/ the 2,048-token BPE ``` ## A note on scale 1.5 billion tokens into an 8M-parameter model is ~9× past Chinchilla-optimal, and into a 500k-parameter model over 150× past it — on purpose. Tiny models keep improving long past the compute-optimal point, and the goal here isn't FLOP efficiency — it's watching cross-entropy fall from 7.6 (uniform over 2,048 tokens) toward 4, and seeing a network that started as pure noise begin to emit English words. Nothing trained here is a useful assistant. That was never the point. Built with 🤗 by [hugging-science](https://huggingface.co/hugging-science).