{% extends "base.html" %} {% set nav = 'train' %} {% block title %}New run — NanoDex{% endblock %} {% block body %}
New training run

Build a model from nothing

Step 1
Size
Step 2
Tokens
Step 3
Name
Step 4
Review

How big should it be?

Every tier is the same architecture — a LlamaForCausalLM decoder-only transformer with SiLU MLPs, RMSNorm, rotary embeddings and grouped-query attention — scaled down in width and depth. The parameter counts below are totals, embeddings included.

{% for t in tiers %} {% endfor %}

How much should it read?

Training data is fineweb-edu — filtered educational web text. More tokens means a sharper model and a longer wait.

500M
200M500M1B1.5B
{% for m in [200, 500, 1000, 1500] %} {% endfor %}
Parameters— 
Optimizer steps—gradient updates
Estimated time— 
Tokens / param— 

Give it a name

This is what you'll see in your model list, and the default repository name when you publish it.

Letters, numbers, -, _ and .. Leave it blank and we'll name it after the size and token budget.

Will publish to {{ user.username }}/…

Ready to launch

Your run joins the shared queue. You can watch it — or close the tab and come back; it keeps going either way.

Model
Name—
Size—
Parameters—
Architecture—
Vocabulary{{ "{:,}".format(vocab) }} BPE
Context{{ seq_len }} tokens
Training
Datasetfineweb-edu
Token budget—
Optimizer steps—
OptimizerAdamW · cosine
Estimated time—
Queue limit{{ max_active }} runs per person
{% endblock %} {% block scripts %} {% endblock %}