{% extends "base.html" %} {% set nav = 'train' %} {% block title %}New run — NanoDex{% endblock %} {% block body %}
Every tier is the same architecture — a LlamaForCausalLM decoder-only transformer with SiLU MLPs, RMSNorm, rotary embeddings and grouped-query attention — scaled down in width and depth. The parameter counts below are totals, embeddings included.
Training data is fineweb-edu — filtered educational web text. More tokens means a sharper model and a longer wait.
This is what you'll see in your model list, and the default repository name when you publish it.
Letters, numbers, -, _ and .. Leave it blank and we'll name it after the size and token budget.
Will publish to {{ user.username }}/…
Your run joins the shared queue. You can watch it — or close the tab and come back; it keeps going either way.