| --- |
| license: mit |
| library_name: pytorch |
| language: |
| - en |
| tags: |
| - nanochat |
| - language-model |
| - pre-1930 |
| - historical |
| - vintage |
| - research |
| --- |
| |
| # BART Experiments |
|
|
| The complete training archive behind [BART](https://huggingface.co/jbduran/bart) — 39 runs, |
| their checkpoints, tokenizers, and evaluations, including every dead end. |
|
|
| 📝 [Read the write-up](https://www.unboundedlab.com/blog/bart) · 🌐 [Unbounded Labs](https://unboundedlab.com) |
|
|
| If you want the model itself, use [bart](https://huggingface.co/jbduran/bart) or |
| [bart-sft](https://huggingface.co/jbduran/bart-sft). This repository is for reproducing or |
| inspecting how they were reached. |
|
|
| ## Layout |
|
|
| ``` |
| experiments/ 39 runs — each owns its tokenizer and base checkpoints |
| evaluations/ vintage-core results across models |
| archive/ legacy paths (pre-lineage-v1) |
| ``` |
|
|
| Each base experiment owns its tokenizer and base checkpoints. SFT runs nest under their exact base |
| parent (`experiments/<base>/sft/<sft-run>/`), and post-training runs nest under their exact SFT |
| parent. |
|
|
| ## Reading experiment names |
|
|
| | Fragment | Meaning | |
| |---|---| |
| | `d12`, `d24`, `d32` | model depth | |
| | `r11`–`r30` | target parameter-to-data ratio | |
| | `ctx4096`, `ctx8192` | max sequence length | |
| | `sssl` | window pattern | |
| | `fulltok`, `randtok` | tokenizer variant | |
|
|
| The prefix tells you the corpus generation: |
|
|
| | Prefix | Corpus | |
| |---|---| |
| | `think-` | [bart-dataset-v1](https://huggingface.co/datasets/jbduran/bart-dataset-v1) | |
| | `thinkcleaned-` | [v2](https://huggingface.co/datasets/jbduran/bart-dataset-v2) | |
| | `clean1930s-` | [v3](https://huggingface.co/datasets/jbduran/bart-dataset-v3) | |
| | `Think.Unbounded-` | v3 plus [midtraining](https://huggingface.co/datasets/zachnorton03/bart-midtrain) | |
|
|
| ## The runs that matter |
|
|
| `Think.Unbounded-d32` is the first d32 run, trained on midtrain mixtures whose ratios were computed |
| by **document count**. Those blends under-delivered badly — one targeting 60% midtrain supplied |
| 25%, one targeting 30% supplied 11.5%. |
|
|
| `Think.Unbounded-d32-v2mix-cont` is the repair: it branches from that run at step 5500, carries the |
| optimizer state, and continues on corrected **token-based** mixtures. It is the model shipped as |
| [bart](https://huggingface.co/jbduran/bart). The two configs side by side are the clearest record |
| of that bug and its fix. |
|
|
| Its `sft/` directory holds six fine-tuning variants. The one shipped as |
| [bart-sft](https://huggingface.co/jbduran/bart-sft) is |
| `pre1930-curriculum-c3-robust-v2`; the others — `karpathy-modern-sft-v1`, |
| `nanochat-default-datamatch-v1`, `c3-robust`, and two `c3-robust-v3` configs — are kept here. |
|
|
| `clean1930s-d24-r12-ctx4096-sssl-fulltok-v1` is the d24 / 1.38B run; its tokenizer is the one every |
| midtrain mixture was built against. |
|
|
| > **Note.** Each `config.json` records repo and dataset names as they were at training time |
| > (`jbduran/think.nano`, `think-dataset-clean-1930s`, `think-midtrain`). Repo renames still resolve |
| > via Hub redirects, but the midtrain mixture subfolders they reference (`mixed/v2/ratio_21/data`) |
| > moved to `mixtures/v2-by-tokens/ratio_21/data` when that dataset was restructured. Configs are |
| > kept as historical records rather than rewritten. |
|
|
| ## Related |
|
|
| [bart](https://huggingface.co/jbduran/bart) · |
| [bart-sft](https://huggingface.co/jbduran/bart-sft) · |
| [bart-dataset-v3](https://huggingface.co/datasets/jbduran/bart-dataset-v3) · |
| [bart-midtrain](https://huggingface.co/datasets/zachnorton03/bart-midtrain) · |
| [bart-dataset-scripts](https://github.com/OwenVoorhees/bart-dataset-scripts) · |
| [bart-midtrain-scripts](https://github.com/OwenVoorhees/bart-midtrain-scripts) |
|
|
| ## License |
|
|
| MIT |
|
|
| --- |
|
|
| Built by [Unbounded Labs](https://unboundedlab.com). |
|
|