File size: 3,724 Bytes
3d2a742
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9b7cebb
3d2a742
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
---
license: mit
library_name: pytorch
language:
- en
tags:
- nanochat
- language-model
- pre-1930
- historical
- vintage
- research
---

# BART Experiments

The complete training archive behind [BART](https://huggingface.co/jbduran/bart) — 39 runs,
their checkpoints, tokenizers, and evaluations, including every dead end.

📝 [Read the write-up](https://www.unboundedlab.com/blog/bart) · 🌐 [Unbounded Labs](https://unboundedlab.com)

If you want the model itself, use [bart](https://huggingface.co/jbduran/bart) or
[bart-sft](https://huggingface.co/jbduran/bart-sft). This repository is for reproducing or
inspecting how they were reached.

## Layout

```
experiments/    39 runs — each owns its tokenizer and base checkpoints
evaluations/    vintage-core results across models
archive/        legacy paths (pre-lineage-v1)
```

Each base experiment owns its tokenizer and base checkpoints. SFT runs nest under their exact base
parent (`experiments/<base>/sft/<sft-run>/`), and post-training runs nest under their exact SFT
parent.

## Reading experiment names

| Fragment | Meaning |
|---|---|
| `d12`, `d24`, `d32` | model depth |
| `r11``r30` | target parameter-to-data ratio |
| `ctx4096`, `ctx8192` | max sequence length |
| `sssl` | window pattern |
| `fulltok`, `randtok` | tokenizer variant |

The prefix tells you the corpus generation:

| Prefix | Corpus |
|---|---|
| `think-` | [bart-dataset-v1](https://huggingface.co/datasets/jbduran/bart-dataset-v1) |
| `thinkcleaned-` | [v2](https://huggingface.co/datasets/jbduran/bart-dataset-v2) |
| `clean1930s-` | [v3](https://huggingface.co/datasets/jbduran/bart-dataset-v3) |
| `Think.Unbounded-` | v3 plus [midtraining](https://huggingface.co/datasets/zachnorton03/bart-midtrain) |

## The runs that matter

`Think.Unbounded-d32` is the first d32 run, trained on midtrain mixtures whose ratios were computed
by **document count**. Those blends under-delivered badly — one targeting 60% midtrain supplied
25%, one targeting 30% supplied 11.5%.

`Think.Unbounded-d32-v2mix-cont` is the repair: it branches from that run at step 5500, carries the
optimizer state, and continues on corrected **token-based** mixtures. It is the model shipped as
[bart](https://huggingface.co/jbduran/bart). The two configs side by side are the clearest record
of that bug and its fix.

Its `sft/` directory holds six fine-tuning variants. The one shipped as
[bart-sft](https://huggingface.co/jbduran/bart-sft) is
`pre1930-curriculum-c3-robust-v2`; the others — `karpathy-modern-sft-v1`,
`nanochat-default-datamatch-v1`, `c3-robust`, and two `c3-robust-v3` configs — are kept here.

`clean1930s-d24-r12-ctx4096-sssl-fulltok-v1` is the d24 / 1.38B run; its tokenizer is the one every
midtrain mixture was built against.

> **Note.** Each `config.json` records repo and dataset names as they were at training time
> (`jbduran/think.nano`, `think-dataset-clean-1930s`, `think-midtrain`). Repo renames still resolve
> via Hub redirects, but the midtrain mixture subfolders they reference (`mixed/v2/ratio_21/data`)
> moved to `mixtures/v2-by-tokens/ratio_21/data` when that dataset was restructured. Configs are
> kept as historical records rather than rewritten.

## Related

[bart](https://huggingface.co/jbduran/bart) ·
[bart-sft](https://huggingface.co/jbduran/bart-sft) ·
[bart-dataset-v3](https://huggingface.co/datasets/jbduran/bart-dataset-v3) ·
[bart-midtrain](https://huggingface.co/datasets/zachnorton03/bart-midtrain) ·
[bart-dataset-scripts](https://github.com/OwenVoorhees/bart-dataset-scripts) ·
[bart-midtrain-scripts](https://github.com/OwenVoorhees/bart-midtrain-scripts)

## License

MIT

---

Built by [Unbounded Labs](https://unboundedlab.com).