README: layout and contents (2026-09-30 archive)
Browse files
README.md
ADDED
|
@@ -0,0 +1,34 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# MoS-8B-Checkpoints
|
| 2 |
+
|
| 3 |
+
Checkpoints of the Qwen3-8B MoS DFlash draft models (clean-800K E5 follow-up arms, 8B self-generated corpus `dataclean-8b-20260912`) and the warm-start inputs they were trained from. Project: MoS (mixture of layer-local expert MLPs added to a shared MLP in each DFlash draft layer).
|
| 4 |
+
|
| 5 |
+
## Layout
|
| 6 |
+
|
| 7 |
+
```
|
| 8 |
+
e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/hf/ portable export: config.json, model.safetensors, export_manifest.json (serving weights)
|
| 9 |
+
e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/ full training checkpoint (DCP model/, optimizer/, lr_scheduler/, rng_rank_*.pt, meta.json) where listed below
|
| 10 |
+
inputs/<name>/ warm-start initializations (config.json, model.safetensors)
|
| 11 |
+
```
|
| 12 |
+
|
| 13 |
+
Milestones: 13,325 steps = 1 epoch (E1); 26,650 E2; 39,975 E3; 53,300 E4; 66,625 E5; 6,662 / 19,986 / 33,310 / 46,634 / 59,958 are half-epoch saves.
|
| 14 |
+
|
| 15 |
+
## Contents (2026-09-30)
|
| 16 |
+
|
| 17 |
+
| Arm | Recipe | Archived |
|
| 18 |
+
|---|---|---|
|
| 19 |
+
| `e5-x4rand-mos` | K4 random-init experts, expert LR 6e-4 (rerun from 2026-09-21) | `hf` at 6,662, 26,650, 33,310, 39,975, 46,634, 53,300, 59,958, 66,625. E1 (13,325) and 19,986 were not archived and are lost |
|
| 20 |
+
| `e5-x4rand-mos-r2` | continuation of the first x4rand run from its E1 weights, Adam reset (stopped, no reference value) | `hf` at 13,325, 19,986, 26,650 |
|
| 21 |
+
| `e5-x8rand-mos` | K8 random-init experts, expert LR 6e-4 | `hf` at 46,634, 53,300, 59,958, 66,625 |
|
| 22 |
+
| `e5-x16rand-mos` | K16 random-init experts, expert LR 6e-4, expert parallelism | `hf` at every save 6,662 … 66,625 |
|
| 23 |
+
| `e5-x16lr4-mos` | K16 warm-start jitter init, expert LR ×4, expert parallelism (rerun paused at step 18,478) | `hf` at 6,662 and 13,325; full checkpoint (with optimizer) at 13,325 = resume point |
|
| 24 |
+
| `e5-x4rand-lr1e3-mos` | ablation: x4rand with expert LR 1e-3, trained to E1 on the 5-epoch schedule | `hf` at 6,662; full checkpoint (with optimizer) at 13,325 |
|
| 25 |
+
|
| 26 |
+
| Input | Content |
|
| 27 |
+
|---|---|
|
| 28 |
+
| `inputs/d0-shared-random-experts-8b-x4` | x4rand init: D0 shared MLP/attention + 4 random experts (seed 20260917) |
|
| 29 |
+
| `inputs/d0-shared-random-experts-8b-x4-half` | half-width ablation init: the x4 init with each expert cut to its first 6,144 hidden units |
|
| 30 |
+
| `inputs/d0-shared-random-experts-8b-x16` | x16rand init: D0 + 16 random experts (seed 20260919) |
|
| 31 |
+
| `inputs/d0-fullwidth-init-8b-jitter` | K4 warm-start init (experts = D0 MLP + 1% jitter); also the source from which the random-expert inits are rebuilt |
|
| 32 |
+
| `inputs/d0-fullwidth-init-8b-x16-jitter` | K16 warm-start init used by `e5-x16lr4-mos` |
|
| 33 |
+
|
| 34 |
+
Results, configs and logs: GitHub `MoS` repository, `experiments/aurora/qwen3-8b-800k-e5/bal01/README.md` (§10). Archive records (file counts and byte totals, verified against the local copies before deletion): `experiments/aurora/qwen3-8b-800k-e5/archive/`.
|