ryan-0608 commited on
Commit
bbe7aed
·
verified ·
1 Parent(s): 052f363

README: layout and contents (2026-09-30 archive)

Browse files
Files changed (1) hide show
  1. README.md +34 -0
README.md ADDED
@@ -0,0 +1,34 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # MoS-8B-Checkpoints
2
+
3
+ Checkpoints of the Qwen3-8B MoS DFlash draft models (clean-800K E5 follow-up arms, 8B self-generated corpus `dataclean-8b-20260912`) and the warm-start inputs they were trained from. Project: MoS (mixture of layer-local expert MLPs added to a shared MLP in each DFlash draft layer).
4
+
5
+ ## Layout
6
+
7
+ ```
8
+ e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/hf/ portable export: config.json, model.safetensors, export_manifest.json (serving weights)
9
+ e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/ full training checkpoint (DCP model/, optimizer/, lr_scheduler/, rng_rank_*.pt, meta.json) where listed below
10
+ inputs/<name>/ warm-start initializations (config.json, model.safetensors)
11
+ ```
12
+
13
+ Milestones: 13,325 steps = 1 epoch (E1); 26,650 E2; 39,975 E3; 53,300 E4; 66,625 E5; 6,662 / 19,986 / 33,310 / 46,634 / 59,958 are half-epoch saves.
14
+
15
+ ## Contents (2026-09-30)
16
+
17
+ | Arm | Recipe | Archived |
18
+ |---|---|---|
19
+ | `e5-x4rand-mos` | K4 random-init experts, expert LR 6e-4 (rerun from 2026-09-21) | `hf` at 6,662, 26,650, 33,310, 39,975, 46,634, 53,300, 59,958, 66,625. E1 (13,325) and 19,986 were not archived and are lost |
20
+ | `e5-x4rand-mos-r2` | continuation of the first x4rand run from its E1 weights, Adam reset (stopped, no reference value) | `hf` at 13,325, 19,986, 26,650 |
21
+ | `e5-x8rand-mos` | K8 random-init experts, expert LR 6e-4 | `hf` at 46,634, 53,300, 59,958, 66,625 |
22
+ | `e5-x16rand-mos` | K16 random-init experts, expert LR 6e-4, expert parallelism | `hf` at every save 6,662 … 66,625 |
23
+ | `e5-x16lr4-mos` | K16 warm-start jitter init, expert LR ×4, expert parallelism (rerun paused at step 18,478) | `hf` at 6,662 and 13,325; full checkpoint (with optimizer) at 13,325 = resume point |
24
+ | `e5-x4rand-lr1e3-mos` | ablation: x4rand with expert LR 1e-3, trained to E1 on the 5-epoch schedule | `hf` at 6,662; full checkpoint (with optimizer) at 13,325 |
25
+
26
+ | Input | Content |
27
+ |---|---|
28
+ | `inputs/d0-shared-random-experts-8b-x4` | x4rand init: D0 shared MLP/attention + 4 random experts (seed 20260917) |
29
+ | `inputs/d0-shared-random-experts-8b-x4-half` | half-width ablation init: the x4 init with each expert cut to its first 6,144 hidden units |
30
+ | `inputs/d0-shared-random-experts-8b-x16` | x16rand init: D0 + 16 random experts (seed 20260919) |
31
+ | `inputs/d0-fullwidth-init-8b-jitter` | K4 warm-start init (experts = D0 MLP + 1% jitter); also the source from which the random-expert inits are rebuilt |
32
+ | `inputs/d0-fullwidth-init-8b-x16-jitter` | K16 warm-start init used by `e5-x16lr4-mos` |
33
+
34
+ Results, configs and logs: GitHub `MoS` repository, `experiments/aurora/qwen3-8b-800k-e5/bal01/README.md` (§10). Archive records (file counts and byte totals, verified against the local copies before deletion): `experiments/aurora/qwen3-8b-800k-e5/archive/`.