# MoS-8B-Checkpoints Checkpoints of the Qwen3-8B MoS DFlash draft models (clean-800K E5 follow-up arms, 8B self-generated corpus `dataclean-8b-20260912`) and the warm-start inputs they were trained from. Project: MoS (mixture of layer-local expert MLPs added to a shared MLP in each DFlash draft layer). ## Layout ``` e5_8b_800k_20260831/arms//checkpoints/iter_NNNNNNN/hf/ portable export: config.json, model.safetensors, export_manifest.json (serving weights) e5_8b_800k_20260831/arms//checkpoints/iter_NNNNNNN/ full training checkpoint (DCP model/, optimizer/, lr_scheduler/, rng_rank_*.pt, meta.json) where listed below inputs// warm-start initializations (config.json, model.safetensors) ``` Milestones: 13,325 steps = 1 epoch (E1); 26,650 E2; 39,975 E3; 53,300 E4; 66,625 E5; 6,662 / 19,986 / 33,310 / 46,634 / 59,958 are half-epoch saves. ## Contents (2026-09-30) | Arm | Recipe | Archived | |---|---|---| | `e5-x4rand-mos` | K4 random-init experts, expert LR 6e-4 (rerun from 2026-09-21) | `hf` at 6,662, 26,650, 33,310, 39,975, 46,634, 53,300, 59,958, 66,625. E1 (13,325) and 19,986 were not archived and are lost | | `e5-x4rand-mos-r2` | continuation of the first x4rand run from its E1 weights, Adam reset (stopped, no reference value) | `hf` at 13,325, 19,986, 26,650 | | `e5-x8rand-mos` | K8 random-init experts, expert LR 6e-4 | `hf` at 46,634, 53,300, 59,958, 66,625 | | `e5-x16rand-mos` | K16 random-init experts, expert LR 6e-4, expert parallelism | `hf` at every save 6,662 … 66,625 | | `e5-x16lr4-mos` | K16 warm-start jitter init, expert LR ×4, expert parallelism (rerun paused at step 18,478) | `hf` at 6,662 and 13,325; full checkpoint (with optimizer) at 13,325 = resume point | | `e5-x4rand-lr1e3-mos` | ablation: x4rand with expert LR 1e-3, trained to E1 on the 5-epoch schedule | `hf` at 6,662; full checkpoint (with optimizer) at 13,325 | | Input | Content | |---|---| | `inputs/d0-shared-random-experts-8b-x4` | x4rand init: D0 shared MLP/attention + 4 random experts (seed 20260917) | | `inputs/d0-shared-random-experts-8b-x4-half` | half-width ablation init: the x4 init with each expert cut to its first 6,144 hidden units | | `inputs/d0-shared-random-experts-8b-x16` | x16rand init: D0 + 16 random experts (seed 20260919) | | `inputs/d0-fullwidth-init-8b-jitter` | K4 warm-start init (experts = D0 MLP + 1% jitter); also the source from which the random-expert inits are rebuilt | | `inputs/d0-fullwidth-init-8b-x16-jitter` | K16 warm-start init used by `e5-x16lr4-mos` | Results, configs and logs: GitHub `MoS` repository, `experiments/aurora/qwen3-8b-800k-e5/bal01/README.md` (§10). Archive records (file counts and byte totals, verified against the local copies before deletion): `experiments/aurora/qwen3-8b-800k-e5/archive/`.