YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
MoS-8B-Checkpoints
Checkpoints of the Qwen3-8B MoS DFlash draft models (clean-800K E5 follow-up arms, 8B self-generated corpus dataclean-8b-20260912) and the warm-start inputs they were trained from. Project: MoS (mixture of layer-local expert MLPs added to a shared MLP in each DFlash draft layer).
Layout
e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/hf/ portable export: config.json, model.safetensors, export_manifest.json (serving weights)
e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/ full training checkpoint (DCP model/, optimizer/, lr_scheduler/, rng_rank_*.pt, meta.json) where listed below
inputs/<name>/ warm-start initializations (config.json, model.safetensors)
Milestones: 13,325 steps = 1 epoch (E1); 26,650 E2; 39,975 E3; 53,300 E4; 66,625 E5; 6,662 / 19,986 / 33,310 / 46,634 / 59,958 are half-epoch saves.
Contents (2026-09-30)
| Arm | Recipe | Archived |
|---|---|---|
e5-x4rand-mos |
K4 random-init experts, expert LR 6e-4 (rerun from 2026-09-21) | hf at 6,662, 26,650, 33,310, 39,975, 46,634, 53,300, 59,958, 66,625. E1 (13,325) and 19,986 were not archived and are lost |
e5-x4rand-mos-r2 |
continuation of the first x4rand run from its E1 weights, Adam reset (stopped, no reference value) | hf at 13,325, 19,986, 26,650 |
e5-x8rand-mos |
K8 random-init experts, expert LR 6e-4 | hf at 46,634, 53,300, 59,958, 66,625 |
e5-x16rand-mos |
K16 random-init experts, expert LR 6e-4, expert parallelism | hf at every save 6,662 … 66,625 |
e5-x16lr4-mos |
K16 warm-start jitter init, expert LR ×4, expert parallelism (rerun paused at step 18,478) | hf at 6,662 and 13,325; full checkpoint (with optimizer) at 13,325 = resume point |
e5-x4rand-lr1e3-mos |
ablation: x4rand with expert LR 1e-3, trained to E1 on the 5-epoch schedule | hf at 6,662; full checkpoint (with optimizer) at 13,325 |
| Input | Content |
|---|---|
inputs/d0-shared-random-experts-8b-x4 |
x4rand init: D0 shared MLP/attention + 4 random experts (seed 20260917) |
inputs/d0-shared-random-experts-8b-x4-half |
half-width ablation init: the x4 init with each expert cut to its first 6,144 hidden units |
inputs/d0-shared-random-experts-8b-x16 |
x16rand init: D0 + 16 random experts (seed 20260919) |
inputs/d0-fullwidth-init-8b-jitter |
K4 warm-start init (experts = D0 MLP + 1% jitter); also the source from which the random-expert inits are rebuilt |
inputs/d0-fullwidth-init-8b-x16-jitter |
K16 warm-start init used by e5-x16lr4-mos |
Results, configs and logs: GitHub MoS repository, experiments/aurora/qwen3-8b-800k-e5/bal01/README.md (§10). Archive records (file counts and byte totals, verified against the local copies before deletion): experiments/aurora/qwen3-8b-800k-e5/archive/.