YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

MoS-8B-Checkpoints

Checkpoints of the Qwen3-8B MoS DFlash draft models (clean-800K E5 follow-up arms, 8B self-generated corpus dataclean-8b-20260912) and the warm-start inputs they were trained from. Project: MoS (mixture of layer-local expert MLPs added to a shared MLP in each DFlash draft layer).

Layout

e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/hf/      portable export: config.json, model.safetensors, export_manifest.json (serving weights)
e5_8b_800k_20260831/arms/<arm>/checkpoints/iter_NNNNNNN/         full training checkpoint (DCP model/, optimizer/, lr_scheduler/, rng_rank_*.pt, meta.json) where listed below
inputs/<name>/                                                   warm-start initializations (config.json, model.safetensors)

Milestones: 13,325 steps = 1 epoch (E1); 26,650 E2; 39,975 E3; 53,300 E4; 66,625 E5; 6,662 / 19,986 / 33,310 / 46,634 / 59,958 are half-epoch saves.

Contents (2026-09-30)

Arm Recipe Archived
e5-x4rand-mos K4 random-init experts, expert LR 6e-4 (rerun from 2026-09-21) hf at 6,662, 26,650, 33,310, 39,975, 46,634, 53,300, 59,958, 66,625. E1 (13,325) and 19,986 were not archived and are lost
e5-x4rand-mos-r2 continuation of the first x4rand run from its E1 weights, Adam reset (stopped, no reference value) hf at 13,325, 19,986, 26,650
e5-x8rand-mos K8 random-init experts, expert LR 6e-4 hf at 46,634, 53,300, 59,958, 66,625
e5-x16rand-mos K16 random-init experts, expert LR 6e-4, expert parallelism hf at every save 6,662 … 66,625
e5-x16lr4-mos K16 warm-start jitter init, expert LR ×4, expert parallelism (rerun paused at step 18,478) hf at 6,662 and 13,325; full checkpoint (with optimizer) at 13,325 = resume point
e5-x4rand-lr1e3-mos ablation: x4rand with expert LR 1e-3, trained to E1 on the 5-epoch schedule hf at 6,662; full checkpoint (with optimizer) at 13,325
Input Content
inputs/d0-shared-random-experts-8b-x4 x4rand init: D0 shared MLP/attention + 4 random experts (seed 20260917)
inputs/d0-shared-random-experts-8b-x4-half half-width ablation init: the x4 init with each expert cut to its first 6,144 hidden units
inputs/d0-shared-random-experts-8b-x16 x16rand init: D0 + 16 random experts (seed 20260919)
inputs/d0-fullwidth-init-8b-jitter K4 warm-start init (experts = D0 MLP + 1% jitter); also the source from which the random-expert inits are rebuilt
inputs/d0-fullwidth-init-8b-x16-jitter K16 warm-start init used by e5-x16lr4-mos

Results, configs and logs: GitHub MoS repository, experiments/aurora/qwen3-8b-800k-e5/bal01/README.md (§10). Archive records (file counts and byte totals, verified against the local copies before deletion): experiments/aurora/qwen3-8b-800k-e5/archive/.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support