YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
MoS-27B-Checkpoints
Epoch checkpoints of the Qwen3.8-27B DFlash2 MoS line (project MoS, branch ryan/mos-improve,
experiments/dflash2/qwen3.8-27b-100k/). Started 2026-09-22, when the general archive repo
ryan-0608/MoS-Aurora-Experiment-Archive reached the Hugging Face 20,000-file limit.
Layout
dflash2_27b_traj_20260915/main-800k/<arm>/epoch<N>/
model.safetensors drafter weights (speculators DFlash2MoSDraftModel)
optimizer_state_dict.pt full AdamW state -> exact resume
scheduler_state_dict.pt, training_state.json, config.json, config.py, train_command.txt
val_metrics.json trainer held-out 10% validation at the end of the epoch
serving/ fixed-500 serving result for this checkpoint
acceptance_summary.json pooled AL = sum completion tokens / sum verify calls
acceptance_trace.jsonl per-request counts
client.log, server.log
Since 2026-09-30 this repo holds the only copies of the main-800k epoch weights (the PVC copies of x4rand E3/E4 and
dense4 E1/E3/E4 were deleted after presence + byte-size checks here).
Also in this repo (2026-09-30, PVC cleanup; the PVC copies of the large files were deleted after presence + byte-size checks here):
dflash2_27b_traj_20260915/inputs/<init>/ config.json, model.safetensors, build-*.log
dflash2_27b_traj_20260915/data/current-800k-self27b-traj/
train.jsonl cleaned 800K self-distilled corpus, 799,395 rows (8,926,961,996 B)
traj.jsonl the same rows as pre-built input_ids / loss_mask (19,323,325,279 B); the 800K runs were
prepared from it (speculators prepare_data, 80 Arrow shards)
manifest.json, public-overlap-27b-800k.json, evidence-archived-20260915/
| init | what | used by |
|---|---|---|
dflash2-27b-zlab-x8rand-b16 |
z-lab import + K=8 random experts (gate/up N(0,0.02), down N(0,1e-3), seed 20260917) | 300-step K8 memory/speed probe 2026-09-29; the K8 4-epoch run was dropped |
dflash2-27b-zlab-x4rand-sel1024-b16 |
x4rand init with the candidate selector widened 256 -> 1024 (new successor columns 0, so step-0 scores are unchanged; seed 20260929) | selector-rank smoke 2026-09-29/30: path AL +0.09% vs x4rand at steps 13k-18k, not pursued |
Arms
| arm (path) | Ryan's name | recipe |
|---|---|---|
dense4 |
27B dense (4-epoch baseline) | dense drafter warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4; cosine over 4 epochs, warmup 0.005; same corpus, 3 verifier + 5 trainer layout, global batch and 73,670 steps per epoch as x4rand; finished 2026-09-29 |
x4rand |
27B 4expert rand LR6 | K=4 full-width expert MLPs per draft layer, random init (gate/up N(0,0.02), down N(0,1e-3)); shared MLP, attention and the rest warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4, expert LR 6e-4, router LR 5e-4; top-2, tau 0.9, balance 0.01; cosine over 4 epochs, warmup 0.005; corpus Current-800K-Self-27B-Traj (719,455 train rows); 73,670 steps per epoch |
Results (target Qwen/Qwen3.8-27B @1d4bf0f2)
Serving protocol: sglang main f5866545, TP2, DFLASH 16 draft tokens, fixed-500 prompts, max 128 new tokens, greedy, concurrency 1, 500/500 completed. Same protocol for every arm.
| arm | epoch | global_step | train-side val AL | serving pooled AL | vs dense4 same epoch |
|---|---|---|---|---|---|
| dense4 | 1 | 73,670 | 5.0262 | 5.0571 | โ |
| dense4 | 2 | 147,335 | 5.0261 | 5.0510 | โ |
| dense4 | 3 | 221,001 | 5.026 (from train.log) | 5.0558 | โ |
| dense4 | 4 | 294,680 | 5.024 (from train.log) | 5.0538 | โ |
| x4rand | 1 | 73,670 | 5.1485 | 5.2314 | +3.45% (dense4 5.0571) |
| x4rand | 2 | 147,335 | 5.2056 | 5.2922 | +4.78% (dense4 5.0510) |
| x4rand | 3 | 221,001 | 5.2202 | 5.3261 | +5.34% (dense4 5.0558) |
| x4rand | 4 | 294,682 | 5.2201 | 5.3113 | +5.09% (dense4 5.0538); best epoch = 3 |
The K4 jitter arm (shared MLP copies + 1% jitter, expert LR 1e-4; E1 only, serving 5.1501, +1.84% vs dense4 E1)
is model-only in ryan-0608/MoS-Aurora-Experiment-Archive under dflash2_27b_traj_20260915/main-800k/mos/epoch1.
Public sets (max tokens 512, c=1, TP2, pooled AL)
x4rand E3 against dense4 E1 (each arm's best fixed-500 @128 epoch); results and traces under
main-800k/{x4rand/epoch3,dense4/epoch1}/serving-public-512/<set>/.
| set (rows) | dense4 E1 | x4rand E3 | x4rand vs dense |
|---|---|---|---|
| gsm8k (1,319) | 6.5957 | 6.8636 | +4.06% |
| humaneval (164) | 5.1535 | 5.2613 | +2.09% |
| math500 (500) | 6.0312 | 6.1622 | +2.17% |
| mbpp_full (500) | 5.6934 | 5.8311 | +2.42% |
| mt_bench_turn1 (80) | 4.0139 | 4.0908 | +1.91% |
Use
Resume training: pass the epoch directory as the checkpoint dir to speculators train.py (same train_command.txt,
--epochs 4). Serve: export_d2_for_sglang.py <epoch dir> <out> then sglang --speculative-algorithm DFLASH --speculative-draft-model-path <out>; the MoS architecture needs apply_serving_dflash2_mos.py on sglang main.