YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

MoS-27B-Checkpoints

Epoch checkpoints of the Qwen3.8-27B DFlash2 MoS line (project MoS, branch ryan/mos-improve, experiments/dflash2/qwen3.8-27b-100k/). Started 2026-09-22, when the general archive repo ryan-0608/MoS-Aurora-Experiment-Archive reached the Hugging Face 20,000-file limit.

Layout

dflash2_27b_traj_20260915/main-800k/<arm>/epoch<N>/
    model.safetensors            drafter weights (speculators DFlash2MoSDraftModel)
    optimizer_state_dict.pt      full AdamW state -> exact resume
    scheduler_state_dict.pt, training_state.json, config.json, config.py, train_command.txt
    val_metrics.json             trainer held-out 10% validation at the end of the epoch
    serving/                     fixed-500 serving result for this checkpoint
        acceptance_summary.json  pooled AL = sum completion tokens / sum verify calls
        acceptance_trace.jsonl   per-request counts
        client.log, server.log

Since 2026-09-30 this repo holds the only copies of the main-800k epoch weights (the PVC copies of x4rand E3/E4 and dense4 E1/E3/E4 were deleted after presence + byte-size checks here).

Also in this repo (2026-09-30, PVC cleanup; the PVC copies of the large files were deleted after presence + byte-size checks here):

dflash2_27b_traj_20260915/inputs/<init>/          config.json, model.safetensors, build-*.log
dflash2_27b_traj_20260915/data/current-800k-self27b-traj/
    train.jsonl        cleaned 800K self-distilled corpus, 799,395 rows (8,926,961,996 B)
    traj.jsonl         the same rows as pre-built input_ids / loss_mask (19,323,325,279 B); the 800K runs were
                       prepared from it (speculators prepare_data, 80 Arrow shards)
    manifest.json, public-overlap-27b-800k.json, evidence-archived-20260915/
init what used by
dflash2-27b-zlab-x8rand-b16 z-lab import + K=8 random experts (gate/up N(0,0.02), down N(0,1e-3), seed 20260917) 300-step K8 memory/speed probe 2026-09-29; the K8 4-epoch run was dropped
dflash2-27b-zlab-x4rand-sel1024-b16 x4rand init with the candidate selector widened 256 -> 1024 (new successor columns 0, so step-0 scores are unchanged; seed 20260929) selector-rank smoke 2026-09-29/30: path AL +0.09% vs x4rand at steps 13k-18k, not pursued

Arms

arm (path) Ryan's name recipe
dense4 27B dense (4-epoch baseline) dense drafter warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4; cosine over 4 epochs, warmup 0.005; same corpus, 3 verifier + 5 trainer layout, global batch and 73,670 steps per epoch as x4rand; finished 2026-09-29
x4rand 27B 4expert rand LR6 K=4 full-width expert MLPs per draft layer, random init (gate/up N(0,0.02), down N(0,1e-3)); shared MLP, attention and the rest warm-started from z-lab/Qwen3.8-27B-DFlash2 @50307d4c; base LR 1e-4, expert LR 6e-4, router LR 5e-4; top-2, tau 0.9, balance 0.01; cosine over 4 epochs, warmup 0.005; corpus Current-800K-Self-27B-Traj (719,455 train rows); 73,670 steps per epoch

Results (target Qwen/Qwen3.8-27B @1d4bf0f2)

Serving protocol: sglang main f5866545, TP2, DFLASH 16 draft tokens, fixed-500 prompts, max 128 new tokens, greedy, concurrency 1, 500/500 completed. Same protocol for every arm.

arm epoch global_step train-side val AL serving pooled AL vs dense4 same epoch
dense4 1 73,670 5.0262 5.0571 โ€”
dense4 2 147,335 5.0261 5.0510 โ€”
dense4 3 221,001 5.026 (from train.log) 5.0558 โ€”
dense4 4 294,680 5.024 (from train.log) 5.0538 โ€”
x4rand 1 73,670 5.1485 5.2314 +3.45% (dense4 5.0571)
x4rand 2 147,335 5.2056 5.2922 +4.78% (dense4 5.0510)
x4rand 3 221,001 5.2202 5.3261 +5.34% (dense4 5.0558)
x4rand 4 294,682 5.2201 5.3113 +5.09% (dense4 5.0538); best epoch = 3

The K4 jitter arm (shared MLP copies + 1% jitter, expert LR 1e-4; E1 only, serving 5.1501, +1.84% vs dense4 E1) is model-only in ryan-0608/MoS-Aurora-Experiment-Archive under dflash2_27b_traj_20260915/main-800k/mos/epoch1.

Public sets (max tokens 512, c=1, TP2, pooled AL)

x4rand E3 against dense4 E1 (each arm's best fixed-500 @128 epoch); results and traces under main-800k/{x4rand/epoch3,dense4/epoch1}/serving-public-512/<set>/.

set (rows) dense4 E1 x4rand E3 x4rand vs dense
gsm8k (1,319) 6.5957 6.8636 +4.06%
humaneval (164) 5.1535 5.2613 +2.09%
math500 (500) 6.0312 6.1622 +2.17%
mbpp_full (500) 5.6934 5.8311 +2.42%
mt_bench_turn1 (80) 4.0139 4.0908 +1.91%

Use

Resume training: pass the epoch directory as the checkpoint dir to speculators train.py (same train_command.txt, --epochs 4). Serve: export_d2_for_sglang.py <epoch dir> <out> then sglang --speculative-algorithm DFLASH --speculative-draft-model-path <out>; the MoS architecture needs apply_serving_dflash2_mos.py on sglang main.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support