Qwen3-8B DFlash speculator (matched baseline)
DFlash draft model, trained identically to the DSpark checkpoint as an A/B baseline, draft model for Qwen/Qwen3-8B, trained with
speculators (PR #677).
Preview/research checkpoint: online training on a 20k ShareGPT slice, seq 2048, 3 epochs.
Validation per-position acceptance (pos1-7): 0.665 / 0.474 / 0.338 / 0.243 / 0.180 / 0.141 / 0.115 | full_acc 0.309.
- Downloads last month
- 18