ctest: Qwen3-8B DFlash speculator โ€” FP8 hidden-states ablation (bf16)

Test checkpoint โ€” not a production release. Trained to validate the FP8-quantized hidden-states transfer backend added in vllm-project/speculators#1028 (draft, reviving #491).

  • Verifier: Qwen/Qwen3-8B
  • Algorithm: dflash
  • Hidden-states precision during training: bf16 (existing bf16 file backend)
  • Training data: 5K magpie + 5K ultrachat samples from inference-optimization/Dataset-Qwen3-235B-Instruct

This is one of a matched pair (...-fp8ablation-bf16 / ...-fp8ablation-fp8) trained identically except for the hidden-states transfer precision, to isolate the effect of FP8 quantization on speculator quality. See the sibling repo for the other precision.

Validation metrics (final epoch)

  • loss_epoch: 0.670712
  • full_acc_epoch: 0.190457
  • position_1_acc_epoch: 0.671505
  • position_2_acc_epoch: 0.470419
  • position_3_acc_epoch: 0.345403
  • position_4_acc_epoch: 0.259921
  • position_5_acc_epoch: 0.199261
  • position_6_acc_epoch: 0.156393
  • position_7_acc_epoch: 0.125580
  • position_8_acc_epoch: 0.105085
  • position_9_acc_epoch: 0.090258
  • position_10_acc_epoch: 0.081353
  • position_11_acc_epoch: 0.075329
  • position_12_acc_epoch: 0.070879
  • position_13_acc_epoch: 0.068388
  • position_14_acc_epoch: 0.065941
  • position_15_acc_epoch: 0.063810
  • eal_epoch: 1.165076

guidellm serving eval

See acceptance.csv in this repo for the full per-subset guidellm breakdown (9 subset rows).

Full ablation writeup, all 6 checkpoints' results, exact commands, and code: see the results package referenced from PR #1028.

Downloads last month
10
Safetensors
Model size
2B params
Tensor type
I64
ยท
BF16
ยท
BOOL
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support