wfangtw's picture
Upload README.md
5e36a06 verified
|
Raw
History Blame Contribute Delete
1.01 kB
metadata
library_name: transformers
tags:
  - dflash
  - speculative-decoding
  - sglang
  - glm-5.2

GLM-5.2 FP8 DFLASH v1

Overview

This is a DFLASH speculative draft model for GLM-5.2 FP8 serving. The checkpoint uses DFLASH block size 16 and is intended to be loaded as the draft model in SGLang speculative decoding.

Training data: SubconsciousDev/Subconscious-Dflash-Training-Dataset-short-glm52-v1.

SGLang Usage

Add these arguments to the SGLang launch command:

--speculative-algorithm DFLASH \
    --speculative-draft-model-path SubconsciousDev/glm-5.2-fp8-dflash-v1 \
    --speculative-num-draft-tokens 16 \

Evaluation

Concurrency is 1.

Benchmark EAGLE (5,1,6) Speedup EAGLE (5,1,6) Accept Length DFLASH Speedup DFLASH Accept Length
GSM8K 2.2x 3.99 2.3x 4.87
MATH500 3.4x 4.92 4.3x 7.44
MBPP 3.5x 5.21 4.9x 9.03
LiveCodeBench 2.9x 4.17 3.1x 5.35
MT-Bench 2.5x 3.71 2.4x 3.96