GLM-5.3-Flash-DFlash2
This repository contains an MXFP8-quantized DFlash 2 draft model for
local-inference-lab/GLM-5.3-Flash-NVFP4.
It is not a standalone language model. A compatible speculative-decoding
server loads it beside the target model and verifies every drafted token
against the target.
The source checkpoint is
incoai/GLM-5.3-Flash-DFlash2
at immutable revision dc77ff1c99eeb2df044ee3d4f0094eb033fee410.
Format
- Linear weights:
float8_e4m3fn - Scale values: biased E8M0 exponents stored as
uint8 - Quantization block: 1×32 values
- Scale layout: row-major and unswizzled
- Excluded module:
lm_head - Draft KV cache quantization: not encoded in the checkpoint
conversion_manifest.json records the immutable source revision, source and
output checksums, tensor coverage, aggregate quantization error, and
per-weight validation statistics.
Validation status
Status: qualified for checkpoint structure, exact format reproduction, loading, and smoke inference under the following conditions:
- Target:
local-inference-lab/GLM-5.3-Flash-NVFP4revision520de24eabf507659eaef7c70f14fd584527facc - Runtime:
voipmonitor/vllm@sha256:ef53437759e3a41d5ee1c4e9045ffdd7df2972faad50d1dc687e3ab479c5867a - Hardware: four NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
- Parallelism: tensor parallel size 4 and decode-context parallel size 1
- Target attention, MoE, linear, and tensor-parallel all-reduce: B12X
- DFlash attention: FlashAttention 2
- DFlash linear: B12X MXFP8
- DFlash proposal length: seven tokens
- DFlash KV cache:
auto(BF16) - CUDA graph mode:
FULLrequested; target and DFlash2 decode are captured, while target GDN prefill remains eager
The runtime detected ModelOpt MXFP8, selected B12xMxfp8LinearKernel for
draft GEMMs and the fused DFlash context K/V projection, and loaded 1.20 GB of
draft weights. With seven draft tokens, the qualified runtime measured a
2.1157-second median time to first token for a 32,320-token prompt and
185.5 output tokens per second at concurrency one. Speculative throughput
depends on prompt content and acceptance length.
The checkpoint is unsupported in vLLM builds that do not contain the DFlash 2 and ModelOpt MXFP8 integration used by the qualified runtime.
Serving
docker run --rm \
--gpus '"device=0,1,2,3"' \
--network host \
--ipc host \
--shm-size 32g \
-e MODEL=local-inference-lab/GLM-5.3-Flash-NVFP4 \
-e SERVED_MODEL_NAME=GLM-5.3-Flash-NVFP4 \
-e PORT=8000 \
-e TP=4 \
-e DCP=1 \
-e MAX_NUM_SEQS=16 \
-e MAX_MODEL_LEN=262144 \
-e MAX_NUM_BATCHED_TOKENS=4096 \
-e SPECULATOR=dflash \
-e NUM_SPECULATIVE_TOKENS=7 \
-e DFLASH_MODEL=local-inference-lab/GLM-5.3-Flash-DFlash2 \
-e DFLASH_MODEL_REVISION= \
-e DFLASH_KV_CACHE_DTYPE=auto \
-e DFLASH_ATTENTION_BACKEND=FLASH_ATTN \
-e ATTENTION_BACKEND=B12X \
-e MOE_BACKEND=b12x \
-e LINEAR_BACKEND=b12x \
-e B12X_PCIE_ALLREDUCE=1 \
-e CUDAGRAPH_MODE=FULL \
-e VLLM_B12X_MOE_FP4_FORCE_A16=0 \
voipmonitor/vllm:jovian-judgement-community-dflash2-20260830-r7
An empty DFLASH_MODEL_REVISION makes the launcher resolve the repository's
main branch. For reproducible deployments, replace the empty value with an
immutable Hugging Face commit hash. The OpenAI-compatible endpoint is
available at http://127.0.0.1:8000/v1.
License and attribution
The source DFlash 2 model is distributed under CC BY-NC-ND 4.0. See the source model card for its use restrictions and attribution information.
@misc{inco2026dflash2,
title = {{DFlash 2: Keep Drafting Parallel}},
author = {{Inco AI}},
year = {2026},
month = {August},
url = {https://inco.ai/blog/dflash2/}
}
- Downloads last month
- 125