GLM-5.3-Flash-DFlash2

This repository contains an MXFP8-quantized DFlash 2 draft model for local-inference-lab/GLM-5.3-Flash-NVFP4. It is not a standalone language model. A compatible speculative-decoding server loads it beside the target model and verifies every drafted token against the target.

The source checkpoint is incoai/GLM-5.3-Flash-DFlash2 at immutable revision dc77ff1c99eeb2df044ee3d4f0094eb033fee410.

Format

  • Linear weights: float8_e4m3fn
  • Scale values: biased E8M0 exponents stored as uint8
  • Quantization block: 1×32 values
  • Scale layout: row-major and unswizzled
  • Excluded module: lm_head
  • Draft KV cache quantization: not encoded in the checkpoint

conversion_manifest.json records the immutable source revision, source and output checksums, tensor coverage, aggregate quantization error, and per-weight validation statistics.

Validation status

Status: qualified for checkpoint structure, exact format reproduction, loading, and smoke inference under the following conditions:

  • Target: local-inference-lab/GLM-5.3-Flash-NVFP4 revision 520de24eabf507659eaef7c70f14fd584527facc
  • Runtime: voipmonitor/vllm@sha256:ef53437759e3a41d5ee1c4e9045ffdd7df2972faad50d1dc687e3ab479c5867a
  • Hardware: four NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
  • Parallelism: tensor parallel size 4 and decode-context parallel size 1
  • Target attention, MoE, linear, and tensor-parallel all-reduce: B12X
  • DFlash attention: FlashAttention 2
  • DFlash linear: B12X MXFP8
  • DFlash proposal length: seven tokens
  • DFlash KV cache: auto (BF16)
  • CUDA graph mode: FULL requested; target and DFlash2 decode are captured, while target GDN prefill remains eager

The runtime detected ModelOpt MXFP8, selected B12xMxfp8LinearKernel for draft GEMMs and the fused DFlash context K/V projection, and loaded 1.20 GB of draft weights. With seven draft tokens, the qualified runtime measured a 2.1157-second median time to first token for a 32,320-token prompt and 185.5 output tokens per second at concurrency one. Speculative throughput depends on prompt content and acceptance length.

The checkpoint is unsupported in vLLM builds that do not contain the DFlash 2 and ModelOpt MXFP8 integration used by the qualified runtime.

Serving

docker run --rm \
  --gpus '"device=0,1,2,3"' \
  --network host \
  --ipc host \
  --shm-size 32g \
  -e MODEL=local-inference-lab/GLM-5.3-Flash-NVFP4 \
  -e SERVED_MODEL_NAME=GLM-5.3-Flash-NVFP4 \
  -e PORT=8000 \
  -e TP=4 \
  -e DCP=1 \
  -e MAX_NUM_SEQS=16 \
  -e MAX_MODEL_LEN=262144 \
  -e MAX_NUM_BATCHED_TOKENS=4096 \
  -e SPECULATOR=dflash \
  -e NUM_SPECULATIVE_TOKENS=7 \
  -e DFLASH_MODEL=local-inference-lab/GLM-5.3-Flash-DFlash2 \
  -e DFLASH_MODEL_REVISION= \
  -e DFLASH_KV_CACHE_DTYPE=auto \
  -e DFLASH_ATTENTION_BACKEND=FLASH_ATTN \
  -e ATTENTION_BACKEND=B12X \
  -e MOE_BACKEND=b12x \
  -e LINEAR_BACKEND=b12x \
  -e B12X_PCIE_ALLREDUCE=1 \
  -e CUDAGRAPH_MODE=FULL \
  -e VLLM_B12X_MOE_FP4_FORCE_A16=0 \
  voipmonitor/vllm:jovian-judgement-community-dflash2-20260830-r7

An empty DFLASH_MODEL_REVISION makes the launcher resolve the repository's main branch. For reproducible deployments, replace the empty value with an immutable Hugging Face commit hash. The OpenAI-compatible endpoint is available at http://127.0.0.1:8000/v1.

License and attribution

The source DFlash 2 model is distributed under CC BY-NC-ND 4.0. See the source model card for its use restrictions and attribution information.

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}
Downloads last month
125
Safetensors
Model size
1B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for local-inference-lab/GLM-5.3-Flash-DFlash2

Quantized
(4)
this model