voipmonitor's picture
Publish MXFP8 conversion of DFlash2 bf582e4
cc006ae verified
|
Raw
History Blame Contribute Delete
4.11 kB
metadata
license: cc-by-nc-nd-4.0
library_name: vllm
pipeline_tag: text-generation
base_model:
  - incoai/GLM-5.3-Flash-DFlash2
  - local-inference-lab/GLM-5.3-Flash-NVFP4
inference: false
tags:
  - dflash
  - dflash2
  - speculative-decoding
  - draft-model
  - modelopt
  - mxfp8
  - vllm

GLM-5.3-Flash-DFlash2

This repository contains an offline MXFP8 conversion of the DFlash 2 draft model for local-inference-lab/GLM-5.3-Flash-NVFP4. It is not a standalone language model. A compatible speculative-decoding server loads it beside the target model and verifies every drafted token against the target.

The source checkpoint is incoai/GLM-5.3-Flash-DFlash2 at immutable revision bf582e4eacc1810f76656d1811693ff6c6737d2a.

Serialized format

  • Linear weights: float8_e4m3fn
  • Scale values: biased E8M0 exponents stored as uint8
  • Quantization block: 1×32 values
  • Scale layout: row-major and unswizzled
  • Excluded module: lm_head
  • Draft KV cache quantization: not encoded in the checkpoint

conversion_manifest.json records the immutable source revision, source and output checksums, tensor coverage, aggregate quantization error, and per-weight validation statistics.

Validation status

Status: qualified for checkpoint structure, exact format reproduction, loading, CUDA-graph capture, and speculative inference under these conditions:

  • Target: local-inference-lab/GLM-5.3-Flash-NVFP4
  • Runtime: voipmonitor/vllm:jovian-judgement-community-20260901-r11
  • Hardware: four NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs
  • Parallelism: tensor parallel size 4 and decode-context parallel size 1
  • Target attention, MoE, linear, and tensor-parallel all-reduce: B12X
  • DFlash attention: FlashAttention 2
  • DFlash linear: B12X MXFP8
  • DFlash proposal length: seven tokens
  • DFlash KV cache: auto (BF16)
  • CUDA graph mode: full decode graphs for the target and DFlash drafter

The qualification re-quantized all 47 eligible weights and verified exact serialized weight and scale equality. All 34 preserved tensors remained bit-identical to the BF16 source. The runtime detected ModelOpt MXFP8, selected B12xMxfp8LinearKernel for draft GEMMs and the fused DFlash context K/V projection, captured target and draft decode graphs, and completed speculative inference.

The checkpoint is unsupported in vLLM builds without the DFlash 2 and ModelOpt MXFP8 integrations used by the qualified runtime.

Serving

docker run --rm \
  --gpus '"device=0,1,2,3"' \
  --network host \
  --ipc host \
  --shm-size 32g \
  -e MODEL=local-inference-lab/GLM-5.3-Flash-NVFP4 \
  -e SERVED_MODEL_NAME=GLM-5.3-Flash-NVFP4 \
  -e PORT=8000 \
  -e TP=4 \
  -e DCP=1 \
  -e MAX_NUM_SEQS=16 \
  -e MAX_MODEL_LEN=262144 \
  -e MAX_NUM_BATCHED_TOKENS=4096 \
  -e SPECULATOR=dflash2 \
  -e NUM_SPECULATIVE_TOKENS=7 \
  -e DFLASH_MODEL=local-inference-lab/GLM-5.3-Flash-DFlash2 \
  -e DFLASH_KV_CACHE_DTYPE=auto \
  -e DFLASH_ATTENTION_BACKEND=FLASH_ATTN \
  -e ATTENTION_BACKEND=B12X \
  -e MOE_BACKEND=b12x \
  -e LINEAR_BACKEND=b12x \
  -e B12X_PCIE_ALLREDUCE=1 \
  -e CUDAGRAPH_MODE=FULL \
  -e VLLM_B12X_MOE_FP4_FORCE_A16=0 \
  voipmonitor/vllm:jovian-judgement-community-20260901-r11

The repository name resolves the main branch, including subsequent model updates. Pin MODEL_REVISION and DFLASH_MODEL_REVISION to immutable Hugging Face commit hashes when deployment reproducibility is required. The OpenAI-compatible endpoint is available at http://127.0.0.1:8000/v1.

License and attribution

The source DFlash 2 model is distributed under CC BY-NC-ND 4.0. See the source model card for its use restrictions and attribution information.

@misc{inco2026dflash2,
  title  = {{DFlash 2: Keep Drafting Parallel}},
  author = {{Inco AI}},
  year   = {2026},
  month  = {August},
  url    = {https://inco.ai/blog/dflash2/}
}