--- pipeline_tag: text-generation library_name: transformers base_model: - thinkingmachines/Inkling-NVFP4 license: apache-2.0 inference: false tags: - dflash - speculative-decoding - speculative-decoding-draft - block-diffusion - draft-model - diffusion-language-model - efficiency - inkling - modelopt - nvfp4 - mxfp8 - fp8 - sglang --- # Inkling-NVFP4-DFlash [Paper](https://arxiv.org/abs/2602.06036) | [Github](https://github.com/z-lab/dflash) | [Blog](https://z-lab.ai/projects/dflash) This repository contains a DFlash draft model for [`thinkingmachines/Inkling-NVFP4`](https://huggingface.co/thinkingmachines/Inkling-NVFP4). It is not a standalone language model. It is intended to be paired with the target model in a speculative decoding server. This is an early preview release, and the drafter is still training. It uses all causal sliding-window attention (SWA) layers. DFlash uses a lightweight block diffusion draft model to propose multiple tokens in parallel. The target model verifies those proposals, improving serving throughput while preserving the target model's output distribution. ## Quick Start Here is an example deployment. Some options, including `trtllm_mha` for draft attention, are pending upstream SGLang support. ```bash sglang serve \ --model-path thinkingmachines/Inkling-NVFP4 \ --tp 8 \ --trust-remote-code \ --quantization modelopt_fp4 \ --fp4-gemm-backend flashinfer_trtllm \ --moe-runner-backend flashinfer_trtllm_routed \ --enable-torch-symm-mem \ --attention-backend fa4 \ --mamba-radix-cache-strategy extra_buffer \ --disable-custom-all-reduce \ --page-size 128 \ --reasoning-parser inkling \ --tool-call-parser inkling \ --enable-multimodal \ --cuda-graph-backend-prefill breakable \ --enable-scattered-sconv \ --mem-fraction-static 0.78 \ --max-running-requests 32 \ --swa-full-tokens-ratio 0.10 \ --mamba-full-memory-ratio 0.10 \ --chunked-prefill-size 16384 \ --watchdog-timeout 900 \ --weight-loader-prefetch-checkpoints \ --weight-loader-prefetch-num-threads 8 \ --cuda-graph-max-bs-decode 32 \ --cuda-graph-max-bs-prefill 4096 \ --cuda-graph-bs-prefill 128 256 384 512 768 1024 1536 2048 2560 3072 3584 4096 \ --kv-cache-dtype mxfp8 \ --speculative-algorithm DFLASH \ --speculative-draft-model-path modal-labs/Inkling-DFlash \ --speculative-dflash-block-size 16 \ --speculative-draft-model-quantization fp8 \ --speculative-draft-attention-backend trtllm_mha \ --speculative-draft-kv-cache-dtype fp8_e4m3 \ --speculative-draft-window-size 4096 \ --host 0.0.0.0 \ --port 30000 ``` ## Benchmark Results This preview includes a preliminary accept-length evaluation at concurrency 1 and block size 16. A full benchmark suite, including throughput and higher-concurrency measurements, will follow. ### Setup - Runtime: SGLang on 8x NVIDIA B300 GPUs, tensor parallel size 8 - Target model: `thinkingmachines/Inkling-NVFP4` - Target weights and KV cache: ModelOpt NVFP4 weights with an `mxfp8` KV cache - Backends: `fa4` target attention and `trtllm_mha` DFlash draft attention - Draft model: FP8 weights with an `fp8_e4m3` KV cache and a 4096-token draft window - Workloads: GSM8K, MATH500, HumanEval, MBPP, and MT-Bench with Inkling's native chat renderer - Decoding: greedy, reasoning effort 0.9, max output length 2048 tokens - Measurement: DFlash block size 16 at concurrency 1, with up to 64 measured generation requests per workload; warmup and warmdown requests are excluded - Accept length: `completion_tokens / spec_verify_ct` per generation turn, averaged across generation turns ### Accept Length Mean DFlash accept length at concurrency 1. | Workload | DFlash block=16 | | --- | --- | | GSM8K | 4.562 | | MATH500 | 4.712 | | HumanEval | 4.959 | | MBPP | 3.907 | | MT-Bench | 2.914 | ## Citation If you find DFlash useful, please cite the original paper: ```bibtex @article{chen2026dflash, title = {{DFlash: Block Diffusion for Flash Speculative Decoding}}, author = {Chen, Jian and Liang, Yesheng and Liu, Zhijian}, journal = {arXiv preprint arXiv:2602.06036}, year = {2026} } ```