Disaggregated Quantization: Specializing LLM Prefill and Decode
Abstract
Prefill and decode reward different approaches to quantization: low-precision arithmetic accelerates prompt processing, while compact weights reduce memory traffic during generation. We propose "disaggregated quantization" (DQ), which specializes computation formats, weights and storage placement to both of these phases. On Qwen 3 and Gemma 3, removing activation quantization specifically on decode improves accuracy on decode-heavy tasks without increasing inference cost. Training separate compute-native prefill weights accelerates prompt processing relative to weight-only inference while matching or exceeding its accuracy at 2-3-bit decode on both decode-heavy and prefill-heavy tasks. With released Qwen3.8-27B GGUF decoders, training an NVFP4 prefiller improves 1-bit accuracy by 32.5 points on MMLU-Pro and 35.3 on MMMU-Pro without modifying the decode checkpoint. To accommodate the additional checkpoint on a single device, offloaded disaggregated prefill (ODP) streams its weights from SSD, amortizing loading over prompt length. On the same 27B model, ODP delivers a 1.78x time-to-first-token speedup over the weight-only baseline at 8K prompt length in llama.cpp. We evaluate accuracy under disaggregated serving in vLLM and further validate shared-weight format disaggregation through post-training quantization on models up to 2.8T parameters.
Community
Disaggregated quantization for LLMs
LLM prefill and decode can be processed by separate weights in different precisions. Boosts speed AND accuracy. Offloading makes it zero-overhead on long sequences for dense models. Applicable to existing quantized checkpoints.
Paper: https://arxiv.org/abs/2609.26333
Code: https://github.com/IST-DASLab/disaggregated-quantization
Qwen3.8 NVFP4 prefillers: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-NVFP4-prefiller
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction (2026)
- Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation (2026)
- Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation (2026)
- Train Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning (2026)
- Budget-Aware Compression Pipeline for Single-GPU LLM Inference: Methods, Trade-offs, and Coupling Effects (2026)
- DASC: Decay-Aware State Compression for Hybrid Linear-Attention Serving (2026)
- Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.26333 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper