Papers
arxiv:2609.34117

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

Published on Sep 28
· Submitted by
Gunho Park
on Oct 7
Authors:
,
,
,
,

Abstract

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned model that directly reuses the prefill-generated KV cache without conversion. Across two MoE backbones and three pruning criteria, this training-free KV cache handoff substantially narrows accuracy gaps relative to the full model in many settings. We also show that benchmark accuracy can conceal substantial pruning-induced changes in generation length. To address these distortions and residual accuracy loss, SlimWise introduces a low-cost distillation stage that trains the decoder to continue from full-model KV caches while updating only a small subset of parameters. Implemented in vLLM, SlimWise supports both prefill-decode (PD) disaggregation and PD-colocated serving. On Qwen3.6-35B-A3B, SlimWise improves decode throughput by up to 1.81x at 50% expert pruning with minimal accuracy loss.

Community

Paper author Paper submitter

Hi HF community! I'm one of the authors of SlimWise, which speeds up MoE serving by pruning experts only during decode.

SlimWise runs prefill with the full model and decode with a pruned model that reuses the full-model KV cache. This training-free handoff recovers much of the accuracy lost to pruning, and a lightweight distillation stage closes the gap further. In vLLM, it achieves up to 1.81× decode throughput at 50% expert pruning on Qwen3.6-35B-A3B with minimal accuracy loss.

Happy to answer questions!

·

What's your favorite flavor ice cream?

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.34117
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.34117 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.34117 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.34117 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.