Kimi-K3-W4AFP8

INT4 (group 128) routed experts with FP8 activations, FP8 attention projections, requantized from the native MXFP4 weights of moonshotai/Kimi-K3. Same tensor format and kernels as vessl/Kimi-K3-W4AFP8, with lower error against the native model.

Two configurations on 32 x GH200

8 nodes x 4 GH200, TP32 / EP32, DSpark speculative decoding (block 2), FP8 KV cache. Real AIME/HMMT prompts with real EOS; numbers are median output tok/s per stream.

W4A16 — moonshotai/Kimi-K3 (native MXFP4) W4A8 — this repo
tok/s per stream, 50 concurrent 29.6 34.1
tok/s per stream, 20 concurrent 40.6 43.6
aggregate tok/s, 50 concurrent 883 949
KL to native, nats/token (95% CI) +0.0011 (0.0006–0.0016) +0.0154 (0.0142–0.0167)
AIME 2026 + HMMT Feb 2026, 8 samples/problem 82.3% 81.3% (−1.0 pt, SE 1.6: not significant)
  • Patches in the measured runs: W4A16 used patches 02 and 04 (no --max-total-tokens cap). W4A8 used 01, 02, 03 and 05, i.e. exactly the launch command below.
  • Pick W4A16 when you want the reference model.
  • Pick W4A8 for about 15% more throughput at 50 streams.
  • Scaling out: two W4A8 replicas (64 GPUs) behind sglang-router, 25 streams each, give 35.7 tok/s per stream.
  • KL: teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.

Setup

  1. Image: SGLang v0.5.20, arm64, CUDA 13.
  2. Clone the repo: git clone https://huggingface.co/Radioheading/Kimi-K3-W4AFP8. You need its apply_patch.py and patches/ below.
  3. Patches: apply the ones listed below inside the image's source tree (cd /sgl-workspace/sglang && git apply <patch>).
  4. W4A8 quantization method: for W4A8 only, run python3 apply_patch.py from this repo. It registers the method (from vessl/Kimi-K3-W4AFP8).
  5. Draft model: download RadixArk/Kimi-K3-DSpark.
patch needed for effect
01-sgl-kernel-sm90a.patch W4A8 (required) Builds the SM90 CUTLASS kernels with sm_90a. aarch64 wheels lack it, and W4A8/FP8 kernels abort with "Arch conditional MMA instruction". Shortcut: copy the prebuilt sgl_kernel_sm90a/common_ops.abi3.so (arm64, CUDA 13, v0.5.20) over site-packages/sgl_kernel/sm90/common_ops.abi3.so. Or apply the patch and rebuild the common_ops_sm90_build target (~5 min).
02-expert-load-filter.patch both Each rank reads only its own experts (SGLANG_EXPERT_LOAD_FILTER=896). Load time drops from ~20 to ~5 min.
03-fp8-kv-prefill-gather.patch both FP8-KV prefill converts only the tokens it reads. Without it, add --max-total-tokens 524288 or prefill can OOM.
04-marlin-ep-block.patch W4A16 Fixes Marlin MoE tile size under EP (SGLANG_MARLIN_EP_BLOCK=1). Without it: 23.6 tok/s at 50 streams.
05-w4a8-count-sort.patch W4A8 Bit-identical, faster MoE routing permutation (K3_W4_SRC2DST=count).

Launch

Run on each of the 8 nodes, with NODE_RANK=0..7 and HEAD = node 0's IP:

export NCCL_MIN_NCHANNELS=24 NCCL_MAX_NCHANNELS=24 NCCL_IB_ADAPTIVE_ROUTING=0 NCCL_IB_QPS_PER_CONNECTION=2
export SGLANG_EXPERT_LOAD_FILTER=896 SGLANG_RAGGED_VERIFY_MODE=static SGLANG_JIT_DEEPGEMM_FAST_WARMUP=1

# W4A8 (this repo)
export K3_W4_SRC2DST=count
MODEL="--model-path Radioheading/Kimi-K3-W4AFP8"
# W4A16 (native) instead:
#   export SGLANG_MARLIN_EP_BLOCK=1
#   MODEL="--model-path moonshotai/Kimi-K3 --moe-runner-backend marlin"

sglang serve $MODEL --served-model-name moonshotai/Kimi-K3 --trust-remote-code \
  --tp-size 32 --ep-size 32 --nnodes 8 --node-rank $NODE_RANK --dist-init-addr $HEAD:20000 \
  --kv-cache-dtype fp8_e4m3 --mamba-ssm-dtype float32 --mem-fraction-static 0.82 \
  --prefill-attention-backend flashmla --decode-attention-backend flashmla \
  --disable-radix-cache --max-running-requests 64 --cuda-graph-max-bs-decode 64 --max-mamba-cache-size 192 \
  --chunked-prefill-size 4096 --max-prefill-tokens 8192 \
  --speculative-algorithm DSPARK --speculative-draft-model-path RadixArk/Kimi-K3-DSpark \
  --speculative-dspark-block-size 2 --speculative-draft-model-quantization unquant \
  --speculative-draft-attention-backend flashinfer --speculative-draft-kv-cache-dtype bfloat16 \
  --enable-linear-replayssm-spec \
  --reasoning-parser kimi_k3 --tool-call-parser kimi_k3 --host 0.0.0.0 --port 30000

The server is OpenAI-compatible: http://$HEAD:30000/v1, model moonshotai/Kimi-K3. It is ready after 8–15 minutes.

License

Derived from moonshotai/Kimi-K3 and distributed under its license (see LICENSE).

Downloads last month
43
Safetensors
Model size
1.4T params
Tensor type
F32
·
BF16
·
F8_E4M3
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Radioheading/Kimi-K3-W4AFP8

Quantized
(53)
this model