Kimi-K3-W4AFP8
INT4 (group 128) routed experts with FP8 activations, FP8 attention projections, requantized from the native MXFP4 weights of moonshotai/Kimi-K3. Same tensor format and kernels as vessl/Kimi-K3-W4AFP8, with lower error against the native model.
Two configurations on 32 x GH200
8 nodes x 4 GH200, TP32 / EP32, DSpark speculative decoding (block 2), FP8 KV cache. Real AIME/HMMT prompts with real EOS; numbers are median output tok/s per stream.
| W4A16 — moonshotai/Kimi-K3 (native MXFP4) | W4A8 — this repo | |
|---|---|---|
| tok/s per stream, 50 concurrent | 29.6 | 34.1 |
| tok/s per stream, 20 concurrent | 40.6 | 43.6 |
| aggregate tok/s, 50 concurrent | 883 | 949 |
| KL to native, nats/token (95% CI) | +0.0011 (0.0006–0.0016) | +0.0154 (0.0142–0.0167) |
| AIME 2026 + HMMT Feb 2026, 8 samples/problem | 82.3% | 81.3% (−1.0 pt, SE 1.6: not significant) |
- Patches in the measured runs: W4A16 used patches 02 and 04 (no
--max-total-tokenscap). W4A8 used 01, 02, 03 and 05, i.e. exactly the launch command below. - Pick W4A16 when you want the reference model.
- Pick W4A8 for about 15% more throughput at 50 streams.
- Scaling out: two W4A8 replicas (64 GPUs) behind
sglang-router, 25 streams each, give 35.7 tok/s per stream. - KL: teacher-forced on 355k tokens of native-model math/proof continuations. For comparison, vessl/Kimi-K3-W4AFP8 scores +0.0187.
Setup
- Image: SGLang v0.5.20, arm64, CUDA 13.
- Clone the repo:
git clone https://huggingface.co/Radioheading/Kimi-K3-W4AFP8. You need itsapply_patch.pyandpatches/below. - Patches: apply the ones listed below inside the image's source tree (
cd /sgl-workspace/sglang && git apply <patch>). - W4A8 quantization method: for W4A8 only, run
python3 apply_patch.pyfrom this repo. It registers the method (from vessl/Kimi-K3-W4AFP8). - Draft model: download RadixArk/Kimi-K3-DSpark.
| patch | needed for | effect |
|---|---|---|
01-sgl-kernel-sm90a.patch |
W4A8 (required) | Builds the SM90 CUTLASS kernels with sm_90a. aarch64 wheels lack it, and W4A8/FP8 kernels abort with "Arch conditional MMA instruction". Shortcut: copy the prebuilt sgl_kernel_sm90a/common_ops.abi3.so (arm64, CUDA 13, v0.5.20) over site-packages/sgl_kernel/sm90/common_ops.abi3.so. Or apply the patch and rebuild the common_ops_sm90_build target (~5 min). |
02-expert-load-filter.patch |
both | Each rank reads only its own experts (SGLANG_EXPERT_LOAD_FILTER=896). Load time drops from ~20 to ~5 min. |
03-fp8-kv-prefill-gather.patch |
both | FP8-KV prefill converts only the tokens it reads. Without it, add --max-total-tokens 524288 or prefill can OOM. |
04-marlin-ep-block.patch |
W4A16 | Fixes Marlin MoE tile size under EP (SGLANG_MARLIN_EP_BLOCK=1). Without it: 23.6 tok/s at 50 streams. |
05-w4a8-count-sort.patch |
W4A8 | Bit-identical, faster MoE routing permutation (K3_W4_SRC2DST=count). |
Launch
Run on each of the 8 nodes, with NODE_RANK=0..7 and HEAD = node 0's IP:
export NCCL_MIN_NCHANNELS=24 NCCL_MAX_NCHANNELS=24 NCCL_IB_ADAPTIVE_ROUTING=0 NCCL_IB_QPS_PER_CONNECTION=2
export SGLANG_EXPERT_LOAD_FILTER=896 SGLANG_RAGGED_VERIFY_MODE=static SGLANG_JIT_DEEPGEMM_FAST_WARMUP=1
# W4A8 (this repo)
export K3_W4_SRC2DST=count
MODEL="--model-path Radioheading/Kimi-K3-W4AFP8"
# W4A16 (native) instead:
# export SGLANG_MARLIN_EP_BLOCK=1
# MODEL="--model-path moonshotai/Kimi-K3 --moe-runner-backend marlin"
sglang serve $MODEL --served-model-name moonshotai/Kimi-K3 --trust-remote-code \
--tp-size 32 --ep-size 32 --nnodes 8 --node-rank $NODE_RANK --dist-init-addr $HEAD:20000 \
--kv-cache-dtype fp8_e4m3 --mamba-ssm-dtype float32 --mem-fraction-static 0.82 \
--prefill-attention-backend flashmla --decode-attention-backend flashmla \
--disable-radix-cache --max-running-requests 64 --cuda-graph-max-bs-decode 64 --max-mamba-cache-size 192 \
--chunked-prefill-size 4096 --max-prefill-tokens 8192 \
--speculative-algorithm DSPARK --speculative-draft-model-path RadixArk/Kimi-K3-DSpark \
--speculative-dspark-block-size 2 --speculative-draft-model-quantization unquant \
--speculative-draft-attention-backend flashinfer --speculative-draft-kv-cache-dtype bfloat16 \
--enable-linear-replayssm-spec \
--reasoning-parser kimi_k3 --tool-call-parser kimi_k3 --host 0.0.0.0 --port 30000
The server is OpenAI-compatible: http://$HEAD:30000/v1, model moonshotai/Kimi-K3. It is ready after 8–15 minutes.
License
Derived from moonshotai/Kimi-K3 and distributed under its license (see LICENSE).
- Downloads last month
- 43
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support
Model tree for Radioheading/Kimi-K3-W4AFP8
Base model
moonshotai/Kimi-K3