Text Generation
Transformers
Safetensors
kimi_k3
feature-extraction
nvfp4
fp4
quantized
Mixture of Experts
modelopt
sglang
custom_code
8-bit precision
Instructions to use PatronusAI/kimi-k3-nvfp4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PatronusAI/kimi-k3-nvfp4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PatronusAI/kimi-k3-nvfp4", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("PatronusAI/kimi-k3-nvfp4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PatronusAI/kimi-k3-nvfp4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PatronusAI/kimi-k3-nvfp4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/PatronusAI/kimi-k3-nvfp4
- SGLang
How to use PatronusAI/kimi-k3-nvfp4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PatronusAI/kimi-k3-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PatronusAI/kimi-k3-nvfp4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PatronusAI/kimi-k3-nvfp4", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use PatronusAI/kimi-k3-nvfp4 with Docker Model Runner:
docker model run hf.co/PatronusAI/kimi-k3-nvfp4
File size: 10,879 Bytes
7218095 8a65030 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 6ec7b0f 9e5aec5 58a284f e52f170 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 6ec7b0f 9e5aec5 cddfb2a 4af146e 9e5aec5 7218095 6ec7b0f 7218095 9e5aec5 6ec7b0f 9e5aec5 6ec7b0f 9e5aec5 7218095 9e5aec5 6ec7b0f 9e5aec5 7218095 9e5aec5 6ec7b0f e9efe32 9e5aec5 e52f170 58a284f e52f170 58a284f 9e5aec5 e52f170 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 6ec7b0f e52f170 6ec7b0f e52f170 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 9e5aec5 7218095 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 | ---
base_model: moonshotai/Kimi-K3
base_model_relation: quantized
library_name: transformers
pipeline_tag: text-generation
tags:
- nvfp4
- fp4
- quantized
- moe
- modelopt
- sglang
---
# Kimi-K3-NVFP4
NVFP4 conversion of [`moonshotai/Kimi-K3`](https://huggingface.co/moonshotai/Kimi-K3),
produced by a **closed-form, bit-exact MXFP4 β NVFP4 cast** with
[NVIDIA TensorRT Model Optimizer](https://github.com/NVIDIA/Model-Optimizer).
Tooling: [`patronus-ai/kimi-k3-nvfp4`](https://github.com/patronus-ai/kimi-k3-nvfp4)
| | |
| --- | --- |
| Cast fidelity | **85,085,650,944 / 85,085,650,944 blocks bit-exact (100.0000%)**, `max_abs_err 0.0` |
| Layers cast | 247,296 (routed MoE experts) |
| **GPQA Diamond** | **92.96%** (pass@1 avg-of-16) |
| **IFBench** | **75.27%** (pass@1 avg-of-5, prompt_loose) |
| **SciCode** | **23.08%** problem / **57.29%** subtask |
| Verified on | SGLang `kimi-k3` branch, 2 nodes Γ 8ΓB200, TP16 |
| Size | ~1.6 TB, 96 shards |
## What this is
| | Source (Kimi-K3) | This checkpoint |
| --- | --- | --- |
| Weight format | MXFP4 (E2M1, block 32, E8M0 scale) | **NVFP4** (E2M1, block 16, E4M3 scale + FP32 per-tensor scale) |
| Weight values | β | **bit-identical to the source** |
| Scope | routed MoE experts (`w1`/`w2`/`w3`) | routed MoE experts (`w1`/`w2`/`w3`) |
Attention, shared experts, dense MLP projections, `lm_head` and vision modules are
carried through unquantized.
## Bit-exactness
The cast is **data-free**: NVFP4 scales are derived from the source E8M0 exponents
rather than re-estimated from data. E8M0 has no mantissa, so every source scale is
already an exact power of two. With `m = k_max - 8`:
```
weight_scale_2 = 2^m (per tensor, FP32)
per_block_scale = 2^(k_j - m) (per 16 elements, E4M3)
=> product = 2^k_j exactly
```
E4M3 represents `2^k` exactly for `k β [-9, 8]`, so any block within 17 binades of
the tensor max reconstructs bit-for-bit. On Kimi-K3 that covered **every** block.
No calibration data was used, and none is needed.
> **Note on precision.** Because the per-block scale is always pinned to an
> exact power of two, this cast deliberately uses none of NVFP4's precision
> advantages at the weight level β neither the finer block-16 granularity nor
> E4M3's 3 mantissa bits.
> That is the cost of bit-exactness, and it is the right trade: quantization is
> irreversible, so a higher-resolution container cannot recover information MXFP4
> already discarded. NVFP4's accuracy edge is only reachable by re-quantizing from
> BF16, which for 2.8T params means materializing ~5.6 TB. What the conversion buys
> is access to the NVFP4 kernel and tooling ecosystem.
## Deployment (SGLang)
Verified on the SGLang **`kimi-k3` branch** (not `main`), 2 nodes Γ 8ΓB200, TP16.
Run the identical command on every node, varying only `--node-rank`.
```bash
# --- per-node environment ---
export SGLANG_HOST_IP=$(ip -o -4 addr show ens1 | awk '{print $4}' | cut -d/ -f1)
export GLOO_SOCKET_IFNAME=ens1 NCCL_SOCKET_IFNAME=ens1
export NCCL_IB_HCA=mlx5_0,mlx5_1,mlx5_2,mlx5_3,mlx5_4,mlx5_9,mlx5_12,mlx5_13
export NCCL_IB_GID_INDEX=7 NCCL_IB_DISABLE=0 NCCL_NVLS_ENABLE=0 NCCL_CROSS_NIC=0
export PATH=/path/to/venv/bin:$PATH # ninja must resolve from SHARED storage
python -m sglang.launch_server \
--model-path /path/to/kimi-k3-nvfp4 \
--trust-remote-code \
--tp-size 16 --nnodes 2 --node-rank ${NODE_RANK} \
--dist-init-addr ${HEAD_IP}:20000 \
--moe-runner-backend marlin \
--mem-fraction-static 0.85 \
--mamba-full-memory-ratio 7.21 \
--disable-flashinfer-autotune \
--watchdog-timeout 3600 --dist-timeout 3600 \
--reasoning-parser kimi_k3 --tool-call-parser kimi_k3 \
--model-loader-extra-config '{"enable_multithread_load": true}' \
--host 0.0.0.0 --port 30000
```
`SGLANG_HOST_IP` is mandatory for multi-node β without it the ranks never rendezvous
and startup hangs with no error.
### Five non-obvious requirements
Each of these independently prevents the model from serving:
1. **`--moe-runner-backend marlin`.** Kimi-K3 uses the `situ` (SiTuGlu) activation,
which is **not implemented** in SGLang's NVFP4 FlashInfer runner
(`_SUPPORTED_FP4_ACTIVATIONS = {silu, relu2, gelu}`). Marlin supports gated
`{silu, situ}` and forwards both SiTu parameters
(`gemm1_alpha`=Ξ²=4.0, `clamp_limit`=linear_Ξ²=25.0).
2. **Patch `modelopt_quant.py`**: `_SUPPORTED_ACT_STRS += ("situ",)`. The assert at
`:2552` runs *before* the per-backend dispatch at `:2556`, so it rejects `situ`
even when the selected backend supports it. Every backend already applies its own
stricter check afterwards.
3. **Raise `UNBALANCED_MODEL_LOADING_TIMEOUT_S`** (`load_model_utils.py`, default
480 s, no CLI flag). 1.6 TB over network storage means per-node load skew
routinely exceeds 8 minutes and aborts an otherwise healthy run.
4. **`--dist-timeout 3600`.** Marlin's `prepare_moe_nvfp4_layer_for_marlin` repacks
247k expert weights, exceeding NCCL's 600 s default; the first rank to finish then
times out alone at `SeqNum=1 _ALLGATHER_BASE`.
5. **Omit `--dcp-size`** unless FlashInfer is new enough β `cutedsl_mla_backend`
passes `enable_dcp`/`cp_world` to `trtllm_batch_decode_with_kv_cache_mla`, which
flashinfer 0.6.15 does not accept.
Patches 2β3 are scripted in
[`serving/patch_situ_nvfp4.py`](https://github.com/patronus-ai/kimi-k3-nvfp4/blob/main/serving/patch_situ_nvfp4.py).
### Querying
K3 is a reasoning model β it emits thinking tokens before the answer. Use a generous
`max_tokens` (β₯512) and read `choices[0].message.content`; the thinking is in
`reasoning_content`. `reasoning_effort` accepts `low` / `high` / `max`.
```bash
curl $URL/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "k3",
"messages": [{"role":"user","content":"What is 17 * 23? Reply with only the number."}],
"max_tokens": 1200, "temperature": 0, "reasoning_effort": "low"}'
```
With too small a budget you get `content: ""` and `finish_reason: "length"` β that is
truncated thinking, not a broken model.
### This is a W4A16 deployment
Marlin dequantizes the 4-bit weights to BF16 **inside the GEMM**, so weights stay
4-bit in HBM (the memory and bandwidth win is kept) but the math runs at BF16
tensor-core rate. Activations are never quantized.
Consequences: no activation calibration is needed or used and this deployment setup
does not deliver FP4 compute throughput. Accuracy should be at least that of native
MXFP4, since the weights are bit-identical and the activations are *higher* precision
than the source's MXFP8.
## Evaluation
All three are judge-free and scored with the official harnesses.
| Benchmark | Score | Setup |
| --- | --- | --- |
| **GPQA Diamond** | **92.96%** | pass@1 avg-of-16, sd 1.24; 198 questions |
| **IFBench** | **75.27%** | pass@1 avg-of-5 `prompt_loose`, sd 0.86; 300 prompts |
| **SciCode** | **23.08%** / 57.29% | problem / subtask; 65 problems, 288 subtasks |
GPQA per-seed: 96.0 / 93.4 / 91.9 / 91.4 / 92.9 / 93.4 / 90.9 / 93.9 /
92.4 / 91.4 / 93.9 / 92.4 / 92.9 / 93.9 / 93.4 / 92.9.
IFBench: `prompt_loose` counts a response only if every constraint on it is met;
`prompt_strict` was 71.07% and instruction-level loose 78.20%.
SciCode: AA test split, `prompt_config=eval/scicode/background`, greedy.
Harness: nemo-skills with prompt `eval/aai/mcq-4choices` and
`++eval_type=multichoice`, matching the
[Artificial Analysis methodology](https://artificialanalysis.ai/methodology/intelligence-benchmarking);
198 Diamond questions, greedy for seed 0 and `temperature=0.6, top_p=0.95` thereafter.
`tokens_to_generate=65536` is load-bearing: at 16384 the model exhausted its budget
inside `reasoning` on ~10% of questions and emitted an empty answer, scoring as wrong
for a purely harness reason β worth ~8.6 points on the greedy pass (87.4% β 96.0%).
Truncation fell from 9.6% to 0.58%.
## Layout
Per quantized module `<base>`:
```
<base>.weight uint8 (out, in/2) NVFP4-packed E2M1 nibbles
<base>.weight_scale E4M3 (out, in/16) per-block scale
<base>.weight_scale_2 fp32 scalar per-tensor scale
```
`config.json` reports `format: nvfp4-pack-quantized`, `quant_method: modelopt`,
`quant_algo: NVFP4`, `group_size: 16`, and an `exclude_modules` list of 55 globs
derived from the modules that actually kept an unpacked 2-D float `.weight`.
> **Branches.** `main` is weight-only (`input_activations: null`). `w4a4` adds
> `input_scale` placeholders of 1.0; they are unused by the marlin path.
## Bundled fix
`modeling_kimi_linear.py` here includes a fix to `KimiDeltaAttention`: `A_log` is
stored at `head_dim` (128) but the upstream release builds it at `num_heads` (96), so
`from_pretrained` fails to load its own weights.
The checkpoint holds **96 trained decay values zero-padded to 128** β verified across
all 69 `A_log` tensors here: elements `[96:128]` are exactly zero in every one. The
FLA kernel indexes by `num_heads`, so the padding is inert.
The bundled file follows
[discussion #150](https://huggingface.co/moonshotai/Kimi-K3/discussions/150): it sizes
`A_log` at `num_heads` and narrows the checkpoint tensor in `_load_from_state_dict`,
raising if the discarded tail is ever non-zero.
An earlier version of this file followed our own
[discussion #144](https://huggingface.co/moonshotai/Kimi-K3/discussions/144), which
enlarges the parameter to `head_dim` instead. That approach loads but cannot run:
`fla/ops/kda/gate.py` computes `A_log.view(H, 1)` with `H = g.shape[-2]` (the head
count), so a 128-element parameter raises
`shape '[96, 1]' is invalid for input of size 128` on the first forward. If you pulled
this checkpoint before 2026-07-31, re-fetch `modeling_kimi_linear.py`.
This matters only for `transformers`-based loading β SGLang narrows to the first
`num_heads` elements independently, which the zero-padding makes lossless.
## Known limitation: the FP4 fast path
The trtllm-gen cubin pool ships SiTu kernels at NVFP4's 16-element scale block
(57 Γ BF16-activation, 45 Γ E2M1-activation), but they are unreachable from SGLang's
NVFP4 runner, which resolves activations via FlashInfer's `ActivationType` (no SiTu
member) and reads a different cubin store. Wiring the dispatch to
`sglang.kernels.ops.moe.trtllm_gen_moe` reaches the JIT build stage and then fails on
CUDA 12.8 with `namespace "cuda" has no member "maximum"` β `cuda::maximum` requires
CUDA 13's CCCL. **Untested beyond that point.**
## Reproducing
```bash
python cast_mxfp4_to_nvfp4_offline.py \
--source_ckpt /path/to/Kimi-K3 \
--output_ckpt /path/to/Kimi-K3-NVFP4 --verify
```
Runs on CPU without instantiating the model (~7 h, bounded by shard I/O).
## License
Derivative of `moonshotai/Kimi-K3`; the upstream model's license and terms apply.
|