File size: 3,849 Bytes
612f4e8 9d2e1b2 612f4e8 37d02d6 9d2e1b2 37d02d6 9d2e1b2 c655dee 9d2e1b2 550c759 9d2e1b2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 | ---
license: other
license_name: modified-mit
license_link: LICENSE
base_model:
- moonshotai/Kimi-K2.6
---
# Model Overview
- **Model Architecture:** Kimi-K2.6
- **Input:** Text, Image, Video
- **Output:** Text
- **Supported Hardware Microarchitecture:** AMD MI300/MI350/MI355 (emulation)
- **ROCm:** 7.2.2
- **PyTorch**: 2.10.0
- **Transformers**: 5.2.0
- **Operating System(s):** Linux
- **Inference Engine:** [vLLM](https://docs.vllm.ai/en/latest/)
- **Model Optimizer:** [AMD-Quark](https://quark.docs.amd.com/latest/index.html) (V0.12)
- **Quantized layers:** `experts` and `shared_experts`
- **Weight quantization:** NVFP4, Static
- **Activation quantization:** NVFP4, Dynamic
- **Calibration Dataset:** [Pile](https://huggingface.co/datasets/mit-han-lab/pile-val-backup)
This model was built with Kimi-K2.6 model by applying [AMD-Quark](https://quark.docs.amd.com/latest/index.html) for NVFP4 quantization.
# Model Quantization
The model was quantized from [moonshotai/Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6) using [AMD-Quark](https://quark.docs.amd.com/latest/index.html). The weights and activations are quantized to NVFP4.
**Quantization scripts:**
```
cd Quark/examples/torch/language_modeling/llm_ptq/
export output_dir=amd/Kimi-K2.6-NVFP4
exclude_layers="*self_attn* *mlp.gate *mlp.gate.linear *lm_head *mlp.gate_proj *mlp.up_proj *mlp.down_proj *mm_projector* *vision_tower*"
python3 quantize_quark.py --model_dir $MODEL_DIR \
--quant_scheme nvfp4 \
--num_calib_data 128 \
--exclude_layers $exclude_layers \
--model_export hf_format \
--output_dir $output_dir \
--trust_remote_code \
--multi_gpu balanced
```
# Deployment
### Use with vLLM
This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend.
## Evaluation
The model was evaluated on GSM8K and MMLU_PRO benchmarks.
### Accuracy
<table>
<tr>
<td><strong>Benchmark</strong>
</td>
<td><strong>Kimi-K2.6 </strong>
</td>
<td><strong>Kimi-K2.6-NVFP4(this model) </strong>
</td>
<td><strong>Recovery</strong>
</td>
</tr>
<tr>
<td>GSM8K (flexible-extract)
</td>
<td>93.93
</td>
<td>93.48
</td>
<td>99.52%
</td>
</tr>
<tr>
<td>MMLU_PRO (exact-extract)
</td>
<td>81.43
</td>
<td>79.21
</td>
<td>97.27%
</td>
</tr>
</table>
### Reproduction
The GSM8K and MMLU_PRO results were obtained using the `lm-evaluation-harness` framework, based on the Docker image `rocm/vllm-dev:nightly_main_20260603`.
Install the lm-eval `(Version: 0.4.12)` in container first.
```
pip install lm-eval[api]
```
#### Launching server
```
export VLLM_ROCM_USE_AITER=1
vllm serve amd/Kimi-K2.6-NVFP4 -tp 8 \
--mm-encoder-tp-mode data \
--tool-call-parser kimi_k2 \
--reasoning-parser kimi_k2 \
--enforce-eager \
--trust-remote-code
```
#### Evaluating model in a new terminal
```
lm_eval \
--model local-completions \
--model_args "model=amd/Kimi-K2.6-NVFP4,kv_cache_dtype=fp8,base_url=http://0.0.0.0:8000/v1/completions,tokenized_requests=False,tokenizer_backend=None,num_concurrent=32" \
--tasks gsm8k \
--num_fewshot 5 \
--batch_size 1
```
```
lm_eval \
--model local-completions \
--model_args "model=amd/Kimi-K2.6-NVFP4,kv_cache_dtype=fp8,base_url=http://0.0.0.0:8000/v1/completions,tokenized_requests=False,tokenizer_backend=None,num_concurrent=32,max_length=16384,timeout=14400" \
--tasks mmlu_pro \
--gen_kwargs "do_sample=True,temperature=1.0,top_p=0.95,max_tokens=4096,max_gen_toks=4096" \
--batch_size auto \
--limit 100
```
# License
Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.
|