File size: 3,849 Bytes
612f4e8
 
 
 
9d2e1b2
 
612f4e8
37d02d6
9d2e1b2
37d02d6
9d2e1b2
c655dee
9d2e1b2
 
 
550c759
 
9d2e1b2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
---
license: other
license_name: modified-mit
license_link: LICENSE
base_model:
- moonshotai/Kimi-K2.6
---

# Model Overview

- **Model Architecture:** Kimi-K2.6
  - **Input:** Text, Image, Video
  - **Output:** Text
- **Supported Hardware Microarchitecture:** AMD MI300/MI350/MI355 (emulation)
- **ROCm:** 7.2.2
- **PyTorch**: 2.10.0
- **Transformers**: 5.2.0
- **Operating System(s):** Linux
- **Inference Engine:** [vLLM](https://docs.vllm.ai/en/latest/)
- **Model Optimizer:** [AMD-Quark](https://quark.docs.amd.com/latest/index.html) (V0.12)
  - **Quantized layers:** `experts` and `shared_experts`
  - **Weight quantization:** NVFP4, Static 
  - **Activation quantization:** NVFP4, Dynamic 
- **Calibration Dataset:** [Pile](https://huggingface.co/datasets/mit-han-lab/pile-val-backup)

This model was built with Kimi-K2.6 model by applying [AMD-Quark](https://quark.docs.amd.com/latest/index.html) for NVFP4 quantization.

# Model Quantization

The model was quantized from [moonshotai/Kimi-K2.6](https://huggingface.co/moonshotai/Kimi-K2.6) using [AMD-Quark](https://quark.docs.amd.com/latest/index.html). The weights and activations are quantized to NVFP4.

**Quantization scripts:**
```
cd Quark/examples/torch/language_modeling/llm_ptq/
export output_dir=amd/Kimi-K2.6-NVFP4
exclude_layers="*self_attn* *mlp.gate *mlp.gate.linear *lm_head *mlp.gate_proj *mlp.up_proj *mlp.down_proj *mm_projector* *vision_tower*"
python3 quantize_quark.py --model_dir $MODEL_DIR \
                          --quant_scheme nvfp4 \
                          --num_calib_data 128 \
                          --exclude_layers $exclude_layers \
                          --model_export hf_format \
                          --output_dir $output_dir \
                          --trust_remote_code \
                          --multi_gpu balanced 
```

# Deployment
### Use with vLLM

This model can be deployed efficiently using the [vLLM](https://docs.vllm.ai/en/latest/) backend.

## Evaluation
The model was evaluated on GSM8K and MMLU_PRO benchmarks. 

### Accuracy

<table>
  <tr>
   <td><strong>Benchmark</strong>
   </td>
   <td><strong>Kimi-K2.6 </strong>
   </td>
   <td><strong>Kimi-K2.6-NVFP4(this model) </strong>
   </td>
   <td><strong>Recovery</strong>
   </td>
  </tr>
  <tr>
   <td>GSM8K (flexible-extract)
   </td>
   <td>93.93
   </td>
   <td>93.48
   </td>
   <td>99.52%
   </td>
  </tr>
    <tr>
   <td>MMLU_PRO (exact-extract)
   </td>
   <td>81.43
   </td>
   <td>79.21
   </td>
   <td>97.27%
   </td>
  </tr>
</table>

### Reproduction

The GSM8K and MMLU_PRO results were obtained using the `lm-evaluation-harness` framework, based on the Docker image `rocm/vllm-dev:nightly_main_20260603`.

Install the lm-eval `(Version: 0.4.12)` in container first.
```
pip install lm-eval[api]
```

#### Launching server
```
export VLLM_ROCM_USE_AITER=1
vllm serve amd/Kimi-K2.6-NVFP4 -tp 8 \
  --mm-encoder-tp-mode data \
  --tool-call-parser kimi_k2 \
  --reasoning-parser kimi_k2 \
  --enforce-eager \
  --trust-remote-code
```

#### Evaluating model in a new terminal
```
lm_eval \
  --model local-completions \
  --model_args "model=amd/Kimi-K2.6-NVFP4,kv_cache_dtype=fp8,base_url=http://0.0.0.0:8000/v1/completions,tokenized_requests=False,tokenizer_backend=None,num_concurrent=32" \
  --tasks gsm8k \
  --num_fewshot 5 \
  --batch_size 1 
```
```
lm_eval \
  --model local-completions \
  --model_args "model=amd/Kimi-K2.6-NVFP4,kv_cache_dtype=fp8,base_url=http://0.0.0.0:8000/v1/completions,tokenized_requests=False,tokenizer_backend=None,num_concurrent=32,max_length=16384,timeout=14400" \
  --tasks mmlu_pro \
  --gen_kwargs "do_sample=True,temperature=1.0,top_p=0.95,max_tokens=4096,max_gen_toks=4096" \
  --batch_size auto \
  --limit 100
```

# License
Modifications Copyright(c) 2026 Advanced Micro Devices, Inc. All rights reserved.