Kimi-K3-NVFP4 / README.md
kylesayrs's picture
Update README.md
71eed69 verified
|
Raw
History Blame Contribute Delete
1.16 kB
---
license: mit
base_model:
- moonshotai/Kimi-K3
library_name: transformers
tags:
- compressed-tensors
- LLM Compressor
- vLLM
---
# RedHatAI/Kimi-K3-NVFP4
This is a quantized version of `moonshotai/Kimi-K3` with MoE layers quantized to NVFP4 for accelerated inference.
## Usage
This model is intended for deployment with vLLM. You can serve the model using
```bash
vllm serve RedHatAI/Kimi-K3-NVFP4 \
--tensor-parallel-size 8 \
--trust_remote_code \
--load-format instanttensor \
--reasoning-parser kimi_k3 \
--language-model-only # optional
```
May require https://github.com/vllm-project/vllm/pull/50500 to run
## Creation Process
This model was compressed using LLM Compressor. For more information, please contact Kyle Sayers via the [vLLM Developers Slack](https://communityinviter.com/apps/vllm-dev/join-vllm-developers-slack) or ksayers@redhat.com.
## Evaluation ##
| Benchmark | `moonshotai/Kimi-K3` | `RedHatAI/Kimi-K3-NVFP4`
| - | - | -|
| GPQA | 93.5 | 91.0 |
```bash
inspect eval hf/Idavidrein/gpqa/diamond \
--model vllm/RedHatAI/Kimi-K3-NVFP4 \
--reasoning-effort high \
--model-base-url http://localhost:8000/v1
```