Kimi-K3-NVFP4 / README.md
kylesayrs's picture
Update README.md
71eed69 verified
|
Raw
History Blame Contribute Delete
1.16 kB
metadata
license: mit
base_model:
  - moonshotai/Kimi-K3
library_name: transformers
tags:
  - compressed-tensors
  - LLM Compressor
  - vLLM

RedHatAI/Kimi-K3-NVFP4

This is a quantized version of moonshotai/Kimi-K3 with MoE layers quantized to NVFP4 for accelerated inference.

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Kimi-K3-NVFP4 \
  --tensor-parallel-size 8 \
  --trust_remote_code \
  --load-format instanttensor \
  --reasoning-parser kimi_k3 \
  --language-model-only  # optional

May require https://github.com/vllm-project/vllm/pull/50500 to run

Creation Process

This model was compressed using LLM Compressor. For more information, please contact Kyle Sayers via the vLLM Developers Slack or ksayers@redhat.com.

Evaluation

Benchmark moonshotai/Kimi-K3 RedHatAI/Kimi-K3-NVFP4
GPQA 93.5 91.0
inspect eval hf/Idavidrein/gpqa/diamond \
  --model vllm/RedHatAI/Kimi-K3-NVFP4 \
  --reasoning-effort high \
  --model-base-url http://localhost:8000/v1