| --- |
| license: mit |
| base_model: |
| - moonshotai/Kimi-K3 |
| library_name: compressed-tensors |
| tags: |
| - quantized |
| - fp8 |
| - fp8_block |
| - llm-compressor |
| --- |
| |
| # RedHatAI/Kimi-K3-FP8-BLOCK |
|
|
| This is a variant of `moonshotai/Kimi-K3` quantized with FP8_BLOCK quantization applied to MLP and experts layers. Attention layers were not quantized due to shape mismatches, namely with `self_attn.b_proj.weight` |
| |
| ## Usage ## |
| ```bash |
| docker run --gpus all \ |
| --privileged --ipc=host -p 8000:8000 \ |
| -v ~/.cache/huggingface:/root/.cache/huggingface \ |
| -e GLOO_SOCKET_IFNAME=$IFACE_NAME \ |
| -e NCCL_SOCKET_IFNAME=$IFACE_NAME \ |
| -e VLLM_ENABLE_K3_LATENT_MOE_TAIL_FUSION=1 \ |
| vllm/vllm-openai:kimi-k3 RedHatAI/Kimi-K3-FP8-BLOCK \ |
| --trust-remote-code \ |
| --load-format fastsafetensors \ |
| --gpu-memory-utilization 0.95 \ |
| --tensor-parallel-size 16 \ |
| --nnodes 2 \ |
| --node-rank 0 \ |
| --master-addr $HEAD_IP \ |
| --no-enable-flashinfer-autotune \ |
| --disable-custom-all-reduce \ |
| --enable-auto-tool-choice \ |
| --tool-call-parser kimi_k3 \ |
| --reasoning-parser kimi_k3 |
| ``` |
| |
| ## Creation Process |
| |
| Quantized using [llm-compressor](https://github.com/vllm-project/llm-compressor): |
| |
| ```python |
| from compressed_tensors.entrypoints.convert import CompressedTensorsDequantizer |
| |
| from llmcompressor import model_free_ptq |
| |
| MODEL_ID = "moonshotai/Kimi-K3" |
| SAVE_DIR = MODEL_ID.rstrip("/").split("/")[-1] + "-FP8-BLOCK" |
|
|
| ignore = [ |
| "re:.*embed_tokens.*", |
| "re:.*self_attn.*", # no attention because (q_proj|k_proj|v_proj|b_proj|f_a_proj) are all fused, and `b_proj` has a weight non-divisible by 128 |
| "re:.*block_sparse_moe\.gate.*", |
| "re:.*self_attention_res_proj.*", |
| "re:.*mlp_res_proj.*", |
| "re:.*output_attn_res_proj.*", |
| "re:.*lm_head.*", |
| "re:.*vision_tower.*", |
| "re:.*mm_projector.*", |
| ] |
| |
| model_free_ptq( |
| model_stub=MODEL_ID, |
| save_directory=SAVE_DIR, |
| scheme="FP8_BLOCK", |
| ignore=ignore, |
| converter=CompressedTensorsDequantizer( |
| MODEL_ID, |
| ignore=ignore, |
| ), |
| max_workers=7, |
| device=[ |
| f"cuda:{i}" |
| for i in range(7) |
| ], |
| ) |
| |
| ``` |