Kimi-K2.7-Code GGUF — Quantized by BatiAI

BatiFlow moonshot MoE

The coding upgrade to Kimi K2.6 — +21.8% on Kimi Code Bench v2, running on a 512GB Mac Studio. IQ3_XXS / IQ4_XS GGUF of moonshotai/Kimi-K2.7-Code (1T total / 32.6B active MoE, DeepSeek-V3-family architecture). Quantized directly from official Moonshot weights — code+multilingual imatrix, BatiAI-signed.

📦 Quantizations

Quant Size Shards Target
IQ3_XXS 394 GB (GiB: 367) 10 M3 Ultra 512GB Mac Studio
IQ4_XS 546 GB (GiB: 509) 13 512GB+ / multi-node / server

Both built from official weights via a Q8_0 intermediate, quantized with a code + EN + KO + ZH imatrix (included: Kimi-K2.7-Code-imatrix.dat). Text-only (the vision tower of the K2.5-family checkpoint is not included; same as other K2 GGUFs).

✅ Verified (this build, IQ3_XXS) — captured greedy runs:

  • Math: 127+58185 (clean reasoning trace)
  • Korean: 서울 소개 + 김치·비빔밥·불고기 각 한 문장 — fluent, zero token-mixing or loops
  • Tool-call: {"tool":"get_weather","args":{"city":"부산"}} — exact JSON

🚀 Usage (llama.cpp — mainline, no fork needed)

hf download batiai/Kimi-K2.7-Code-GGUF "Kimi-K2.7-Code-IQ3_XXS-*.gguf" --local-dir ./k27

# llama.cpp auto-loads all shards from the first one
./llama-cli -m ./k27/Kimi-K2.7-Code-IQ3_XXS-00001-of-00010.gguf -ngl 99 -c 16384 \
  -p "Refactor this function and add tests."

Recommended sampling (Moonshot): --temp 1.0 --top-p 0.95 (thinking mode). Architecture is deepseek2 — supported by mainline llama.cpp out of the box. Ollama tags (batiai/kimi-k2.7-code) follow shortly.

📜 License

Modified MIT (Moonshot) — commercial use permitted; products exceeding 100M MAU / $20M monthly revenue must display "Kimi K2.7" attribution. Full text at the base model repo. Quantized weights redistributed under the same terms.

✨ What BatiAI did

  • Direct from official Moonshot weights (never a re-quant of third-party GGUFs)
  • Q8_0 intermediate + diverse imatrix (code/EN/KO/ZH) for balanced fidelity
  • Verified: load ✅ · math ✅ · Korean ✅ · tool-call JSON ✅ — BatiAI metadata-signed

BatiAI · on-device frontier AI · https://flow.bati.ai

Downloads last month
-
GGUF
Model size
1T params
Architecture
deepseek2
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for batiai/Kimi-K2.7-Code-GGUF

Quantized
(29)
this model