--- license: apache-2.0 base_model: Kwaipilot/KAT-Coder-V2.5-Dev tags: - gguf - llama.cpp - turboquant - qwen35moe - quantized - imatrix - mtp - speculative-decoding - coding - agent - consumer-gpu - offmoreal --- # KAT-Coder-V2.5-Dev MaxQuality MTP GGUF ## ⚡ Speed at a glance | Variant | Context | Generation speed | CPU offload | Fits entirely in 16 GB VRAM | |---|---:|---:|---|:---:| | **`Q2_K-AllGPU` MTP** | 10K | **~156 tok/s** | **none — whole model + MTP head resident in VRAM** | ✅ | | `Q4_K_M` MTP | 10K | ~65 tok/s | ~20 expert layers on CPU | ❌ | `Q2_K-AllGPU` is the one to reach for on a single 16 GB consumer card: routed experts quantized to Q2_K (with the more error-sensitive `down` projection bumped one step to Q3_K to claw back quality), zero CPU offload. Expect a real, measured **~10-15% quality drop** versus the Q4 file in exchange for that speed — this is the most aggressive quant in this repo, not a free lunch. ## Tested hardware and software - Linux - NVIDIA GeForce RTX 5060 Ti, 16 GB VRAM - AMD Ryzen 9 5950X, 16 cores / 32 threads - 64 GB system RAM - [TheTom/llama-cpp-turboquant](https://github.com/TheTom/llama-cpp-turboquant), commit `c26cbdffc` (ships working `--spec-type draft-mtp` support for this architecture) :::warning MTP acceptance rate is lower for this model than for its sibling release We grafted the same donor MTP head (from `Qwen/Qwen3.6-35B-A3B`, verbatim, no retraining — see below) onto both this model and our [Ornith-1.0-35B MTP release](https://huggingface.co/offmonreal/Ornith-1.0-35B-MaxQuality-MTP-GGUF). Measured back-to-back on identical hardware, identical launch flags, identical 10K-token prompt: KAT-Coder accepts **50.9%** of draft tokens (328 proposed / 167 accepted) versus Ornith's **67.5%** (268 / 181). The head was never trained for either model, so acceptance rate depends on how far each fine-tune's internal representations drifted from the shared base — that drift is apparently larger for KAT-Coder. Speed numbers above are still a real win over no-MTP, just a smaller one than on Ornith. ::: Same quality-first GGUF quants as [KAT-Coder-V2.5-Dev-MaxQuality-iMatrix-GGUF](https://huggingface.co/offmonreal/KAT-Coder-V2.5-Dev-MaxQuality-iMatrix-GGUF), with one addition: a **Multi-Token Prediction (MTP) draft head grafted in**, giving self-contained speculative decoding — no separate draft-model file needed, everything is in one GGUF. The head was not trained for KAT-Coder. KAT-Coder is a fine-tune of `Qwen/Qwen3.6-35B-A3B`, and that base model ships a real, trained MTP head that KAT-Coder's own release doesn't include. We extracted that head (verbatim, no retraining) and grafted it in as an extra transformer block, using only the ~900 MiB of donor weights that layer needs — not the full base model. Every other tensor in these files is untouched KAT-Coder. ## Files | File | Size | BPW | Recipe | Intended use | |---|---:|---:|---|---| | `KAT-Coder-V2.5-Dev_Q4_K_M.gguf` | 20.55 GiB | 4.88 | Same as Q4_K_M iMatrix | Higher quality, MTP for a speed bonus, requires CPU offload on 16 GB cards | | `KAT-Coder-V2.5-Dev_Q2_K-AllGPU.gguf` | 13.09 GiB | 3.03 (trunk) | Routed-expert `gate`/`up` at Q2_K, `down` bumped to Q3_K, embeddings/output at Q5_K, attention at Q4_K | **AllGPU**: the whole model, MTP head included, fits in a 16 GB card with zero CPU offload — the fastest file in this repo, at a real quality cost from the aggressive routed-expert Q2_K | Both are calibrated with the same iMatrix as the base release (`calibration_datav5.txt`, 802 chunks). `Q4_K_M.gguf` is byte-identical to the non-MTP release except for the added MTP block; `Q2_K-AllGPU.gguf` has a changed base recipe (see above). ## MTP quantization The grafted block (`blk.40` in the GGUF — the model's own layer count plus one) is copied verbatim from the donor at its native precision: `Q8_0` for the routed/shared expert and attention projections, `BF16` for the router, `F32` for norms. It is **not** requantized to match the trunk's Q4/Q2 precision — the head is small (under 1 GiB) and precision there disproportionately affects acceptance rate, so we left it alone. ## Recommended launch commands ### Q2_K-AllGPU MTP — everything on GPU, no CPU offload ```bash llama-server \ --jinja --host 0.0.0.0 --port 8080 \ -m KAT-Coder-V2.5-Dev_Q2_K-AllGPU.gguf \ --spec-type draft-mtp --spec-draft-n-max 4 \ --n-gpu-layers 99 --n-cpu-moe 0 \ --kv-offload \ --ctx-size 131072 --parallel 1 \ --flash-attn on \ --cache-type-k turbo3 --cache-type-v turbo3 \ --batch-size 32768 --ubatch-size 1024 --cache-reuse 256 ``` No `-ot` line at all — the entire trunk plus MTP head is small enough to sit in VRAM outright. Context is capped at 131072 (half of the Q4 file's 256K) to leave headroom for the KV cache on GPU; push past that and the server needs CPU offload like the file below. ### Q4_K_M MTP (256K context, CPU offload) ```bash llama-server \ --jinja --host 0.0.0.0 --port 8080 \ -m KAT-Coder-V2.5-Dev_Q4_K_M.gguf \ --spec-type draft-mtp --spec-draft-n-max 2 \ --n-gpu-layers 99 --n-cpu-moe 0 \ -ot "blk\.(2[0-9]|3[0-9])\.ffn_.*_exps\.weight=CPU" \ --ctx-size 262144 --parallel 1 \ --flash-attn on \ --cache-type-k turbo3 --cache-type-v turbo3 \ --batch-size 262144 --ubatch-size 1024 --cache-reuse 256 ``` ### Context-compaction agent — prompt-throughput profile (MTP disabled) Not every workload wants MTP. Context-compaction / summarization agents feed the model a huge prompt and care about *ingestion* speed, not generation speed — MTP's draft-then-verify overhead only adds latency there for no benefit, so this profile leaves `--spec-type` off entirely (the file still has the MTP head baked in, it's just unused). Fewer layers need CPU offload as a result (only 4, vs. 20 for the Q4 generation profile), and `--ubatch-size` is pushed up to 2048 to maximize prefill throughput instead of favoring low-latency decode. ```bash llama-server \ --jinja --host 0.0.0.0 --port 8080 \ -m KAT-Coder-V2.5-Dev_Q2_K-AllGPU.gguf \ --n-gpu-layers 99 --n-cpu-moe 0 \ -ot "blk\.3[6-9]\.ffn_.*_exps\.weight=CPU" \ --ctx-size 220000 --parallel 1 \ --flash-attn on \ --cache-type-k turbo3 --cache-type-v turbo3 \ --batch-size 8192 --ubatch-size 2048 --cache-reuse 256 ``` Measured on a real 825-message conversation compaction job (630,995 characters of transcript, 633,977 with prompt overhead, 214,478 tokens by the model's own tokenizer — 2.96 characters/token): **228 s prompt processing, ~941 tok/s prefill throughput.** This is prefill (prompt-ingestion) speed, not generation speed — the two are not comparable to the tok/s figures elsewhere in this card. ## Upstream model card This release is based on the original KAT-Coder-V2.5-Dev model card reproduced in full below. Its original license declaration and complete model card are preserved unchanged. --- license: apache-2.0 language: - en - zh library_name: transformers pipeline_tag: text-generation tags: - code - agent - agentic-coding - moe - coding base_model: - Qwen3.6-35B-A3B ---
| This repository contains the model weights and configuration files for the post-trained KAT-Coder-V2.5-Dev in the Hugging Face Transformers format. The artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. Note: this open-weight release ships only the language-model weights and operates as a text-only model; the vision/multimodal components are not included and are unavailable. |
| Benchmark | KAT-Coder-V2.5-Dev | Qwen3.5-27B | Qwen3.6-35BA3B | Gemma4-31B | Qwen3.5-35BA3B | Ornith-1.0-35B | Gemma4-26BA4B | Qwen3-Coder-30B |
| Coding Agent | ||||||||
| SWE-bench Verified | 69.40 | 68.60 | 64.40 | 60.60 | 58.60 | 55.80 | 35.80 | 31.80 |
| SWE-bench Multilingual | 63.00 | 57.67 | 57.00 | 49.33 | 47.67 | 51.67 | 27.33 | 20.67 |
| SWE-bench Pro | 45.96 | 42.13 | 40.63 | 32.97 | 38.03 | 34.47 | 9.58 | 19.84 |
| Terminal-Bench 2.1 | 41.02 32.60 / 49.44 |
34.84 41.57 / 28.10 |
32.02 34.83 / 29.20 |
32.59 30.34 / 34.83 |
26.12 26.44 / 25.80 |
35.98 35.96 / 36.00 |
20.94 27.27 / 14.60 |
13.50 10.11 / 16.90 |
| PinchBench | 93.43 | 90.71 | 92.21 | 85.53 | 88.75 | 91.62 | 82.01 | 72.3 |
| Scicode | 44.20 | 25.58 | 37.53 | 33.19 | 27.73 | 30.34 | 30.84 | 18.27 |
| KAT-Code-Bench | 46.21 | 44.83 | 42.76 | 37.93 | 35.86 | 33.10 | 22.06 | 15.17 |
1. Evaluation method. All metrics presented in the table are reproduced in-house: we download the public model checkpoints, deploy them via vLLM or SGLang, and evaluate under a unified standardized pipeline. No officially reported results of the respective models are directly adopted in this table. Each model is tested only once on each evaluation set; retests are conducted only if obvious errors are found.
2. Evaluation configuration.
* SWE-bench Verified / Multilingual / Pro, KAT-Code-Bench: agent=claude_code@2.1.195, pass@k=1, temperature=1.0, top_p=0.95, 256k ctx.
* Terminal-Bench 2.1: agent=terminus-2 / claude_code, pass@k=1, temperature=0.7, top_p=1.0, 256k ctx.
* PinchBench: agent=openclaw@2026.3.13, pass@k=1, temperature=0.7, top_p=1.0, 256k ctx.
* Scicode: pass@k=1, temperature=0.6, top_p=1.0, 256k ctx.
3. Anomaly description.
* Qwen3.6-35BA3B: We found that on the SWE-bench Verified, SWE-bench Multilingual, and SWE-bench Pro test sets, our test results this time have an approximate 10 pp gap compared with the official results. We believe this is mainly caused by the harness version and some optimizations made to the test sets by the Qwen team, and it should not be an issue with the model itself.
* Qwen3.5-35BA3B: We observed frequent hallucinations during evaluation, including attempts to invoke the unavailable MultiEdit tool under the current agent environment, which negatively impacts the final metric.
* Gemma4-26B-A4B-it: Two main factors degrade evaluation performance: context overflow (exceeding the 256k context limit) and hallucinated calls to the unsupported MultiEdit tool in this evaluation setup.
The above deviations arise from mismatches between model tool preference and the allowed toolset in the evaluation harness, rather than inherent capability limitations of the models.
The proportion of samples that pass the unit test (1 for pass, 0 for fail) within a batch of samples during RL training.
## Quickstart For streamlined integration, we recommend using KAT-Coder-V2.5-Dev via APIs. Below is a guide to use KAT-Coder-V2.5-Dev via OpenAI-compatible API. ### Serving KAT-Coder-V2.5-Dev KAT-Coder-V2.5-Dev can be served via APIs with popular inference frameworks. In the following, we show example commands to launch OpenAI-Compatible API servers for KAT-Coder-V2.5-Dev model. #### SGLang SGLang is a fast serving framework for large language models and vision language models. sglang>=0.5.10 is recommended for KAT-Coder-V2.5-Dev, which can be installed using the following command in a fresh environment: ```shell uv pip install sglang[all] ``` The following will create API endpoints at `http://localhost:8000/v1`: > **Note:** This open-weight release ships only the language-model weights (no vision tower). If your SGLang version attempts to build the multimodal/vision components at load time, startup may fail on missing vision weights; in that case, run with the version's text/language-model-only option (see `python -m sglang.launch_server --help`). - Standard Version: The following command can be used to create an API endpoint with maximum context length 262,144 tokens using tensor parallel on 8 GPUs. ```shell python -m sglang.launch_server \ --model-path Kwaipilot/KAT-Coder-V2.5-Dev \ --port 8000 \ --tp-size 8 \ --mem-fraction-static 0.8 \ --context-length 262144 \ --reasoning-parser qwen3 ``` - Tool Use: To support tool use, you can use the following command. ```shell python -m sglang.launch_server \ --model-path Kwaipilot/KAT-Coder-V2.5-Dev \ --port 8000 \ --tp-size 8 \ --mem-fraction-static 0.8 \ --context-length 262144 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder ``` #### vLLM vLLM is a high-throughput and memory-efficient inference and serving engine for LLMs. vllm>=0.19.0 is recommended for KAT-Coder-V2.5-Dev, which can be installed using the following command in a fresh environment: ```shell uv pip install vllm --torch-backend=auto ``` The following will create API endpoints at `http://localhost:8000/v1`: > **Note:** This open-weight release ships only the language-model weights, so the `--language-model-only` flag is **required**. It tells vLLM to skip the vision encoder and multimodal profiling; without it, vLLM attempts to initialize vision-tower weights that are not present in the checkpoint and startup fails. - Standard Version: The following command can be used to create an API endpoint with maximum context length 262,144 tokens using tensor parallel on 8 GPUs. ```shell vllm serve Kwaipilot/KAT-Coder-V2.5-Dev \ --port 8000 \ --tensor-parallel-size 8 \ --max-model-len 262144 \ --reasoning-parser qwen3 \ --language-model-only ``` - Tool Call: To support tool use, you can use the following command. ```shell vllm serve Kwaipilot/KAT-Coder-V2.5-Dev \ --port 8000 \ --tensor-parallel-size 8 \ --max-model-len 262144 \ --reasoning-parser qwen3 \ --enable-auto-tool-choice \ --tool-call-parser qwen3_coder \ --language-model-only ``` #### KTransformers KTransformers is a flexible framework for experiencing cutting-edge LLM inference optimizations with CPU-GPU heterogeneous computing. For running KAT-Coder-V2.5-Dev with KTransformers, see the KTransformers Deployment Guide. #### Hugging Face Transformers Hugging Face Transformers contains a lightweight server which can be used for quick testing and moderate load deployment. The latest transformers is required for KAT-Coder-V2.5-Dev. Installing `accelerate` is also required for multi-GPU (sharded) loading: ```shell pip install "transformers[serving]" accelerate ``` Then, run transformers serve to launch a server with API endpoints at `http://localhost:8000/v1`; it will place the model on accelerators if available: ```shell transformers serve Kwaipilot/KAT-Coder-V2.5-Dev --port 8000 ``` ## Using KAT-Coder-V2.5-Dev via the Chat Completions API The chat completions API is accessible via standard HTTP requests or OpenAI SDKs. Here, we show examples using the OpenAI Python SDK. Before starting, make sure it is installed and the API key and the API base URL is configured, e.g.: ```shell pip install -U openai # Set the following accordingly export OPENAI_BASE_URL="http://localhost:8000/v1" export OPENAI_API_KEY="EMPTY" ``` ### Text-Only Input ```python from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ {"role": "user", "content": "Type \"I love KAT-Coder-V2.5-Dev\" backwards"}, ] chat_response = client.chat.completions.create( model="Kwaipilot/KAT-Coder-V2.5-Dev", messages=messages, max_tokens=81920, temperature=1.0, top_p=0.95, presence_penalty=1.5, extra_body={ "top_k": 20, }, ) print("Chat response:", chat_response) ``` ## Instruct (or Non-Thinking) Mode KAT-Coder-V2.5-Dev will think by default before response. You can obtain direct response from the model without thinking by configuring the API parameters. For example, ```python from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [ {"role": "user", "content": "Write a Python function that returns the n-th Fibonacci number."}, ] chat_response = client.chat.completions.create( model="Kwaipilot/KAT-Coder-V2.5-Dev", messages=messages, max_tokens=32768, temperature=0.7, top_p=0.8, presence_penalty=1.5, extra_body={ "top_k": 20, "chat_template_kwargs": {"enable_thinking": False}, }, ) print("Chat response:", chat_response) ``` ## Preserve Thinking By default, only the thinking blocks generated in handling the latest user message is retained, resulting in a pattern commonly as interleaved thinking. KAT-Coder-V2.5-Dev has been additionally trained to preserve and leverage thinking traces from historical messages. You can enable this behavior by setting the preserve_thinking option: ```python from openai import OpenAI # Configured by environment variables client = OpenAI() messages = [...] chat_response = client.chat.completions.create( model="Kwaipilot/KAT-Coder-V2.5-Dev", messages=messages, max_tokens=32768, temperature=0.7, top_p=0.8, presence_penalty=1.5, extra_body={ "top_k": 20, "chat_template_kwargs": {"preserve_thinking": True}, }, ) print("Chat response:", chat_response) ``` This capability is particularly beneficial for agent scenarios, where maintaining full reasoning context can enhance decision consistency and, in many cases, reduce overall token consumption by minimizing redundant reasoning. Additionally, it can improve KV cache utilization, optimizing inference efficiency in both thinking and non-thinking modes. ## Processing Ultra-Long Texts KAT-Coder-V2.5-Dev natively supports context lengths of up to 262,144 tokens. For long-horizon tasks where the total length (including both input and output) exceeds this limit, we recommend using RoPE scaling techniques to handle long texts effectively, e.g., YaRN. YaRN is currently supported by several inference frameworks, e.g., transformers, vllm, ktransformers and sglang. In general, there are two approaches to enabling YaRN for supported frameworks: - Modifying the model configuration file: In the config.json file, change the rope_parameters fields in text_config to: ```json { "mrope_interleaved": true, "mrope_section": [ 11, 11, 10 ], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144 } ``` - Passing command line arguments: For vllm, you can use ```shell VLLM_ALLOW_LONG_MAX_MODEL_LEN=1 vllm serve ... --hf-overrides '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --max-model-len 1010000 ``` For sglang and ktransformers, you can use ```shell SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1 python -m sglang.launch_server ... --json-model-override-args '{"text_config": {"rope_parameters": {"mrope_interleaved": true, "mrope_section": [11, 11, 10], "rope_type": "yarn", "rope_theta": 10000000, "partial_rotary_factor": 0.25, "factor": 4.0, "original_max_position_embeddings": 262144}}}' --context-length 1010000 ``` ## Citation If you find our work helpful, feel free to give us a cite. ```bibtex @misc{katcoder_v25_2026, title={{KAT-Coder-V2.5 Technical Report}}, author={{KwaiKAT Team}}, year={2026}, month={July}, eprint={2607.05471}, archivePrefix={arXiv}, primaryClass={cs.AI}, url={https://arxiv.org/pdf/2607.05471} } ```