Text Generation
GGUF
Safetensors
qwen3_5_moe
Mixture of Experts
apex
quantized
imatrix
torch-imatrix
mtp
speculative-decoding
coding
agentic-coding
llama.cpp
conversational
Instructions to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF # Run inference directly in the terminal: llama cli -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF # Run inference directly in the terminal: ./llama-cli -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Use Docker
docker model run hf.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
- LM Studio
- Jan
- vLLM
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
- Ollama
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Ollama:
ollama run hf.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
- Unsloth Studio
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF to start chatting
- Pi
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Docker Model Runner:
docker model run hf.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
- Lemonade
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Run and chat with the model
lemonade run user.KAT-Coder-V2.5-Dev-MTP-APEX-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
| license: apache-2.0 | |
| base_model: Kwaipilot/KAT-Coder-V2.5-Dev | |
| base_model_relation: quantized | |
| pipeline_tag: text-generation | |
| library_name: gguf | |
| tags: | |
| - gguf | |
| - moe | |
| - apex | |
| - quantized | |
| - imatrix | |
| - torch-imatrix | |
| - mtp | |
| - speculative-decoding | |
| - qwen3_5_moe | |
| - coding | |
| - agentic-coding | |
| - llama.cpp | |
| # KAT-Coder-V2.5-Dev β APEX GGUF with MTP head | |
| MoE-aware, mixed-precision **APEX** quantization of | |
| [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev), an | |
| agentic-coding fine-tune on the Qwen3.6-35B-A3B backbone (`qwen3_5_moe`: 40 layers, 256 routed | |
| + 1 shared expert, hybrid GatedDeltaNet linear-attention with periodic full attention). | |
| Kwaipilot reports SWE-bench Verified 69.40 for this size class, against Qwen3.6-35B-A3B's 64.40. | |
| **What makes this build different: it has a working MTP head.** KAT-Coder ships | |
| `mtp_num_hidden_layers: 0` β no multi-token-prediction head at all, so no speculative decoding | |
| is possible out of the box, and that is true of the vendor release and of every other quant of | |
| this model I am aware of. This repo transplants Qwen3.6-35B-A3B's own trained MTP head onto | |
| KAT-Coder and quantizes the result. Measured below: **2.03x on a hard agentic-coding suite, | |
| with correctness unchanged.** | |
| ## Recommended speculative settings | |
| ``` | |
| llama-server -m KAT-Coder-V2.5-Dev-MTP-APEX-i-quality-v2.gguf \ | |
| --spec-type draft-mtp --spec-draft-n-max 3 | |
| ``` | |
| Full draft-depth x temperature surface, measured on `-v2`. One long-context prompt (~82 KB), | |
| `max_tokens=400`, 3 seeds per cell, median tok/s, `top_p 0.95 top_k 20`, ctx 32768, build | |
| `6d05498`. Speedup is against the no-drafter baseline **at the same temperature**: | |
| | depth | T=0 (greedy) | T=0.25 | T=0.5 | T=0.75 | T=1.0 | | |
| |---|---:|---:|---:|---:|---:| | |
| | no drafter | 59.3 | 59.4 | 59.4 | 59.4 | 59.2 | | |
| | `n-max 1` | 71.3 (1.20x) | 75.4 (1.27x) | 75.4 (1.27x) | 73.7 (1.24x) | 71.9 (1.21x) | | |
| | `n-max 2` | 76.5 (1.29x) | 76.2 (1.28x) | 77.7 (1.31x) | 75.2 (1.27x) | 71.9 (1.21x) | | |
| | **`n-max 3`** | **85.7 (1.44x)** | **78.7 (1.33x)** | **78.8 (1.33x)** | **77.0 (1.30x)** | **74.9 (1.26x)** | | |
| | `n-max 5` | 68.4 (1.15x) | 61.9 (1.04x) | 66.5 (1.12x) | 66.4 (1.12x) | 62.0 (1.05x) | | |
| | `n-max 8` | 51.2 (0.86x) | 55.9 (0.94x) | 53.8 (0.91x) | 54.5 (0.92x) | 48.6 (0.82x) | | |
| Draft acceptance for the same cells: | |
| | depth | T=0 | T=0.25 | T=0.5 | T=0.75 | T=1.0 | | |
| |---|---:|---:|---:|---:|---:| | |
| | `n-max 1` | 0.814 | 0.820 | 0.842 | 0.791 | 0.736 | | |
| | `n-max 2` | 0.685 | 0.692 | 0.696 | 0.631 | 0.620 | | |
| | `n-max 3` | 0.692 | 0.588 | 0.605 | 0.586 | 0.552 | | |
| | `n-max 5` | 0.455 | 0.415 | 0.441 | 0.414 | 0.388 | | |
| | `n-max 8` | 0.290 | 0.297 | 0.308 | 0.305 | 0.260 | | |
| **Depth 3 wins at every temperature**, and the falloff past it is steep: acceptance decays roughly | |
| geometrically with depth while draft compute grows linearly, so at `n-max 8` the wasted work | |
| exceeds the gain and it runs **slower than no drafter at all** (0.82-0.94x). If you tune one knob, | |
| tune this one. | |
| The no-drafter row is flat across temperature (59.2-59.4), which is the control: all the variation | |
| above it is acceptance, not decode speed. | |
| Absolute tok/s here is lower than the from-scratch figures further down (59.3 vs ~70.8 baseline) | |
| because this sweep uses a long-context prompt β decode slows as context grows. The *ratios* are | |
| the transferable part. | |
| Every number in both tables is throughput. Correctness is unaffected by construction: the target | |
| model verifies each drafted token against its own computation, so a rejected draft never reaches | |
| the output. Speculative decoding cannot change what the model says, only how fast it says it. | |
| ## Files | |
| | file | size | what it is | | |
| |---|---:|---| | |
| | `KAT-Coder-V2.5-Dev-MTP-APEX-i-quality-v2.gguf` | 19.24 GiB | **recommended default** β supersedes the file below. Same tier, 56 MiB smaller, better PPL. See "v2" | | |
| | `KAT-Coder-V2.5-Dev-APEX-dynamic-v2.gguf` | 11.91 GiB | supersedes the dynamic file below; better PPL on both eval corpora | | |
| | `KAT-Coder-V2.5-Dev-MTP-APEX-i-quality.gguf` | 20.72 GB | superseded by `-v2`; kept so earlier results stay reproducible | | |
| | `KAT-Coder-V2.5-Dev-APEX-dynamic.gguf` | 11.86 GiB | superseded by `-v2`; role-aware allocation targeting a 16GB card, no MTP head | | |
| | `kat-coder.imatrix` | 192 MB | importance matrix used to build both; standard GGUF format, works with any `llama-quantize --imatrix` | | |
| | `model-00014-of-mtp.safetensors` | 1.69 GB | the bf16 MTP head shard, if you want to do the transplant yourself | | |
| | `config.json`, `model.safetensors.index.json` | β | already merged to reference all 14 shards | | |
| | `TOOL_CALLING.md` | β | reproducible chained tool-calling transcript | | |
| Pick `MTP-APEX-i-quality` unless you are VRAM-constrained, in which case `APEX-dynamic` is | |
| the best quality available at its size. To run either you need only the `.gguf`. The safetensors shard and index are there for | |
| anyone who wants to redo the splice against the original checkpoint: drop those three files | |
| into a local copy of | |
| [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) | |
| alongside its own 13 `model-*-of-00013.safetensors` shards. KAT-Coder's original shards are | |
| untouched and byte-identical; this is a drop-in addition, not a re-upload of the 65 GB | |
| checkpoint. | |
| ## The MTP head β honest framing | |
| This is the **naive, un-fine-tuned transplant**. The head's weights are copied straight from | |
| stock Qwen3.6 with zero training on KAT-Coder's own hidden-state distribution. It is | |
| architecturally clean rather than a hack β hidden_size, expert count, and attention/MoE shapes | |
| match exactly, because KAT-Coder is itself built on the Qwen3.6-35B-A3B backbone β but a | |
| fine-tuning pass (freezing KAT-Coder's 40 layers, training only the ~850M-parameter head) is a | |
| planned follow-up. These numbers are the baseline before that. | |
| ### Draft-acceptance rate | |
| `llama-server --spec-type draft-mtp --spec-draft-n-max 2`, 5 real from-scratch coding prompts, | |
| `n_predict=400` each: | |
| | prompt | draft tokens | accepted | rate | tok/s | | |
| |---|---:|---:|---:|---:| | |
| | binary search tree (insert/delete/traverse) | 286 | 255 | 89.2% | 105.5 | | |
| | CSV parser (stdlib only) | 114 | 84 | 73.7% | 94.6 | | |
| | retry decorator w/ exponential backoff | 188 | 152 | 80.9% | 100.6 | | |
| | merge two sorted linked lists | 74 | 58 | 78.4% | 97.6 | | |
| | thread-safe LRU cache | 360 | 218 | 60.6% | 85.9 | | |
| | **overall** | **1022** | **767** | **75.1%** | β | | |
| The same quant tier without any MTP head runs ~70.8 tok/s solo. Acceptance tracks | |
| predictability: BST traversal boilerplate hits 89%, the LRU cache with genuine locking and | |
| eviction logic is hardest at 61%. | |
| ### Agentic coding benchmarks | |
| Two task suites run through `opencode`'s real tool-calling loop, MTP build vs the same quant | |
| without MTP. **Correctness was identical β 100% on both suites either way** (59/59 and 28/28 | |
| test cases), which is what the theory predicts: the MTP head only drafts, the main model always | |
| verifies against its own computation, so a bad draft is rejected and never reaches the output. | |
| Only speed moves. | |
| | hard-tier task | no MTP | with MTP | speedup | | |
| |---|---:|---:|---:| | |
| | batch_scheduler | 22.5s | 20.8s | 1.08x | | |
| | buddy_allocator | 123.5s | 63.5s | 1.94x | | |
| | circuit_breaker | 50.7s | 14.0s | 3.62x | | |
| | expr_parser_advanced | 54.7s | 21.1s | 2.59x | | |
| | inventory_fix_and_extend | 16.4s | 12.3s | 1.33x | | |
| | **total** | **267.8s** | **131.7s** | **2.03x** | | |
| | tier-4 task | no MTP | with MTP | speedup | | |
| |---|---:|---:|---:| | |
| | lru_ttl_cache | 20.5s | 25.2s | 0.81x | | |
| | rate_limiter_fix_and_extend | 19.7s | 12.1s | 1.63x | | |
| | weighted_interval_scheduling | 22.5s | 17.4s | 1.29x | | |
| | **total** | **62.7s** | **54.7s** | **1.15x** | | |
| The single slower result (`lru_ttl_cache`, 0.81x) is most likely run-to-run variance in how | |
| much iteration that attempt needed rather than a systematic MTP cost β every other task | |
| improved. | |
| ## v2 | |
| Two allocation defects were found and fixed. Both were reported or surfaced after the original | |
| upload; the `-v2` files are the corrected builds and the originals are retained so earlier numbers | |
| stay reproducible. | |
| **1. State coefficients were quantised.** `blk.N.ssm_alpha` and `blk.N.ssm_beta` β 60 tensors, | |
| [2048,32], 3.93 M params total β were absent from the allocation recipe and silently inherited the | |
| base quant type (Q4_K here, Q3_K in the dynamic tier). They now ship **F32**. These parameterise a | |
| gate: `beta = sigmoid(ssm_beta . x)`, `alpha = softplus(ssm_alpha . x + dt)`, and the result | |
| multiplies the state path, so error there compounds along the sequence rather than adding to one | |
| layer's output. Cost of pinning them: ~14 MB on a 19 GiB file. Reported by | |
| [@anjeysapkovski](https://huggingface.co/Myric/KAT-Coder-V2.5-Dev-MTP-APEX-GGUF/discussions/2). | |
| **2. The MTP block was never allocated.** `blk.40` β the transplanted head, 805 M params of | |
| experts plus `nextn.eh_proj` β was also missing from the recipe and took the base type by accident, | |
| landing four expert tensors on Q6_K. A drafter's precision buys acceptance rate, never output | |
| quality, so those were bits spent where the return is zero by construction. `blk.40` is now pinned | |
| Q4_K throughout. Measured effect on acceptance: none (within run-to-run spread at every depth), | |
| and the file is 56 MiB smaller. | |
| | tier | published | v2 | code-heavy PPL | wiki.test PPL | | |
| |---|---:|---:|---:|---:| | |
| | MTP i-quality | 19,765 MiB | **19,709 MiB** | 2.1247 -> **2.1231** | 7.1169 -> **7.1128** | | |
| | dynamic | 12,141 MiB | 12,193 MiB | 2.4479 -> **2.4415** | 7.5435 -> **7.5160** | | |
| Both v2 files pass a recipe-vs-artifact audit: every requested tensor type honoured, and zero | |
| tensors quantised that the recipe never asked about (the original files had 71 such tensors). | |
| ## The imatrix | |
| `llama-imatrix` is impractical on this architecture: GatedDeltaNet linear-attention is a serial | |
| state-space recurrence, and the imatrix collection callback breaks the GPU path, falling back to | |
| a single CPU thread. `kat-coder.imatrix` was instead generated with a band-serialized PyTorch | |
| forward pass over a general text/code corpus. It is a standard GGUF-format importance matrix | |
| (`in_sum2` + `counts` per tensor), bit-compatible with `llama-quantize --imatrix`, so it is | |
| directly reusable if you want to build your own tiers. | |
| One checkpoint quirk worth knowing if you do: KAT-Coder stores MoE experts as fully separate | |
| per-expert tensors (`mlp.experts.{e}.{gate,up,down}_proj.weight` Γ 256) rather than the fused | |
| `gate_up_proj`/`down_proj` layout the HF module expects. Both imatrix generation and GGUF | |
| conversion have to fuse them. | |
| ## Allocation | |
| APEX structural mode: routed experts are bucketed into edge / near / middle depth bands (edge = | |
| first and last layers, generally most load-bearing) with decreasing precision toward the middle. | |
| For i-quality that is routed edge β Q6_K, near β Q4_K, middle β Q3_K. Attention, shared-expert, | |
| and router tensors stay at higher fixed precision (Q6_KβQ8_0). | |
| ### Measured quality | |
| Wikitext perplexity across the tiers that were built. The two marked *hosted* are in this repo; | |
| the other two were measured for comparison and are not uploaded: | |
| | tier | file size | bits/weight | wikitext PPL | vs bf16 | | | |
| |---|---:|---:|---:|---:|---| | |
| | bf16 (reference) | 64.61 GiB | 16.01 | 7.028 | β | | | |
| | APEX-i-quality | 19.23 GiB | 4.65 | 7.117 | +1.27% | recipe used by the hosted MTP build | | |
| | **APEX-dynamic** | 12.13 GiB | 2.94 | 7.543 | +7.33% | **hosted** | | |
| | IQ3_XXS (stock, imatrix-guided) | 12.68 GB | ~2.98 | 7.881 | +12.13% | stock baseline, for comparison | | |
| Two things worth drawing out. i-quality lands within 1.3% of full bf16 at under a third of the | |
| size. And **APEX-dynamic beats the stock IQ3_XXS baseline on both axes at once** β smaller | |
| (12.13 vs 12.68 GiB) *and* lower perplexity (7.543 vs 7.881) β which is why it is the one | |
| hosted alongside the flagship rather than the stock tier. | |
| That result is cleaner than the same method got on Laguna-XS, where it traded size for a small | |
| PPL cost. Whether role/depth-aware allocation beats a well-tuned stock quant appears to depend | |
| on the architecture and calibration, and is not something to assume without measuring. | |
| The hosted MTP build uses the i-quality recipe; its file is slightly larger than the 19.23 GiB | |
| row because it carries the MTP head as well. | |
| ## Attribution | |
| - KAT-Coder base model: **Kwaipilot** β [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) (Apache-2.0). | |
| - MTP head weights: **Qwen** β [Qwen/Qwen3.6-35B-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) (Apache-2.0), transplanted unmodified. | |
| - APEX recipe & toolkit: **LocalAI** β [localai-org/apex-quant](https://github.com/localai-org/apex-quant). | |
| - Quantization engine: **llama.cpp** ([ggml-org](https://github.com/ggml-org/llama.cpp)). | |
| Unofficial community quantization and experiment; not affiliated with or endorsed by Kwaipilot | |
| or Qwen. | |