--- license: apache-2.0 base_model: Kwaipilot/KAT-Coder-V2.5-Dev tags: - nvfp4 - compressed-tensors - gb10 - veloGB10 - code - moe --- # KAT-Coder-V2.5-Dev — NVFP4 for veloGB10 (NVIDIA DGX Spark / GB10) > **Original model:** [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) by the KwaiKAT team, > [Apache-2.0](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev/blob/main/LICENSE). This artifact contains > **modified weights** (NVFP4/FP8 quantization); the original model is © its authors. NVFP4/FP8-quantized variant of **KAT-Coder-V2.5-Dev**. **This quantized artifact exists first and foremost to run with the [veloGB10](https://github.com/sf-stav/veloGB10) inference engine** (`gb10_inference`) on NVIDIA DGX Spark / GB10 systems — it is produced, validated, and tuned for veloGB10's kernels. The weights use the standard HuggingFace `compressed-tensors` layout (`nvfp4-pack-quantized`), so this artifact can be used **for any purpose, with any framework that reads the format**. It is in particular fit and *proven* to work with veloGB10 — greedy-lossless speculative decoding (where applicable), two-node TP=2 serving, and the engine's full correctness test suite all pass on this artifact. ## Specifications | | | |---|---| | Base model | [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) | | Architecture | MoE hybrid (40 GatedDeltaNet/full-GQA layers), hidden 2048, 16 Q / 2 KV heads | | Attention | hybrid GatedDeltaNet (linear attention, fixed-size recurrent state) + periodic full GQA | | Experts | 256 routed experts, 8 active per token, plus shared expert | | MTP head | none in this checkpoint — speculative decoding uses n-gram draft or future MTP heads | | Context | 262,144 tokens | | Vocabulary | 248,320 | | Quantization | `nvfp4-mixed` — all GEMMs NVFP4 except GDN in/out projections (FP8); router kept in bf16 | | Format | `nvfp4-pack-quantized` (compressed-tensors compatible) | | On-disk size | ~19 GB | ## Running with veloGB10 Full instructions: **[veloGB10 README](https://github.com/sf-stav/veloGB10)** (build or download a release binary, then). The binary loads its PTX kernel artifacts from the current directory — run it from wherever the binary + `*.ptx` live. **Single user — maximum speed:** ```bash gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED --port=9000 \ --max-seq-len=32768 --max-batch=1 --max-tokens=4096 \ --prefix-cache=on --mtp=off --default-presence-penalty 1.5 ``` **Four concurrent users — maximum aggregate throughput:** ```bash gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED --port=9000 \ --max-seq-len=32768 --max-batch=4 --max-tokens=4096 \ --prefix-cache=on --default-presence-penalty 1.5 ``` **Two-node TP=2 (maximum speed and capacity):** On the second node (no model copy, no configuration — the head ships weights, config, and calibration at sync): ```bash gb10_inference --node --port 29500 ``` On the head node: ```bash GB10_TP_SHARD_MIXERS=1 gb10_inference --server --model-dir=/path/to/Kwaipilot-KAT-Coder-V2.5-Dev-NVFP4-MIXED \ --tp --nodes :29500 --port 9000 \ --max-seq-len=32768 --max-batch=1 --max-tokens=4096 \ --prefix-cache=on --default-presence-penalty 1.5 ``` ## Measured on veloGB10 (greedy, NVFP4) *Pending — single-node and TP=2 throughput numbers are being gathered and will be added here once the full gate suite completes. See the [veloGB10 README](https://github.com/sf-stav/veloGB10) for the latest measured tables across the supported model family.* ## License & credits - **Original model:** [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) — © KwaiKAT team, [Apache License 2.0](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev/blob/main/LICENSE). - **This artifact:** quantized from the original with veloGB10's offline quantizer (`nvfp4-mixed`). Weights are modified relative to the original; tokenizer and chat template are unchanged from the original model. - **Engine:** [veloGB10](https://github.com/sf-stav/veloGB10) (Apache-2.0).