--- license: apache-2.0 base_model: - Kwaipilot/KAT-Coder-V2.5-Dev pipeline_tag: image-text-to-text library_name: transformers tags: - code - agent - agentic-coding - moe - vision-language - vllm language: - en - zh --- # KAT-Coder-V2.5-Dev-VL A **vision-enabled** build of [`Kwaipilot/KAT-Coder-V2.5-Dev`](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev). The upstream release declares a full multimodal architecture — `Qwen3_5MoeForConditionalGeneration`, a `vision_config`, a `Qwen3VLProcessor` preprocessor, image/video token IDs, and a chat template that handles images — but ships **no vision weights**, so it can only run text-only and requires `--language-model-only` under vLLM. This build restores the vision tower so the declared architecture actually works end-to-end. **No language-model weights were modified**: text behaviour is bit-identical to upstream. > **Read the [Vision capability](#vision-capability-read-this) section before using this for anything > visual.** Coarse visual understanding works; fine-grained reading (OCR) does not. | | | |---|---| | Parameters | 35 B total / 3 B active (MoE, 256 experts, top-8) | | Layers | 40 (30 linear-attention + 10 full-attention, hybrid) | | Vision tower | 27 blocks, hidden 1152 → 2048, patch 16, spatial merge 2 (446.6 M params) | | Context | 262 144 | | Precision | bfloat16 | | Size on disk | 70.2 GB | | License | Apache-2.0 | A quantized build for Blackwell GPUs is available at [`beyoru/KAT-Coder-V2.5-Dev-VL-Flash`](https://huggingface.co/beyoru/KAT-Coder-V2.5-Dev-VL-Flash). --- ## Vision capability — READ THIS The vision tower was never co-trained with this checkpoint's language model, which was heavily RL-tuned for agentic coding. The vision→LM projection therefore carries **coarse features reliably and fine features unreliably**. Measured on a served endpoint, greedy decoding: | probe | score | notes | |---|---|---| | Dominant colour (4 solid-colour images) | **4 / 4** | red, blue, green, yellow — all correct | | Shape (circle / square / triangle) | **3 / 3** | all correct | | Scene description (real icon) | plausible | correctly identified colour and rough form | | **Text/number reading (OCR)** | **0 / 4** | `CAT`→`CCT`, `42`→`48`, `HELLO`→`Hiro`, `7492`→`ZOE` | The OCR failures are all **near-misses onto visually similar glyphs**, which is the signature of a tower that resolves stroke-level shape but not glyph identity. In practice: - ✅ Reasonable for: "is there a chart here", colour/layout/shape questions, coarse scene gist, "does this screenshot show an error dialog or a code editor". - ❌ **Not usable for**: reading code from screenshots, OCR, reading UI labels, dense-document understanding, chart-value extraction, anything where a wrong character changes the answer. If your workload is text-only, prefer upstream `Kwaipilot/KAT-Coder-V2.5-Dev` — this build offers no advantage there (the LM weights are identical) and costs 0.9 GB extra VRAM. --- ## Ready-made configs | file | what | status | |---|---|---| | [`docker-compose.vllm.yaml`](./docker-compose.vllm.yaml) | vLLM, 1 GPU, vision enabled | **verified** | | [`docker-compose.sglang-flashqla.yaml`](./docker-compose.sglang-flashqla.yaml) | SGLang + FlashQLA GDN prefill | untested | FlashQLA applies here because 30 of the 40 layers are `linear_attention` (Gated DeltaNet) — exactly the layers its TileLang kernel accelerates during prefill. It is opt-in via a single flag (`--linear-attn-prefill-backend flashqla`); the default path is byte-identical to stock SGLang. **Windowed-MTP does not apply to this model.** `mtp_num_hidden_layers = 0`, so the checkpoint carries no MTP/NEXTN head to self-speculate with. That technique needs `Qwen/Qwen3.6-35B-A3B` (mtp = 1). ## Serving ### vLLM (recommended, verified) Verified working configuration: | | | |---|---| | Image | `vllm/vllm-openai:v0.26.0` | | GPU | 1 × H200 (143 GB), tensor-parallel 1 | | Load time | ~2 min cold (14 shards, 65.4 GiB) | | Attention | FlashAttention 3 (LM), FLASH_ATTN (ViT), FlashInfer GDN for linear attention | | MoE backend | TRITON (auto-selected) | ```bash docker run -d --name kat-vl \ --gpus '"device=0"' --ipc=host \ -v /path/to/models:/models \ -p 8000:8000 \ vllm/vllm-openai:v0.26.0 \ --model /models/KAT-Coder-V2.5-Dev-VL \ --served-model-name kat-vl \ --max-model-len 32768 \ --max-num-batched-tokens 8192 \ --gpu-memory-utilization 0.90 \ --trust-remote-code ``` > **`--max-num-batched-tokens 8192` is required.** This is a hybrid attention + Gated-DeltaNet > architecture. With prefix caching on (the default in recent vLLM), vLLM forces > `mamba_cache_mode='align'` and raises the attention block size to **2096** so that "attention page > size >= mamba page size". The default `max_num_batched_tokens` is **2048**, so startup dies with > `AssertionError: In Mamba cache align mode, block_size (2096) must be <= max_num_batched_tokens > (2048)`. Any value >= 2096 works. Or without Docker (`vllm>=0.26.0`): ```bash uv pip install vllm --torch-backend=auto vllm serve beyoru/KAT-Coder-V2.5-Dev-VL \ --served-model-name kat-vl \ --max-model-len 32768 \ --gpu-memory-utilization 0.90 \ --trust-remote-code ``` > **Do NOT pass `--language-model-only`.** That flag is required for the upstream text-only release; > passing it here disables the vision tower and you get the upstream model back. **Memory.** bf16 weights are 70.2 GB, so a single 80 GB card cannot hold them plus a usable KV cache. Use one 141/143 GB card, or `--tensor-parallel-size 2` on 2 × 80 GB. Raise `--max-model-len` only after checking free memory — the model's native 262 144 context needs a large KV allocation. **Long context.** For contexts beyond the native window, apply YaRN exactly as documented for `Qwen/Qwen3.6-35B-A3B`; the RoPE parameters here are unchanged from upstream. Note that all major frameworks implement *static* YaRN, so enabling it can hurt short-prompt quality — turn it on only when you actually need long contexts. ### SGLang `sglang>=0.5.10`. Because this build **does** contain the vision tower, do **not** pass the text/language-model-only flag that the upstream release requires. ### Transformers Works out of the box — verified on `transformers==5.10.1`. ```python import torch from transformers import AutoModelForImageTextToText, AutoProcessor model = AutoModelForImageTextToText.from_pretrained( "beyoru/KAT-Coder-V2.5-Dev-VL", dtype=torch.bfloat16, device_map="auto") processor = AutoProcessor.from_pretrained("beyoru/KAT-Coder-V2.5-Dev-VL") ``` Note: this checkpoint stores MoE experts in the **per-expert (split)** layout (`...mlp.experts.{0..255}.{gate,up,down}_proj.weight`), inherited unchanged from upstream, while transformers 5.x uses a **fused** layout internally. The conversion is automatic at load time — no manual step is needed, and no weights are lost. --- ## Usage The API is OpenAI-compatible. This is a **thinking model**: it emits a reasoning block terminated by `` before the answer, so give it a generous `max_tokens` (≥256) or you will only ever receive truncated reasoning. Strip everything up to and including `` to get the final answer, or disable thinking with `chat_template_kwargs: {"enable_thinking": false}`. ```python from openai import OpenAI client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY") # text r = client.chat.completions.create( model="kat-vl", messages=[{"role": "user", "content": "Write a Python function to reverse a linked list."}], max_tokens=512, temperature=0.0, ) print(r.choices[0].message.content.split("")[-1].strip()) # image + text r = client.chat.completions.create( model="kat-vl", messages=[{"role": "user", "content": [ {"type": "image_url", "image_url": {"url": "data:image/png;base64,"}}, {"type": "text", "text": "Describe this image in one sentence."}, ]}], max_tokens=512, temperature=0.0, ) print(r.choices[0].message.content.split("")[-1].strip()) ``` Sampling parameters follow upstream: `temperature=1.0, top_p=0.95, top_k=20` for open-ended generation; `temperature=0.0` for deterministic coding tasks. --- ## What this build is and is not - **Is:** the upstream language model, byte-for-byte, plus a working vision tower and a patched weight index so the declared multimodal architecture loads. - **Is not:** a multimodally trained model. The vision path was not fine-tuned against this LM. Treat its visual outputs as coarse hints, never as ground truth. - No distillation, no re-training, no quantization, no merging of language-model weights. ## Acknowledgements - [`Kwaipilot/KAT-Coder-V2.5-Dev`](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) — the base model and all language-model weights (Apache-2.0). - [`Qwen/Qwen3.6-35B-A3B`](https://huggingface.co/Qwen/Qwen3.6-35B-A3B) — the architecture this model derives from (Apache-2.0). Both are Apache-2.0; this build is released under the same licence.