--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation base_model: Kwaipilot/KAT-Coder-V2.5-Dev base_model_relation: finetune datasets: - TorpedoSoftware/Roblox-Luau-Reasoning-v1.0 - TorpedoSoftware/RobloxQA-v2.0 - TorpedoSoftware/RobloxQA-v1.0 - TorpedoSoftware/LuauLeetcode - Pinkstack/luaucoder-instructions-v3-SFT - khtsly/luau-stack-hq - Roblox/luau_corpus tags: - code - luau - roblox - qwen3.5 - moe - nvfp4 - blackwell - fine-tune - comulative language: - en - zh --- # KAT-Coder-V2.5_JKL-Luau-NVFP4 **Roblox Luau–specialized fine-tune** of [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev), released as **NVFP4** (W4A4 compressed-tensors) for efficient inference on **NVIDIA Blackwell** GPUs. | | | |--|--| | **Base model** | [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev) (~35B-A3B Qwen3.5-MoE, text/language release) | | **Fine-tune** | Supervised LoRA SFT on a Roblox / Luau mix (1 epoch) | | **This artifact** | Merged bf16 weights → **NVFP4** post-training quantization | | **Trained by** | [@dylanjkl](https://huggingface.co/dylanjkl) at **Comulative Limited** (UK) | | **Hardware** | **1× NVIDIA RTX PRO 6000 Blackwell 96GB** (sm_120) | > This is **not** an official Kwaipilot release. It is an independent specialty fine-tune and quant by Comulative Limited. --- ## Intended uses **Primary:** local / self-hosted **Roblox Luau coding assistant** and **task executor**: - Idiomatic Luau modules (`--!strict`, services, remotes, DataStore patterns) - Roblox client/server split, RemoteEvents / RemoteFunctions, replication-aware snippets - Code completion / continuation on existing Luau scripts - Roblox engine Q&A grounded in specialty SFT data **Recommended deployment pattern:** use a stronger **planning / review** model for architecture and security, and this model as a **fast local executor** for Luau implementation (low latency, private weights, Blackwell-friendly NVFP4). **Not intended for:** unsupervised production game economy / anti-cheat design without human review; non-Roblox general agenting as a drop-in frontier replacement; vision/multimodal tasks (language-only lineage). --- ## What we did (training pipeline) ### 1. Base Started from **KAT-Coder-V2.5-Dev** (Qwen3.5 MoE coding model, language-only open weights; vision tower declared in config but not shipped upstream). ### 2. Domain SFT data mix Built a chat-formatted SFT mix (**42,302** train rows after filtering/dedup; ~57M tokens) from public **Hugging Face datasets** (local mirror under `data/`). #### Hugging Face datasets used in training | Hub dataset | Role in mix | Approx. SFT rows (source tag) | |-------------|-------------|-------------------------------| | [TorpedoSoftware/Roblox-Luau-Reasoning-v1.0](https://huggingface.co/datasets/TorpedoSoftware/Roblox-Luau-Reasoning-v1.0) | Prompt → CoT + Luau code + explanation (`train`) | **14,840** `luau-reasoning` | | [Pinkstack/luaucoder-instructions-v3-SFT](https://huggingface.co/datasets/Pinkstack/luaucoder-instructions-v3-SFT) | Instruction SFT (filtered; cap ~8k quality rows) | **7,927** `pinkstack-sft` | | [khtsly/luau-stack-hq](https://huggingface.co/datasets/khtsly/luau-stack-hq) | Curated Luau corpus → fill-in / continuation tasks | **5,944** `stackhq-completion` | | [TorpedoSoftware/RobloxQA-v2.0](https://huggingface.co/datasets/TorpedoSoftware/RobloxQA-v2.0) | Engine/API knowledge from **`train`** (MCQ + direct QA variants) | **4,573** `robloxqa2-mcq` + **2,290** `robloxqa2-direct` | | [TorpedoSoftware/LuauLeetcode](https://huggingface.co/datasets/TorpedoSoftware/LuauLeetcode) | Algorithmic Luau problems (`train`) | **2,336** `luau-leetcode` | | [TorpedoSoftware/RobloxQA-v1.0](https://huggingface.co/datasets/TorpedoSoftware/RobloxQA-v1.0) | Older QA; **deduped against RobloxQA-v2 test** | **2,281** `robloxqa1-mcq` | | [Roblox/luau_corpus](https://huggingface.co/datasets/Roblox/luau_corpus) | Official Luau Data Sharing fragments → continuation (`train`) | **2,111** `luaucorpus-completion` | **Total train examples:** 42,302 (plus 400 held-out mix rows for training-time val). #### Evaluation dataset (not used for gradient updates) | Hub dataset | Use | |-------------|-----| | [TorpedoSoftware/RobloxQA-v2.0](https://huggingface.co/datasets/TorpedoSoftware/RobloxQA-v2.0) **`test`** (3,000 MCQ) | **Held out** for baseline / future eval only | #### Downloaded but **not** used in this SFT mix | Hub dataset | Notes | |-------------|--------| | [TorpedoSoftware/roblox-info-dump](https://huggingface.co/datasets/TorpedoSoftware/roblox-info-dump) | Roblox/Luau docs scrape present under `data/`; **not** mixed into the epoch-1 SFT JSONL | Formatting used the base model chat template, with **prompt tokens masked** (train on completions). Optional system prompts mixed in (~30% Roblox-assistant style). Decompiled-looking completion snippets filtered out. Pinkstack rows required code fences + minimum reasoning length. ### 3. LoRA fine-tune (1 epoch) | Hyperparameter | Value | |----------------|--------| | Method | LoRA (PEFT), bf16 base | | Rank / alpha | r=64, α=128, dropout 0.05 | | Targets | Attention (q/k/v/o), linear-attn projections, shared-expert MLP; **not** per-routed experts | | Context | 4096 | | Effective batch | 16 (microbatch 2 × grad accum 8) | | Epochs shipped here | **1** (stopped at step **2642 / 5284** of a 2-epoch schedule) | | Optim | AdamW fused, LR 1e-4 cosine, warmup 40 | | Hardware | 1× RTX PRO 6000 96GB, Windows 11 + WSL2 Ubuntu | A **full Trainer checkpoint** (adapter + optimizer) was frozen at epoch 1 so a second epoch can be resumed later if desired. This Hub repo is the **merged + NVFP4** product of that epoch-1 adapter, not the raw LoRA. ### 4. Merge & cleanup - Merged LoRA into base bf16 weights - **Stripped** randomly initialized `visual.*` keys so the checkpoint matches the language-only upstream release ### 5. NVFP4 quantization - Tooling: **llm-compressor** oneshot PTQ - Scheme: **NVFP4** (W4A4, `nvfp4-pack-quantized`) - Calibration: **256** samples @ 4096 from the Luau SFT mix - `moe_calibrate_all_experts=True` - Ignores: `lm_head`, visual towers, router gates, embeddings, linear-attn (see `recipe.yaml`) **Why NVFP4:** native-friendly **4-bit float** path for **Blackwell** inference stacks (e.g. vLLM on modern NVIDIA data-center / pro GPUs). Smaller footprint (~22GB weights here) vs full bf16 (~70GB class). --- ## Base model evaluation note Unmodified base **KAT-Coder-V2.5-Dev** on RobloxQA-v2.0 test (3000 questions), MMLU-style **log-prob forced choice**, bf16, HF Transformers: **Baseline: 87.60% (2628 / 3000)** — measured 2026-08-01. Post–fine-tune / post-NVFP4 RobloxQA numbers for *this* checkpoint may be published later; treat the above as the starting point the specialty run was designed to preserve or beat on knowledge while improving **Luau generation**. --- ## How to use ### Transformers / llm-compressor compressed load Follow current [llm-compressor](https://github.com/vllm-project/llm-compressor) / Transformers docs for **NVFP4 compressed-tensors** checkpoints. Ensure a stack that understands `quantization_config` with format `nvfp4-pack-quantized`. ### vLLM (Blackwell) Prefer a recent vLLM build with NVFP4 + Qwen3.5 MoE support. Language-only serving flags may still apply (upstream was language-only; config may still mention vision). Example pattern (versions change quickly—check your vLLM release notes): ```bash vllm serve Comulative/KAT-Coder-V2.5_JKL-Luau-NVFP4 \ --quantization modelopt_fp4 \ # plus whatever your vLLM build requires for NVFP4 / Qwen3.5 MoE ``` Use the model’s chat template (`chat_template.jinja` / tokenizer config). For agentic coding UIs, Qwen-style reasoning / tool parsers may apply depending on runtime. ### Prompting tips - Prefer **narrow tasks**: one module, one bugfix, one remote contract - State **client vs server**, remote names, and data shapes - Ask for `--!strict` Luau when you want typed modules - Review **security** (RemoteEvent trust, DataStore races) with a stronger planner model before shipping --- ## Training summary card ```text Base: Kwaipilot/KAT-Coder-V2.5-Dev (Qwen3.5-MoE ~35B-A3B) SFT: LoRA r=64, 1 epoch, Luau/Roblox mix ~42k examples Merge: bf16 merge + strip visual.* Quant: NVFP4 (llm-compressor), 256 calib samples, all MoE experts Trainer: @dylanjkl / Comulative Limited (UK) GPU: 1× NVIDIA RTX PRO 6000 Blackwell 96GB ``` --- ## Limitations & risks - **Domain SFT** can narrow style; may be weaker than base on non-Roblox agentic coding - **1 epoch** only in this release; second epoch optional, not included - **NVFP4** is a lossy quant vs bf16 merge—validate on your tasks - **No Studio execution loop** in training—always test Luau in Roblox - May produce insecure remote/DataStore patterns; **do not** deploy unreviewed - MoE + hybrid linear attention runtimes are still evolving—pin tested engine versions --- ## License Weights are a derivative of [Kwaipilot/KAT-Coder-V2.5-Dev](https://huggingface.co/Kwaipilot/KAT-Coder-V2.5-Dev). Unless otherwise required by the base model license, this distribution is provided under **Apache-2.0**. Review the base model card and license for any additional terms. --- ## Citation / attribution ```text Base model: Kwaipilot/KAT-Coder-V2.5-Dev Fine-tune & NVFP4 release: Comulative Limited (UK) / @dylanjkl Hub: Comulative/KAT-Coder-V2.5_JKL-Luau-NVFP4 Hardware: 1× NVIDIA RTX PRO 6000 Blackwell 96GB ``` --- ## Maintainers - **[@dylanjkl](https://huggingface.co/dylanjkl)** — Comulative Limited (United Kingdom) - Organization: **[Comulative](https://huggingface.co/Comulative)** For issues with *this* fine-tune/quant, contact the Comulative maintainers—not Kwaipilot—unless the bug is reproducible on the unmodified base.