⚡ Tals-coder-flash-02 (27B Ternary + Grafted Native MTP + Multimodal Vision)

The Fastest 27B Coding Model on Consumer Hardware

The Flash Lightning Speed

"Why wait 5 minutes for a reasoning model to finish an existential crisis in its chain-of-thought? Tals-coder-flash-02 goes at lightning speed with 60 tok/s on a budget 16GB GPU."

Proof Metrics

Hardware Testing Notice: All speed benchmarks and token generation speeds reported below were tested on an NVIDIA GeForce RTX 4060 Ti 16GB at stock factory settings without any overclocking temp 0.2 top p 0.85.


🏎️ Pure Speed: Breaking the Physical Bandwidth Limit

On paper, an NVIDIA RTX 4060 Ti 16GB has a narrow 128-bit memory bus with 288 GB/s bandwidth. Streaming an 8 GB 27B model on that bus physically caps standard autoregressive generation to ~25 tokens/sec.

Tals-coder-flash-02 shatters this barrier:

  • ⚡ 55–60 tokens/second sustained on RTX 4060 Ti 16GB (verified in production!).
  • 🚀 Up to 76.7 tokens/second peak burst during boilerplate and repetitive code blocks.
  • 💨 120–150+ tokens/second projected on high-bandwidth hardware (RTX 3090, 4090, A100).
  • 🎯 Zero Overthinking: No rambling monologues, no 2,000-token loops questioning itself. It cuts straight to clean, functional code.

🧬 Custom GGUF Quantization & Grafting

Tals-coder-flash-02 is a custom GGUF quantization of prism-ml/Ternary-Bonsai-2-27B-gguf:

  • Surgically Grafted Native MTP: We permanently grafted the 65th Multi-Token Prediction (MTP) draft head and unrotated token embeddings directly into the main weights (Tals-coder-flash-02.gguf). You do not need to juggle separate target and drafter model files—it's an all-in-one standalone file with native speculative decoding!
  • Multimodal Vision Projector: Includes the high-fidelity Tals-coder-flash-02-mmproj.gguf CLIP Q8_0 vision tower (0.63 GB) for analyzing UI designs, charts, diagrams, and debugging screenshots.
  • 100% VRAM Execution: Runs 27B parameters, 128,000 context, and multimodal vision completely on GPU (~13.8–14.1 GB / 16.0 GB total) with zero CPU/PCIe spillover.

📊 Benchmark & Performance

Tested on NVIDIA GeForce RTX 4060 Ti 16GB (PCIe 4.0 x8, 288 GB/s bandwidth, stock factory settings without overclocking):

Configuration Speculative Acceptance Generation Speed Context Window
Base Model Only (No MTP) N/A (0%) 24.8 – 25.8 tok/s Up to 262k
Tals-coder-flash-02 (MTP n_max=1) 77.8% 40.7 – 42.5 tok/s Up to 262k
Tals-coder-flash-02 (MTP n_max=2) (Optimal) 68.9% – 74.2% 53.0 – 65.5 tok/s Up to 262k
Boilerplate / Code Bursts 85.2% Up to 76.7 tok/s Up to 262k
Estimated RTX 4090 / 3090 (1,008 GB/s) 75% – 85% 120 – 155 tok/s Full 262k

📦 Model Files in Repository

Filename Size Description
Tals-coder-flash-02.gguf 8.25 GB Unified 27B Custom GGUF Quant with Grafted Native MTP Head (All 65 layers).
Tals-coder-flash-02-mmproj.gguf 0.63 GB High-fidelity CLIP Q8_0 multimodal vision projector.
bonsai2-pascal.patch 259 KB Source patch for custom Hadamard kernels & MTP timing.
flash.gif 5.05 MB High-speed Flash badge visual asset.

🛠️ Optimal llama.cpp Runtime Setup

Because this model uses ternary quantization in a Hadamard-rotated basis, use the optimized llama.cpp build with Hadamard kernel support (Prism ML / Pascal branch, or apply bonsai2-pascal.patch).

Critical Environment Variables (Mandatory)

Before starting the server, set these environment variables to enable Pascal CUDA kernels and prevent CPU busy-polling:

Windows (PowerShell / CMD):

set GGML_CUDA_ROWLANE=1
set GGML_CUDA_RL_N4_LONG=1
set CUDA_DEVICE_SCHEDULE=BLOCKING_SYNC

Linux (Bash):

export GGML_CUDA_ROWLANE=1
export GGML_CUDA_RL_N4_LONG=1
export CUDA_DEVICE_SCHEDULE=BLOCKING_SYNC

⚡ Recommended Run Configurations

1. Default Production Mode (128k Context + Grafted MTP + Vision)

Run the single unified model with vision on a 16GB GPU:

./llama-server \
  -m "Tals-coder-flash-02.gguf" \
  --mmproj "Tals-coder-flash-02-mmproj.gguf" \
  --mmproj-offload \
  --image-min-tokens 1024 \
  -ngl 99 \
  -c 128000 \
  -ctk q8_0 -ctv q4_0 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --load-mode mmap --reasoning-preserve \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --jinja \
  --host 127.0.0.1 --port 8080

2. High-Precision Coding Mode (64k Context + FP16 KV Cache)

Maximum precision for complex mathematical derivations:

./llama-server \
  -m "Tals-coder-flash-02.gguf" \
  --mmproj "Tals-coder-flash-02-mmproj.gguf" \
  -ngl 99 \
  -c 65536 \
  -ctk f16 -ctv f16 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --host 127.0.0.1 --port 8080

3. Extreme Context Mode (262k Native Context on 16GB VRAM)

Process entire codebases and large book-length documents with zero CPU offload:

./llama-server \
  -m "Tals-coder-flash-02.gguf" \
  -ngl 99 \
  -c 262144 \
  -ctk q4_0 -ctv q4_0 \
  -fa on -np 1 -t 4 -tb 4 --poll 0 \
  --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
  --backend-sampling --host 127.0.0.1 --port 8080

💻 OpenCode / Cline / Continue Configuration

Set your tool's API endpoint to:

  • Base URL: http://127.0.0.1:8080/v1
  • Model Name: tals-coder-flash-02
  • API Key: not-needed (any string)
  • Temperature: 0.0 – 0.2 (Deterministic coding & reasoning)

🙏 Credits & Acknowledgements

  • PrismML: The original creators of the Ternary-Bonsai-2-27B-gguf architecture and pioneering 1.58-bit ternary Hadamard quantization kernels.
  • Ukisai: For Swift-Bonsai-2, adapting and quantizing the ternary base weights.
  • killy369 / Kilian: For training and providing the native MTP speculative draft weights and 64k pruned vocabulary.
  • Qwen Team: For the underlying Qwen 3.5 architecture.

⚖️ Disclaimer & Responsible Use

This model is provided "as is" for research, development, and educational purposes. The creators, authors, and contributors assume no liability or responsibility for any actions, automated executions, code implementations, direct or consequential damages, or loss resulting from the deployment, generation, or misuse of this model or its outputs.

Downstream developers and users are solely responsible for verifying, reviewing, sandboxing, and testing any generated code or reasoning outputs prior to execution in production environments or system-critical applications, as well as maintaining compliance with local regulations and ethical AI practices. Please use responsibly.

Distributed under the Apache 2.0 license.

Downloads last month
784
GGUF
Model size
29B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for talxcc/Tals-coder-flash-02

Base model

Qwen/Qwen3.8-27B
Quantized
(30)
this model