Instructions to use talxcc/Tals-coder-flash-02 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use talxcc/Tals-coder-flash-02 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf talxcc/Tals-coder-flash-02 # Run inference directly in the terminal: llama cli -hf talxcc/Tals-coder-flash-02
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf talxcc/Tals-coder-flash-02 # Run inference directly in the terminal: llama cli -hf talxcc/Tals-coder-flash-02
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf talxcc/Tals-coder-flash-02 # Run inference directly in the terminal: ./llama-cli -hf talxcc/Tals-coder-flash-02
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf talxcc/Tals-coder-flash-02 # Run inference directly in the terminal: ./build/bin/llama-cli -hf talxcc/Tals-coder-flash-02
Use Docker
docker model run hf.co/talxcc/Tals-coder-flash-02
- LM Studio
- Jan
- vLLM
How to use talxcc/Tals-coder-flash-02 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "talxcc/Tals-coder-flash-02" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "talxcc/Tals-coder-flash-02", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/talxcc/Tals-coder-flash-02
- Ollama
How to use talxcc/Tals-coder-flash-02 with Ollama:
ollama run hf.co/talxcc/Tals-coder-flash-02
- Unsloth Desktop
- Pi
How to use talxcc/Tals-coder-flash-02 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf talxcc/Tals-coder-flash-02
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "talxcc/Tals-coder-flash-02" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use talxcc/Tals-coder-flash-02 with Docker Model Runner:
docker model run hf.co/talxcc/Tals-coder-flash-02
- Lemonade
How to use talxcc/Tals-coder-flash-02 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull talxcc/Tals-coder-flash-02
Run and chat with the model
lemonade run user.Tals-coder-flash-02-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use talxcc/Tals-coder-flash-02 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf talxcc/Tals-coder-flash-02
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default talxcc/Tals-coder-flash-02
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use talxcc/Tals-coder-flash-02 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf talxcc/Tals-coder-flash-02
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "talxcc/Tals-coder-flash-02" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- ⚡ Tals-coder-flash-02 (27B Ternary + Grafted Native MTP + Multimodal Vision)
- 🏎️ Pure Speed: Breaking the Physical Bandwidth Limit
- 🧬 Custom GGUF Quantization & Grafting
- 📊 Benchmark & Performance
- 📦 Model Files in Repository
- 🛠️ Optimal llama.cpp Runtime Setup
- ⚡ Recommended Run Configurations
- 💻 OpenCode / Cline / Continue Configuration
- 🙏 Credits & Acknowledgements
- ⚖️ Disclaimer & Responsible Use
- 🏎️ Pure Speed: Breaking the Physical Bandwidth Limit
⚡ Tals-coder-flash-02 (27B Ternary + Grafted Native MTP + Multimodal Vision)
The Fastest 27B Coding Model on Consumer Hardware
"Why wait 5 minutes for a reasoning model to finish an existential crisis in its chain-of-thought? Tals-coder-flash-02 goes at lightning speed with 60 tok/s on a budget 16GB GPU."
![]() |
![]() |
Hardware Testing Notice: All speed benchmarks and token generation speeds reported below were tested on an NVIDIA GeForce RTX 4060 Ti 16GB at stock factory settings without any overclocking temp 0.2 top p 0.85.
🏎️ Pure Speed: Breaking the Physical Bandwidth Limit
On paper, an NVIDIA RTX 4060 Ti 16GB has a narrow 128-bit memory bus with 288 GB/s bandwidth. Streaming an 8 GB 27B model on that bus physically caps standard autoregressive generation to ~25 tokens/sec.
Tals-coder-flash-02 shatters this barrier:
- ⚡ 55–60 tokens/second sustained on RTX 4060 Ti 16GB (verified in production!).
- 🚀 Up to 76.7 tokens/second peak burst during boilerplate and repetitive code blocks.
- 💨 120–150+ tokens/second projected on high-bandwidth hardware (RTX 3090, 4090, A100).
- 🎯 Zero Overthinking: No rambling monologues, no 2,000-token loops questioning itself. It cuts straight to clean, functional code.
🧬 Custom GGUF Quantization & Grafting
Tals-coder-flash-02 is a custom GGUF quantization of prism-ml/Ternary-Bonsai-2-27B-gguf:
- Surgically Grafted Native MTP: We permanently grafted the 65th Multi-Token Prediction (MTP) draft head and unrotated token embeddings directly into the main weights (
Tals-coder-flash-02.gguf). You do not need to juggle separate target and drafter model files—it's an all-in-one standalone file with native speculative decoding! - Multimodal Vision Projector: Includes the high-fidelity
Tals-coder-flash-02-mmproj.ggufCLIP Q8_0 vision tower (0.63 GB) for analyzing UI designs, charts, diagrams, and debugging screenshots. - 100% VRAM Execution: Runs 27B parameters, 128,000 context, and multimodal vision completely on GPU (~13.8–14.1 GB / 16.0 GB total) with zero CPU/PCIe spillover.
📊 Benchmark & Performance
Tested on NVIDIA GeForce RTX 4060 Ti 16GB (PCIe 4.0 x8, 288 GB/s bandwidth, stock factory settings without overclocking):
| Configuration | Speculative Acceptance | Generation Speed | Context Window |
|---|---|---|---|
| Base Model Only (No MTP) | N/A (0%) | 24.8 – 25.8 tok/s | Up to 262k |
Tals-coder-flash-02 (MTP n_max=1) |
77.8% | 40.7 – 42.5 tok/s | Up to 262k |
Tals-coder-flash-02 (MTP n_max=2) (Optimal) |
68.9% – 74.2% | 53.0 – 65.5 tok/s | Up to 262k |
| Boilerplate / Code Bursts | 85.2% | Up to 76.7 tok/s | Up to 262k |
| Estimated RTX 4090 / 3090 (1,008 GB/s) | 75% – 85% | 120 – 155 tok/s | Full 262k |
📦 Model Files in Repository
| Filename | Size | Description |
|---|---|---|
Tals-coder-flash-02.gguf |
8.25 GB | Unified 27B Custom GGUF Quant with Grafted Native MTP Head (All 65 layers). |
Tals-coder-flash-02-mmproj.gguf |
0.63 GB | High-fidelity CLIP Q8_0 multimodal vision projector. |
bonsai2-pascal.patch |
259 KB | Source patch for custom Hadamard kernels & MTP timing. |
flash.gif |
5.05 MB | High-speed Flash badge visual asset. |
🛠️ Optimal llama.cpp Runtime Setup
Because this model uses ternary quantization in a Hadamard-rotated basis, use the optimized llama.cpp build with Hadamard kernel support (Prism ML / Pascal branch, or apply bonsai2-pascal.patch).
Critical Environment Variables (Mandatory)
Before starting the server, set these environment variables to enable Pascal CUDA kernels and prevent CPU busy-polling:
Windows (PowerShell / CMD):
set GGML_CUDA_ROWLANE=1
set GGML_CUDA_RL_N4_LONG=1
set CUDA_DEVICE_SCHEDULE=BLOCKING_SYNC
Linux (Bash):
export GGML_CUDA_ROWLANE=1
export GGML_CUDA_RL_N4_LONG=1
export CUDA_DEVICE_SCHEDULE=BLOCKING_SYNC
⚡ Recommended Run Configurations
1. Default Production Mode (128k Context + Grafted MTP + Vision)
Run the single unified model with vision on a 16GB GPU:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
--mmproj "Tals-coder-flash-02-mmproj.gguf" \
--mmproj-offload \
--image-min-tokens 1024 \
-ngl 99 \
-c 128000 \
-ctk q8_0 -ctv q4_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--load-mode mmap --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --jinja \
--host 127.0.0.1 --port 8080
2. High-Precision Coding Mode (64k Context + FP16 KV Cache)
Maximum precision for complex mathematical derivations:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
--mmproj "Tals-coder-flash-02-mmproj.gguf" \
-ngl 99 \
-c 65536 \
-ctk f16 -ctv f16 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 8080
3. Extreme Context Mode (262k Native Context on 16GB VRAM)
Process entire codebases and large book-length documents with zero CPU offload:
./llama-server \
-m "Tals-coder-flash-02.gguf" \
-ngl 99 \
-c 262144 \
-ctk q4_0 -ctv q4_0 \
-fa on -np 1 -t 4 -tb 4 --poll 0 \
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.2 \
--backend-sampling --host 127.0.0.1 --port 8080
💻 OpenCode / Cline / Continue Configuration
Set your tool's API endpoint to:
- Base URL:
http://127.0.0.1:8080/v1 - Model Name:
tals-coder-flash-02 - API Key:
not-needed(any string) - Temperature:
0.0–0.2(Deterministic coding & reasoning)
🙏 Credits & Acknowledgements
- PrismML: The original creators of the Ternary-Bonsai-2-27B-gguf architecture and pioneering 1.58-bit ternary Hadamard quantization kernels.
- Ukisai: For Swift-Bonsai-2, adapting and quantizing the ternary base weights.
- killy369 / Kilian: For training and providing the native MTP speculative draft weights and 64k pruned vocabulary.
- Qwen Team: For the underlying Qwen 3.5 architecture.
⚖️ Disclaimer & Responsible Use
This model is provided "as is" for research, development, and educational purposes. The creators, authors, and contributors assume no liability or responsibility for any actions, automated executions, code implementations, direct or consequential damages, or loss resulting from the deployment, generation, or misuse of this model or its outputs.
Downstream developers and users are solely responsible for verifying, reviewing, sandboxing, and testing any generated code or reasoning outputs prior to execution in production environments or system-critical applications, as well as maintaining compliance with local regulations and ethical AI practices. Please use responsibly.
Distributed under the Apache 2.0 license.
- Downloads last month
- 784
We're not able to determine the quantization variants.

