DeepSeek Sharp Chat Templates
A drop-in Jinja chat template that applies a terseness system prompt to DeepSeek instruct models. Cuts filler without dropping correctness, so the model communicates more per token and stays on-task on long coding and knowledge-work turns.
License: apache-2.0
What it does
This is a DeepSeek instruct template with a terseness block force-appended after the model's own system prompt. Nothing you pass in is replaced — the block is appended, so the model keeps its own instructions.
The terseness core:
Never: open with preamble or pleasantries; restate the question; add filler
transitions; hedge with niceties; or repeat a point you've already made.
Always: keep essential steps, caveats, uncertainties, and specifics — never drop
correctness or a needed warning for brevity. Keep the final answer lean. Use the
least structure that conveys it (plain prose when short; lists or code only when
they earn their place). If genuinely uncertain, say so and explain why.
If a user request is genuinely ambiguous, ask a sharp question, don't guess.
Where the deliverable is mostly code or a structured artifact there is less padding to remove, so the effect is largest on explanation- and reasoning-heavy turns. The template deliberately protects substance — it never drops correctness for brevity.
Results
Measured on DeepSeek-Coder-V2-Lite across 6 blind code tasks, terseness vs. the stock template:
| task | stock | terse | delta |
|---|---|---|---|
| Bug hunt | 648 tok | 330 tok | -49% |
| Path traversal | 1388 tok | 603 tok | -56% |
| Optimize | 1052 tok | 965 tok | -8% |
| LRU cache | 1500 tok (truncated) | completed | avoids cut-off |
Token use drops 40–56% on ramble-heavy tasks with equal code quality; the model stops hitting output limits mid-implementation. On legitimately complex tasks it keeps the full answer rather than truncating. More benchmarks and model coverage forthcoming.
Use
MLX / oMLX / transformers: drop chat_template.jinja into the model directory
(overwrite the model's own template).
huggingface-cli download apuebla/DeepSeek-Sharp-Chat-Templates --local-dir DeepSeek-Sharp-Chat-Templates
llama.cpp / llama-server / koboldcpp:
llama-server -m your_deepseek.gguf --jinja --chat-template-file chat_template.jinja
vLLM / SGLang: set chat_template to the contents of chat_template.jinja.
Turning terseness off
Pass {"terse": false} via chat_template_kwargs (or edit the terse default near the
top of the template) to serve the model with only its own system prompt.
Supported engines
MLX · oMLX · llama.cpp / llama-server / koboldcpp · vLLM · SGLang · LM Studio — any engine that loads Hugging Face Jinja chat templates.
Sampling parameters
Start from DeepSeek's official defaults; the template does not change sampling. For
coding, a low temperature (0.2–0.4) with top_p ~0.9 suits the terse output. Keep
top_k ~20 to truncate long-tail token noise.
Credits
Base template patterns (tool format, KV-cache safety, error escalation): froggeric/Qwen-Fixed-Chat-Templates. Terseness approach inspired by a Qwen terse-template project. More coverage and model support forthcoming.
License
apache-2.0