DeepSeek Sharp Chat Templates

A drop-in Jinja chat template that applies a terseness system prompt to DeepSeek instruct models. Cuts filler without dropping correctness, so the model communicates more per token and stays on-task on long coding and knowledge-work turns.

License: apache-2.0


What it does

This is a DeepSeek instruct template with a terseness block force-appended after the model's own system prompt. Nothing you pass in is replaced — the block is appended, so the model keeps its own instructions.

The terseness core:

Never: open with preamble or pleasantries; restate the question; add filler
transitions; hedge with niceties; or repeat a point you've already made.
Always: keep essential steps, caveats, uncertainties, and specifics — never drop
correctness or a needed warning for brevity. Keep the final answer lean. Use the
least structure that conveys it (plain prose when short; lists or code only when
they earn their place). If genuinely uncertain, say so and explain why.
If a user request is genuinely ambiguous, ask a sharp question, don't guess.

Where the deliverable is mostly code or a structured artifact there is less padding to remove, so the effect is largest on explanation- and reasoning-heavy turns. The template deliberately protects substance — it never drops correctness for brevity.

Results

Measured on DeepSeek-Coder-V2-Lite across 6 blind code tasks, terseness vs. the stock template:

task stock terse delta
Bug hunt 648 tok 330 tok -49%
Path traversal 1388 tok 603 tok -56%
Optimize 1052 tok 965 tok -8%
LRU cache 1500 tok (truncated) completed avoids cut-off

Token use drops 40–56% on ramble-heavy tasks with equal code quality; the model stops hitting output limits mid-implementation. On legitimately complex tasks it keeps the full answer rather than truncating. More benchmarks and model coverage forthcoming.

Use

MLX / oMLX / transformers: drop chat_template.jinja into the model directory (overwrite the model's own template).

huggingface-cli download apuebla/DeepSeek-Sharp-Chat-Templates --local-dir DeepSeek-Sharp-Chat-Templates

llama.cpp / llama-server / koboldcpp:

llama-server -m your_deepseek.gguf --jinja --chat-template-file chat_template.jinja

vLLM / SGLang: set chat_template to the contents of chat_template.jinja.

Turning terseness off

Pass {"terse": false} via chat_template_kwargs (or edit the terse default near the top of the template) to serve the model with only its own system prompt.

Supported engines

MLX · oMLX · llama.cpp / llama-server / koboldcpp · vLLM · SGLang · LM Studio — any engine that loads Hugging Face Jinja chat templates.

Sampling parameters

Start from DeepSeek's official defaults; the template does not change sampling. For coding, a low temperature (0.2–0.4) with top_p ~0.9 suits the terse output. Keep top_k ~20 to truncate long-tail token noise.

Credits

Base template patterns (tool format, KV-cache safety, error escalation): froggeric/Qwen-Fixed-Chat-Templates. Terseness approach inspired by a Qwen terse-template project. More coverage and model support forthcoming.

License

apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support