|
Download docs/modules/chat.md from PYTHAI/bankml: direct link, hf CLI and curl.
- Browser
- Download file 7.56 kB
-
https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/chat.md
- Command line
-
hf download hf://spaces/PYTHAI/bankml/docs/modules/chat.md
-
curl -L -o chat.md https://huggingface.co/spaces/PYTHAI/bankml/resolve/main/docs/modules/chat.md
7.56 kB
| # `bankML/chat.rs` — chat templates, byte-identical to llama.cpp b11192 | |
| ## Summary | |
| A conversation becomes a prompt through the model's chat template, a Jinja program stored in the GGUF | |
| (`tokenizer.chat_template`). bankml does not run Jinja. `chat.rs` writes out the rules of each template it | |
| reproduces, and identifies the model's template by the sha256 of its text. A template it does not know is refused. | |
| Three templates are reproduced (`TEMPLATES`): | |
| | sha256 prefix | `Template` | carried by | | |
| |---|---|---| | |
| | `30a75d10e60b57e2` | `Qwen3` | Bonsai 1.7B and 8B (Qwen3), thinking off | | |
| | `872be49dbb638044` | `SmolLm2` | SmolLM2-Instruct: ChatML with a default system message | | |
| | `9fe579a2c222698c` | `ChatMl` | plain ChatML, no default system message (`mindx-genN`) | | |
| On the Qwen3 template the generation prompt ends in an empty `<think>` block (thinking off). On the two ChatML | |
| templates every message renders as `<|im_start|>{role}\n{content}<|im_end|>\n` whatever its role, and | |
| `reasoning_content` is dropped, as llama-server renders them. | |
| Callers: the native engine (`native.rs`: `template_of` at open, `render` for every prompt, `generation_prompt` for | |
| the grammar's prefill); `bankml serve --native` (`POST /apply-template`, `/v1/chat/completions`, `/api/chat`); the | |
| Ollama layer (`ollama.rs`); `bankml chat-template` and `bankml generate` (`main.rs`); and `grammar.rs` / `schema.rs`, | |
| which choose the JSON grammar's root by template. | |
| ## Technical usage | |
| ```rust | |
| pub const TEMPLATE_SHA256: &str = "30a75d10e60b57e2"; | |
| pub enum Template { Qwen3, SmolLm2, ChatMl } | |
| pub const TEMPLATES: [(&str, Template, &str); 3]; | |
| impl Template { | |
| pub fn render(self, msgs: &[Message]) -> Result<String, String> | |
| pub fn generation_prompt(self) -> &'static str | |
| } | |
| pub struct Message { pub role: String, pub content: String, pub reasoning: Option<String> } | |
| impl Message { pub fn new(role: &str, content: &str) -> Self } | |
| pub fn template_of(gguf: &std::path::Path) -> Result<Template, String> | |
| pub fn check_template(gguf: &std::path::Path) -> Result<(), String> | |
| pub fn messages_from_json(v: &Json) -> Result<Vec<Message>, String> | |
| pub fn render(msgs: &[Message]) -> Result<String, String> // the Qwen3 template | |
| ``` | |
| - `template_of` guards the file, hashes `tokenizer.chat_template` and matches the first 16 hex digits against | |
| `TEMPLATES`. The error names the hash it found and the templates it knows. | |
| - `messages_from_json` reads OpenAI-style `[{"role", "content", "reasoning_content"?}, …]`. An empty | |
| `reasoning_content` is dropped before templating, as llama-server does. | |
| - `render` (Qwen3) reproduces the template's details: a first system message; the last real user query (scanning | |
| back, the first user message that is not a wrapped `<tool_response>`); assistant turns after it keep their | |
| `<think>` block, those before it lose it; runs of tool messages grouped into one user turn of | |
| `<tool_response>` blocks; Python's `split` and `strip` semantics where the template uses them. | |
| - `generation_prompt()` is what the template appends for the assistant's turn: | |
| `"<|im_start|>assistant\n<think>\n\n</think>\n\n"` on Qwen3, `"<|im_start|>assistant\n"` on the ChatML templates. | |
| ```sh | |
| echo '[{"role":"system","content":"You are Savante."},{"role":"user","content":"Hi"}]' \ | |
| | bankml chat-template .models/Bonsai-8B-Q1_0.gguf | |
| ``` | |
| prints: | |
| ```text | |
| <|im_start|>system | |
| You are Savante.<|im_end|> | |
| <|im_start|>user | |
| Hi<|im_end|> | |
| <|im_start|>assistant | |
| <think> | |
| </think> | |
| ``` | |
| ## How it is verified | |
| - Unit tests: `a_short_conversation`, `python_split_semantics`, `chatml_templates`. | |
| - `oracle_chat_template` (`#[ignore]`, in the gate): `testing/template_oracle.py` records llama-server's own | |
| `/apply-template` for 317 conversations (system prompts first, later or absent; assistant turns with and without | |
| `<think>` blocks around the last real query; `reasoning_content`; runs of tool results; user messages that look | |
| like tool responses; special markers and Unicode in content; 300 random conversations). Every prompt must be | |
| byte-identical: **317 of 317** (0.3.6 gate record). | |
| - `oracle_chat_template_chatml`: SmolLM2-Instruct's template and mindx-gen39's ChatML, each against llama-server | |
| rendering that model's own, **317 of 317** each (docs/oracles.md §5d). | |
| - Indirectly: the conversation oracle (`oracle_native_serve`, 9 of 9 turns, including the prompt cache's reuse) and | |
| every greedy, sampling and JSON oracle start from a rendered prompt. | |
| The oracle found one server behaviour the template alone would not predict: an empty `reasoning_content` is dropped. | |
| ## Advantages and efficiency | |
| - **No template engine.** No Jinja interpreter and no crate: each template is a short, readable Rust function, and | |
| the sha256 pin means a model whose template changed is refused instead of rendered wrongly. | |
| - **Exact prompts.** A byte-identical prompt is what lets the prompt cache reuse what llama-server would reuse, and | |
| lets every token-level oracle compare like with like. | |
| - **Cheap.** Rendering is string concatenation over the messages; the template is identified once, when the engine | |
| opens the model. | |
| - **Rust practice.** Zero dependencies, no `unsafe`; out-of-scope input is an `Err` with the reason, not a guess. | |
| - **Next** (docs/TODO.md, docs/OLLAMA.md): tool calls through the template (O6; JSON schemas, its first step, came in | |
| 0.3.5), and the Llama 3.x template with that architecture (0.6.0). | |
| ## Limitations | |
| - Only the three templates above. Any other is refused by its hash. | |
| - Refused rather than guessed: tool definitions and `tool_calls` (the template serialises them with its own | |
| `tojson`); non-text content; a conversation ending with an assistant message (llama-server treats that as a | |
| prefill, which is server logic beyond the template); an empty message list. | |
| - On Qwen3, a role other than system, user, assistant and tool renders nothing, as in the template. | |
| - Ollama's persona `SYSTEM` for `mindx-genN` lives in its Modelfile, not in the GGUF: here it is the caller's system | |
| message (`bankml create` handles the Modelfile). | |
| ## Design notes | |
| - History: the chat template was step two of P3 (the Bonsai / Qwen3 template only). 0.3.4 (O4) mapped templates | |
| per model, each pinned by the sha256 of its text, adding SmolLM2-Instruct's and the plain ChatML of `mindx-genN`. | |
| - The SmolLM2-Instruct default system message, inserted when the first message is not a system message, is | |
| `You are a helpful AI assistant named SmolLM, trained by Hugging Face`. | |
| - `mindx-genN` is SmolLM2-135M fine-tuned by mindXtrain; its GGUF carries plain ChatML. | |
| - Each template's oracle is llama-server's `/apply-template` on a GGUF that carries that template | |
| (`testing/template_oracle.py`); the oracle shows both ChatML templates render a tool result like any other role | |
| and drop `reasoning_content`. | |
| - Why an empty `reasoning_content` is dropped: llama-server removes it before templating, so the assistant | |
| content's own `<think>` block is split out; a present-but-empty string would have prevented that split. | |
| ## See also | |
| - [../oracles.md](../oracles.md) §1c, §5d — the chat-template oracle | |
| - [../usage.md](../usage.md) §13 — `bankml chat-template` | |
| - [../TECHNICAL.md](../TECHNICAL.md) §III.7 — from a template to a token | |
| - [../OLLAMA.md](../OLLAMA.md) — O4 templates, O6 tool calls | |
| - [../TODO.md](../TODO.md) | |
| - Sibling pages: [tokenizer.md](tokenizer.md), [grammar.md](grammar.md), [schema.md](schema.md), | |
| [native.md](native.md), [serve.md](serve.md), [ollama.md](ollama.md), [create.md](create.md) | |