Chiboard 1.1 tokenizer (prompt contract v3)
This is the tokenizer for prompt contract qwen35-chiboard-field-tokens-v3.
Pre-tokenization and decoding are the Qwen3.5 base tokenizer's.
- The normalizer, pre-tokenizer (regex
[\p{L}\p{M}]+), post-processor, decoder and BPE vocabulary/merges are copied verbatim fromQwen/Qwen3.5-0.8B-Base@dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68. Thattokenizer.jsonhas sha256fe000e3ed39ed12b8d2481d527d44f93c65d37e87645d2dcc80d1bf9d50d2927. - The merges keep the same string storage format.
- The normalizer, pre-tokenizer (regex
Added tokens are Chiboard 1 (T1)'s 37 added tokens (ids 248044–248080, including the four Chiboard field tokens) plus two new screen-slot tokens:
Token ID `< chiboard_screen_ref `< chiboard_screen The ids shared with the base tokenizer are identical. Tokenizer length is 248083; the model keeps 248320 embedding rows. No existing token id changed.
Why not T1's pre-tokenizer
T1's tokenizer.json matches what transformers 5's Qwen2Tokenizer class rebuilds: the Qwen2 regex (\p{L}+, which leaves combining marks out of words) and its default ByteLevel decoder. T1 was therefore trained with Qwen2 pre-tokenization. Its GGUF, however, runs llama.cpp's qwen35 pre-tokenizer.
Chiboard 1.1 trains and runs with the same Qwen3.5 pre-tokenization. On a 200K-document train sample, the two pre-tokenizers give different token ids for 0.0075% of documents. All of them contain combining marks (emoji U+FE0F, Indic/Thai signs, decomposed diacritics).
The first v3 build (@177cdc85, tokenizer.json sha256 2e62a571…) inherited T1's pre-tokenizer and is superseded.
Loading
tokenizer_config.json sets tokenizer_class: TokenizersBackend. With it, transformers 5.x loads tokenizer.json verbatim, so AutoTokenizer.from_pretrained produces exactly the same encodings as tokenizers.Tokenizer.from_file("tokenizer.json").
Do not load this tokenizer as Qwen2Tokenizer or Qwen3_5Tokenizer: those classes rebuild the pre-tokenizer and decoder from class constants.
Guard (trainer, evaluator, pipeline):
import json
from tokenizers import Tokenizer
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained(path)
want = json.loads(Tokenizer.from_file(f"{path}/tokenizer.json").to_str())
got = json.loads(tok.backend_tokenizer.to_str())
assert all(got[k] == want[k] for k in ("normalizer", "pre_tokenizer", "post_processor", "decoder"))
With this config, llama.cpp convert_hf_to_gguf.py computes chkhsh d30d75d9059f1aa2c19359de71047b3ae408c70875e8a3ccf8c5fba56c9d8af4, which maps to qwen35.
Other files:
tokenizer_config.jsonis T1's, with three changes:tokenizer_classisTokenizersBackend, the two new tokens are appended toextra_special_tokens, and two save-time keys are removed.chat_template.jinjais byte-identical to T1's. Contract v3 uses no chat template.
Serialization
<|chiboard_screen_ref|>{ref}<|chiboard_screen|>{screen}<|chiboard_context|>{ctx}<|chiboard_pinyin|>{raw}<|chiboard_display|>{display}<|chiboard_output|>
- Completion:
{target}<|endoftext|>. Loss applies to the completion only. - There is no BOS, no chat template and no EOS on the prompt side.
- Each payload is
[app_id][<|vision_start|><|image_pad|>×N<|vision_end|>], and both parts are optional. app_idcomes from the Chiboard canonical vocabulary. The valueunknownis omitted from the text.- The prompt carries
Npre-expanded<|image_pad|>tokens per image. Tokenize it directly with this tokenizer. - Do not pass a pre-expanded prompt through
Qwen3VLProcessor.__call__, which would expand each pad again.
Model tree for johnbean393/chiboard-1.1-tokenizer
Base model
Qwen/Qwen3.5-0.8B-Base