Chiboard 1.1 tokenizer (prompt contract v3)

This is the tokenizer for prompt contract qwen35-chiboard-field-tokens-v3.

  • Pre-tokenization and decoding are the Qwen3.5 base tokenizer's.

    • The normalizer, pre-tokenizer (regex [\p{L}\p{M}]+), post-processor, decoder and BPE vocabulary/merges are copied verbatim from Qwen/Qwen3.5-0.8B-Base@dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68. That tokenizer.json has sha256 fe000e3ed39ed12b8d2481d527d44f93c65d37e87645d2dcc80d1bf9d50d2927.
    • The merges keep the same string storage format.
  • Added tokens are Chiboard 1 (T1)'s 37 added tokens (ids 248044–248080, including the four Chiboard field tokens) plus two new screen-slot tokens:

    Token ID
    `< chiboard_screen_ref
    `< chiboard_screen

    The ids shared with the base tokenizer are identical. Tokenizer length is 248083; the model keeps 248320 embedding rows. No existing token id changed.

Why not T1's pre-tokenizer

T1's tokenizer.json matches what transformers 5's Qwen2Tokenizer class rebuilds: the Qwen2 regex (\p{L}+, which leaves combining marks out of words) and its default ByteLevel decoder. T1 was therefore trained with Qwen2 pre-tokenization. Its GGUF, however, runs llama.cpp's qwen35 pre-tokenizer.

Chiboard 1.1 trains and runs with the same Qwen3.5 pre-tokenization. On a 200K-document train sample, the two pre-tokenizers give different token ids for 0.0075% of documents. All of them contain combining marks (emoji U+FE0F, Indic/Thai signs, decomposed diacritics).

The first v3 build (@177cdc85, tokenizer.json sha256 2e62a571…) inherited T1's pre-tokenizer and is superseded.

Loading

tokenizer_config.json sets tokenizer_class: TokenizersBackend. With it, transformers 5.x loads tokenizer.json verbatim, so AutoTokenizer.from_pretrained produces exactly the same encodings as tokenizers.Tokenizer.from_file("tokenizer.json").

Do not load this tokenizer as Qwen2Tokenizer or Qwen3_5Tokenizer: those classes rebuild the pre-tokenizer and decoder from class constants.

Guard (trainer, evaluator, pipeline):

import json
from tokenizers import Tokenizer
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained(path)
want = json.loads(Tokenizer.from_file(f"{path}/tokenizer.json").to_str())
got = json.loads(tok.backend_tokenizer.to_str())
assert all(got[k] == want[k] for k in ("normalizer", "pre_tokenizer", "post_processor", "decoder"))

With this config, llama.cpp convert_hf_to_gguf.py computes chkhsh d30d75d9059f1aa2c19359de71047b3ae408c70875e8a3ccf8c5fba56c9d8af4, which maps to qwen35.

Other files:

  • tokenizer_config.json is T1's, with three changes: tokenizer_class is TokenizersBackend, the two new tokens are appended to extra_special_tokens, and two save-time keys are removed.
  • chat_template.jinja is byte-identical to T1's. Contract v3 uses no chat template.

Serialization

<|chiboard_screen_ref|>{ref}<|chiboard_screen|>{screen}<|chiboard_context|>{ctx}<|chiboard_pinyin|>{raw}<|chiboard_display|>{display}<|chiboard_output|>
  • Completion: {target}<|endoftext|>. Loss applies to the completion only.
  • There is no BOS, no chat template and no EOS on the prompt side.
  • Each payload is [app_id][<|vision_start|><|image_pad|>×N<|vision_end|>], and both parts are optional.
  • app_id comes from the Chiboard canonical vocabulary. The value unknown is omitted from the text.
  • The prompt carries N pre-expanded <|image_pad|> tokens per image. Tokenize it directly with this tokenizer.
  • Do not pass a pre-expanded prompt through Qwen3VLProcessor.__call__, which would expand each pad again.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for johnbean393/chiboard-1.1-tokenizer

Finetuned
(128)
this model