YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Baseten Tokenizer for Kimi K3

High-performance, Rust-backed BPE tokenization for inference. This repo contains the tokenizer.json. The Python package is named basetenkenizer.

Fast tokenization for Kimi K3

Kimi K3 tokenization latency comparison

Install

pip install basetenkenizer

Loading by Hugging Face model ID downloads tokenizer.json on first use and then uses the local Hugging Face cache. Private or gated models use the usual HF_TOKEN environment variable.

Kimi K3

Kimi K3 uses XTML and a Python chat renderer rather than a Jinja chat template. Raw text can still be encoded normally:

from basetenkenizer import Tokenizer

tokenizer = Tokenizer.from_model("baseten/kimi-k3-tokenizer")
ids = tokenizer.encode("Hello from Kimi K3").ids

For rendered chat, pass the official Kimi K3 renderer's EncodeSegment objects to encode_segments as (segment.text, segment.allow_special). A minimal user message followed by an assistant-generation prefix looks like this:

user_message = "Explain speculative decoding."

segments = [
    ("<|open|>", True),
    ("message", False),
    (" role", False),
    ('="', False),
    ("user", False),
    ('"', False),
    ("<|sep|>", True),
    (user_message, False),
    ("<|close|>", True),
    ("message", False),
    ("<|sep|>", True),
    ("<|end_of_msg|>", True),
    ("<|open|>", True),
    ("message", False),
    (" role", False),
    ('="', False),
    ("assistant", False),
    ('"', False),
    ("<|sep|>", True),
    ("<|open|>", True),
    ("response", False),
    ("<|sep|>", True),
]

encoding = tokenizer.encode_segments(segments)
input_ids = encoding.ids

Use Kimi K3's official build_chat_segments renderer for production conversations involving tools, images, reasoning content, response schemas, or multiple message types. Do not concatenate its segments before encoding: doing so loses the special-token safety boundary.

Why encode_segments?

Each segment is a (text, allow_special) pair:

  • allow_special=True recognizes tokenizer control tokens emitted by a trusted chat renderer.
  • allow_special=False treats control-token-looking strings in user, tool, or attribute content as ordinary text.

tiktoken_safe=True is the default. It reproduces the chunk boundaries used by legacy tiktoken tokenizers, including on very long inputs, so token IDs stay compatible. Set it to False only when exact tiktoken parity is not required and whole-segment BPE encoding is intentional.

Post-processing, truncation, and padding are applied after all segment IDs have been joined.

Encode text with basetenkenizer

from basetenkenizer import Tokenizer

tokenizer = Tokenizer.from_model("deepseek-ai/DeepSeek-V3.2")
encoding = tokenizer.encode(
    "A very long prompt that is now much faster.",
    add_special_tokens=False,
)

print(encoding.ids)
print(tokenizer.decode(encoding.ids))

Tokenizer.from_file("tokenizer.json") loads a local tokenizer. Encoding objects expose ids, attention_mask, type_ids, and special_tokens_mask; selected fields can be moved into NumPy arrays with encoding.into_numpy(...).

Kimi K2.7 Code with basetenkenizer

Load the published tokenizer directly from its model repository:

from basetenkenizer import Tokenizer

tokenizer = Tokenizer.from_model("moonshotai/Kimi-K2.7-Code")
ids = tokenizer.encode("def hello():\n    return 'world'").ids

For a chat prompt, preserve the boundary between template control tokens and untrusted message content with encode_segments:

user_message = "Write a Python HTTP server."

segments = [
    ("<|im_user|>user<|im_middle|>", True),
    (user_message, False),
    ("<|im_end|>", True),
    ("<|im_assistant|>assistant<|im_middle|><think>", True),
]

encoding = tokenizer.encode_segments(segments)
input_ids = encoding.ids

This is the minimal Kimi K2.7 Code user/assistant shape. Applications using system messages, tools, images, or existing assistant messages should render the complete official model template and retain the same control-text versus message-text boundaries.

Use with Transformers

Call patch_transformers before loading a tokenizer:

import basetenkenizer

basetenkenizer.patch_transformers()

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("openai/gpt-oss-120b")
tokens = tokenizer("Hello, world!")

Transformers v4 and v5 are supported. Pass patch_transformers(apply_chat_template=True) to also use the native renderer for supported render-only apply_chat_template(..., tokenize=False) calls; unsupported templates automatically fall back to Transformers.

Baseten Tokenizer is focused on inference and does not implement every Hugging Face Tokenizers training or alignment feature.

License

The basetenkenizer package is licensed under the MIT License.

View basetenkenizer on PyPI

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support