L16-V65536 chat tokenizer
A 65,536-token multilingual SentencePiece tokenizer for 16 languages. It adds the Llama-3 chat special tokens in reserved slots.
At a glance
- Algorithm: SentencePiece BPE (65,536 vocab), with byte fallback, split digits and identity normalization. Character coverage is 0.99995.
- Languages (16): Russian, Spanish, Chinese (Simplified), Indonesian, Korean, Japanese, Turkish, Persian, German, Hungarian, Arabic, Vietnamese, Finnish, Greek, Thai and Hindi.
- Not English. English is not among the training languages, so English text takes about 2× as many tokens as with the Llama-3 tokenizer.
- Loads as:
LlamaTokenizer, from the fasttokenizer.jsonor the SentencePiecetokenizer.model. - No automatic BOS/EOS. It adds neither
<s>nor</s>by itself. In the pretraining mixture it was built for, each document is its tokens followed by</s>.
Special tokens
| id | token | Llama-3 equivalent |
|---|---|---|
| 0 | <unk> |
|
| 1 | <s> |
<|begin_of_text|> |
| 2 | </s> |
<|end_of_text|> |
| 3 | <pad> |
<|finetune_right_pad_id|> |
| 4 | <|start_header_id|> |
same |
| 5 | <|end_header_id|> |
same |
| 6 | <|eom_id|> |
same |
| 7 | <|eot_id|> |
same |
| 8 | <|python_tag|> |
same |
| 9–103 | <|reserved_005|> … <|reserved_099|> |
free slots |
Changes from the original L16_V65536 tokenizer
- Only ids 4–8 were renamed (from
<|reserved_000|>…<|reserved_004|>). - The rename is applied consistently in
tokenizer.json,tokenizer_config.jsonand thetokenizer.modelproto. Piece types and scores are unchanged. - The vocab size and every other token id are identical.
- Ordinary text tokenizes identically, checked on 1,200 documents across all 16 languages.
CHANGES.json records the renames and the sha256 of the original files. spm_training_config.json holds the original training settings.
Usage
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("amphora/L16-V65536-chat-tokenizer")
ids = tok("<|start_header_id|>user<|end_header_id|>\n\nПривет!<|eot_id|>")["input_ids"]
No chat template is included.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support