L16-V65536 chat tokenizer

A 65,536-token multilingual SentencePiece tokenizer for 16 languages. It adds the Llama-3 chat special tokens in reserved slots.

At a glance

  • Algorithm: SentencePiece BPE (65,536 vocab), with byte fallback, split digits and identity normalization. Character coverage is 0.99995.
  • Languages (16): Russian, Spanish, Chinese (Simplified), Indonesian, Korean, Japanese, Turkish, Persian, German, Hungarian, Arabic, Vietnamese, Finnish, Greek, Thai and Hindi.
  • Not English. English is not among the training languages, so English text takes about 2× as many tokens as with the Llama-3 tokenizer.
  • Loads as: LlamaTokenizer, from the fast tokenizer.json or the SentencePiece tokenizer.model.
  • No automatic BOS/EOS. It adds neither <s> nor </s> by itself. In the pretraining mixture it was built for, each document is its tokens followed by </s>.

Special tokens

id token Llama-3 equivalent
0 <unk>
1 <s> <|begin_of_text|>
2 </s> <|end_of_text|>
3 <pad> <|finetune_right_pad_id|>
4 <|start_header_id|> same
5 <|end_header_id|> same
6 <|eom_id|> same
7 <|eot_id|> same
8 <|python_tag|> same
9–103 <|reserved_005|> … <|reserved_099|> free slots

Changes from the original L16_V65536 tokenizer

  • Only ids 4–8 were renamed (from <|reserved_000|>…<|reserved_004|>).
  • The rename is applied consistently in tokenizer.json, tokenizer_config.json and the tokenizer.model proto. Piece types and scores are unchanged.
  • The vocab size and every other token id are identical.
  • Ordinary text tokenizes identically, checked on 1,200 documents across all 16 languages.

CHANGES.json records the renames and the sha256 of the original files. spm_training_config.json holds the original training settings.

Usage

from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("amphora/L16-V65536-chat-tokenizer")
ids = tok("<|start_header_id|>user<|end_header_id|>\n\nПривет!<|eot_id|>")["input_ids"]

No chat template is included.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support