tokenizer_test / README.md
ljcamargo's picture
Upload README.md with huggingface_hub
a60a379 verified
|
Raw
History Blame Contribute Delete
1.32 kB
metadata
library_name: tokenizers
tags:
  - tokenizer
  - multilingual
  - byte-level-bpe
  - indigenous-languages
  - tachiwin

Tachiwin multilingual tokenizer

A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.

Corpus weighting

Component Target
Modern + old exotic-language data 70%
English 10%
Spanish 10%
Code 10%

The complete available exotic-language corpus is used as the 70% anchor. Its existing modern/old composition is preserved.

Tokenizer

  • Model: BPE
  • Vocabulary target: 256,000
  • Initial alphabet: complete ByteLevel alphabet
  • ByteLevel GPT-2 regex: disabled
  • Unicode normalizer: none
  • Special tokens: 94
  • Human-language tags: 60

Corpus size

Total materialized corpus:

50,027,186 bytes (0.047 GiB)

Important training note

The Hugging Face BPE trainer does not expose an internal resumable merge-state checkpoint. The recipe therefore treats the completed tokenizer.json as the training checkpoint:

  • corpus preparation is resumable;
  • recipe/statistics/checksums are stored in recipe/;
  • if tokenizer.json already exists, subsequent runs skip BPE training;
  • an interrupted BPE computation itself must be restarted.

This avoids changing the training procedure merely to obtain artificial checkpointing.