File size: 1,993 Bytes
b9213f4 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 | ---
library_name: tokenizers
tags:
- tokenizer
- multilingual
- byte-level-bpe
- indigenous-languages
- tachiwin
---
# Tachiwin multilingual tokenizer
A multilingual ByteLevel-BPE tokenizer trained with Hugging Face Tokenizers.
## Corpus weighting
| Component | Target |
|---|---:|
| Modern + old exotic-language data | 70% |
| English | 10% |
| Spanish | 10% |
| Code | 10% |
The complete available exotic-language corpus is used as the 70% anchor.
Its existing modern/old composition is preserved.
## Tokenizer
- Model: BPE, trained via `Tokenizer.train_from_iterator` over a streaming
line generator (constant memory regardless of corpus size)
- Vocabulary target: 256,000
- Initial alphabet: complete ByteLevel alphabet
- ByteLevel GPT-2 regex: disabled
- Unicode normalizer: none
- Special tokens: 282
- Human-language tags: 248
## Corpus size
Total materialized corpus: 457,300,912 bytes (0.426 GiB)
## Evaluation
The tokenizer was evaluated against the Tachiwin language catalogue.
Languages without text samples are skipped. For each language, all
available samples are concatenated ONLY within that language for aggregate
fertility statistics. Metrics: characters/token, tokens/character, UTF-8
bytes/token, tokens/UTF-8 byte, exact round-trip preservation.
## Important training note
The Hugging Face BPE trainer does not expose an internal resumable
merge-state checkpoint. The recipe therefore treats the completed
`tokenizer.json` as the training checkpoint:
- corpus preparation is resumable (both the exotic streaming pass and the
capped external-corpus downloads reuse existing shards);
- recipe/statistics/checksums are stored in `recipe/`;
- if `tokenizer.json` already exists, subsequent runs skip BPE training;
- an interrupted BPE computation itself must be restarted.
## Repository evaluation artifacts
- evaluation/catalogue.json
- evaluation/language_fertility.csv
- evaluation/language_fertility.json
- evaluation/evaluation_summary.json
|