Tokenizer Validation Report
Model: geez_en_shared.model, vocab size: 120000
Per-language metrics (sampled)
| Language | Lines sampled | Fertility (tok/word) | Compression (chars/tok) | UNK rate | Roundtrip failure rate |
|---|---|---|---|---|---|
| amharic | 2000 | 1.389 | 3.594 | 0.0 | 0.009 |
| english | 2000 | 1.237 | 5.0 | 0.0 | 0.022 |
| geez | 1813 | 1.2 | 2.833 | 0.0 | 0.0 |
| tigre | 2000 | 1.285 | 3.181 | 0.0 | 0.0 |
| tigrinya | 2000 | 1.376 | 3.446 | 0.0 | 0.008 |
Note on roundtrip failure rate: these are not reconstruction bugs. Spot-checking
actual mismatches shows they're SentencePiece's default normalizer doing its job β
e.g. a real ellipsis character β¦ gets normalized to three dots ..., and a
leading BOM (ο»Ώ) gets stripped. decode(encode(x)) reflects the normalized
form, not raw input bytes; this is standard SentencePiece behavior, not data loss.
Morphologically hard words (from the MoVoC paper)
amharic
α αα°α α©α-> ['βα α', 'α°α α©', 'α'] (3 pieces, roundtrip OK)α¨ααα°α₯α¨α-> ['βα¨αα', 'α°α₯', 'α¨α'] (3 pieces, roundtrip OK)ααα ααα΅-> ['βααα α', 'αα΅'] (2 pieces, roundtrip OK)
tigrinya
α£αα΅α¨ααα-> ['βα£αα΅', 'α¨αα', 'α'] (3 pieces, roundtrip OK)ααα³α«α΅-> ['βααα³', 'α«α΅'] (2 pieces, roundtrip OK)
Bilin (byn) byte-fallback check
- probe:
α₯αα ααα-> 2 pieces, roundtrip OK - No Bilin corpus exists to train on (see SOURCES.md) β this only confirms byte_fallback keeps unseen Ge'ez-script text lossless, not that Bilin gets meaningful subword coverage.