geez-en-shared-tokenizer / validation_report.md
Hailay's picture
Upload validation_report.md with huggingface_hub
baadd6c verified
|
Raw
History Blame Contribute Delete
1.69 kB

Tokenizer Validation Report

Model: geez_en_shared.model, vocab size: 120000

Per-language metrics (sampled)

Language Lines sampled Fertility (tok/word) Compression (chars/tok) UNK rate Roundtrip failure rate
amharic 2000 1.389 3.594 0.0 0.009
english 2000 1.237 5.0 0.0 0.022
geez 1813 1.2 2.833 0.0 0.0
tigre 2000 1.285 3.181 0.0 0.0
tigrinya 2000 1.376 3.446 0.0 0.008

Note on roundtrip failure rate: these are not reconstruction bugs. Spot-checking actual mismatches shows they're SentencePiece's default normalizer doing its job β€” e.g. a real ellipsis character … gets normalized to three dots ..., and a leading BOM (ο»Ώ) gets stripped. decode(encode(x)) reflects the normalized form, not raw input bytes; this is standard SentencePiece behavior, not data loss.

Morphologically hard words (from the MoVoC paper)

amharic

  • αŠ αˆαˆ°α‰ αˆ©αˆ -> ['β–αŠ αˆ', 'αˆ°α‰ αˆ©', 'ም'] (3 pieces, roundtrip OK)
  • α‹¨αˆ›α‹­αˆ°α‰₯αˆ¨α‹ -> ['β–α‹¨αˆ›α‹­', 'ሰα‰₯', 'αˆ¨α‹'] (3 pieces, roundtrip OK)
  • αˆ˜αˆαˆ…αˆ­αŠα‰΅ -> ['β–αˆ˜αˆαˆ…αˆ­', 'αŠα‰΅'] (2 pieces, roundtrip OK)

tigrinya

  • αŠ£α‹­α‰΅αŠ¨α‹αŠ•αŠ• -> ['β–αŠ£α‹­α‰΅', 'αŠ¨α‹αŠ•', 'αŠ•'] (3 pieces, roundtrip OK)
  • αˆαˆ•α‹³αˆ«α‰΅ -> ['β–αˆαˆ•α‹³', 'αˆ«α‰΅'] (2 pieces, roundtrip OK)

Bilin (byn) byte-fallback check

  • probe: α‰₯αˆŠαŠ• α‰‹αŠ•α‰‹ -> 2 pieces, roundtrip OK
  • No Bilin corpus exists to train on (see SOURCES.md) β€” this only confirms byte_fallback keeps unseen Ge'ez-script text lossless, not that Bilin gets meaningful subword coverage.