ByT5 multilingual P2G (small) โ€” 300M params

Phoneme-to-grapheme inversion of the harmonized retrain. Input: <lang>: phoneme tokens. Output: word spelling. Trained on a 4.12M-pair harmonized corpus (same as the G2P sibling).

Results (4k stratified test sample)

score
micro exact 0.483
macro exact 0.581

P2G is harder than G2P (phoneme โ†’ spelling is not one-to-one), but this model is useful for AAC scenarios where a user sounds out a word they can't spell.

Files

  • HF-format weights at root (~1.2 GB)
  • onnx/ โ€” validated encoder+decoder pair

Licence

CC BY-SA 4.0

Downloads last month
33
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for willwade/byt5-p2g-multilingual

Quantized
(6)
this model