Buckets:
| configs: | |
| - config_name: yor | |
| data_files: | |
| - split: train | |
| path: yor_train.csv | |
| # - split: test | |
| # path: yor_test.csv | |
| - split: dev | |
| path: yor_dev.csv | |
| # - config_name: ara | |
| # data_files: | |
| # - split: test | |
| # path: ara_test.csv | |
| <center> <h2> Are LLMs Good Text Diacritizers? An Arabic and Yorùbá Case Study </h2> </center> | |
| <center> <b> Hawau Olamide Toyin, Samar Magdy, Hanan Aldarmaki </b> </center> | |
| We investigate the effectiveness of large language models (LLMs) for text diacritization in two typologically distinct languages: Arabic and Yoruba. To enable a rigorous evaluation, we introduce a novel multilingual dataset <strong>MultiDiac</strong> | |
| , with diverse samples that capture a range of diacritic ambiguities. We evaluate 14 LLMs varying in size, accessibility, and language coverage, and benchmarked them against 6 specialized diacritization models. Additionally, we fine-tune four small open-source models using LoRA for Yoruba. Our results show that many off-the-shelf LLMs outperform specialized diacritiztion models for both Arabic and Yoruba, but smaller models suffer from hallucinations. Fine-tuning on a small dataset can help improve diacritization performance and reduce hallucination rates. | |
| #### Cite this work: | |
| ``` | |
| @misc{toyin2025llmsgoodtextdiacritizers, | |
| title={Are LLMs Good Text Diacritizers? An Arabic and Yor\`ub\'a Case Study}, | |
| author={Hawau Olamide Toyin and Samar M. Magdy and Hanan Aldarmaki}, | |
| year={2025}, | |
| eprint={2506.11602}, | |
| archivePrefix={arXiv}, | |
| primaryClass={cs.CL}, | |
| url={https://arxiv.org/abs/2506.11602}, | |
| } | |
| ``` |
Xet Storage Details
- Size:
- 1.64 kB
- Xet hash:
- 3f7bb566706ed8e849bb7df6bb6f17f0cfcb8ad58b459d75bfebcb8d8be508f7
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.