A large-scale synthetic Arabic OCR dataset comprising 843,622 book-style document images across 10 fonts, designed to advance VLM for Arabic Texts
Robotics and Internet-of-Things
riotu-lab
AI & ML interests
None yet
Organizations
None yet
SARD: Synthetic Arabic Recognition Dataset
A large-scale synthetic Arabic OCR dataset comprising 843,622 book-style document images across 10 fonts, designed to advance VLM for Arabic Texts
Aranizer | Arabic Tokenization with SentencePiece & PBE
Collection of Arabic Tokenizers with different sizes based on SentencePiece & PBE Encodings suitable for training LLMs
ArabianLLM Series | Native Arabic Large Language Models
This collection is related to native Arabic Large Language Models.. It represent different sizes of GPT trained Model for Test Generative