File size: 2,222 Bytes
0722ae4 0b999a6 0722ae4 0b999a6 0722ae4 0b999a6 0722ae4 0b999a6 e841b14 0b999a6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 | ---
title: Arabic Tokenizer Comparison
emoji: ๐ค
colorFrom: green
colorTo: indigo
sdk: gradio
sdk_version: 6.22.0
app_file: app.py
pinned: false
license: apache-2.0
short_description: "How many tokens does your Arabic actually cost?"
---
# Arabic Tokenizer Comparison โ ูู
ุฑู
ุฒุงู ูููููู ุงููุต ุงูุนุฑุจูุ
Every model card quotes parameter count and context length. None of them tell
you that the same Arabic paragraph costs **50% more context** on one tokenizer
than another โ a tax paid on every prompt, every document and every embedding
you ever run.
Paste Arabic โ MSA, dialect, or code-switched โ and see what nine tokenizers
actually spend on it.
## The point isn't vocabulary size
**Mistral-7B and Emhotob both have ~32K vocabularies.** On Modern Standard
Arabic, Mistral needs **4.5ร more tokens** for identical text. The difference
isn't how many entries a vocabulary has, it's what those entries are spent on.
Measured on a 1,024-word MSA paragraph:
| tokenizer | vocab | tokens/word | vs Emhotob |
|---|---|---|---|
| **Emhotob 32K** (Arabic-only) | 32,000 | **1.391** | 1.00ร |
| Qwen3.6 | 248,044 | 1.885 | 1.36ร |
| Gemma-4 | 262,144 | 2.109 | 1.52ร |
Emhotob beats vocabularies **8ร its size** on Arabic โ and loses badly on
English, which it was never built for. That trade is the whole design.
The token-split view shows why. On the same sentence, Mistral emits **one token
per Arabic letter**, Gemma-4 breaks words into fragments (`ุฃ` `ุนู` `ูุช`), and
Emhotob keeps them whole.
## What to try
- **Dialect**: Emhotob's lead narrows from 1.52ร to ~1.19ร. Dialectal
orthography is where an MSA-trained vocabulary is weakest.
- **Code-switched**: Gemma-4 *wins*. Half the text is English.
- **English**: Emhotob costs 2.2ร what Gemma-4 does. Right tool, wrong job.
## About Emhotob
The Arabic-only byte-level BPE behind
[`50M-2048-Emhotob`](https://huggingface.co/oddadmix/50M-2048-Emhotob) and the
[Nawah](https://huggingface.co/oddadmix/Nawah-50M-RAG-Support-2K) family. At
hidden size 512 its 32,000 entries cost 16.4M embedding parameters; a 262K
vocabulary would cost 134M โ more than twice the entire model.
ยฉ KAND CA 2026 โ PROJECT NAWAH
|