alexmt-gemma4-e4b-sft
Context-aware English ↔ dialectal Arabic translation for conversational turns, trained on Alexandria.
Supervised fine-tuning (full fine-tune, bf16, 2 epochs, lr 2e-5, effective batch 64) of Gemma-4-E4B-it on the Alexandria train split.
Data. Alexandria train split, 9 varieties (EG, JO, LB, MR, OM, PS, SA, SY, YE), both directions. The train split is divided into 5 conversation-level folds; SFT uses folds 1-4, and fold 0's source turns feed the preference pairs. Each example includes up to 3 prior turns as context, the domain, the participants, and the speaker/addressee gender direction.
Results
Alexandria public test split, 9 varieties, macro-average over varieties; greedy decoding, max 128 new tokens, our own prompt/decoding setup (not the Alexandria paper's).
| model | en→dialect spBLEU | chrF++ | len/ref | dialect→en spBLEU | chrF++ | len/ref |
|---|---|---|---|---|---|---|
| Gemma-4-E4B-it, zero-shot (no training) | 24.48 | 41.19 | 1.04 | 40.71 | 59.00 | 1.02 |
| alexmt-gemma4-e4b-sft (this model) | 29.06 | 44.10 | 0.99 | 50.61 | 65.81 | 0.98 |
Please read these before comparing: single training seed; one data fold; the DPO learning rate was compared on the test split itself (a held-out-dev selection is in progress), so small differences between the DPO rows are not established. Scores come from this setup and are not directly comparable to the Alexandria paper's tables (zero-shot models, different decoding) or to the AlexandriaX shared task (a different, private test set). Only overlap metrics (spBLEU, chrF++) have been computed; dialect fidelity has not been evaluated.
Files. model-kv-shared-unused.safetensors holds the k/v projections and k-norms of layers 24-41, which Gemma-4-E4B shares from earlier layers and never uses in the forward pass. transformers omits them on save; they are copied unchanged from the base model so that vLLM (which loads strictly) can load this checkpoint. processor_config.json is also copied from the base model for the same reason.
Usage
The model expects Alexandria's JSON prompt as a single user turn and answers with
{"translation": "..."}. Apply the tokenizer's chat template with add_generation_prompt=True, enable_thinking=False.
Translate the given turn of a conversation from English to Egyptian Arabic (Cairene) Dialect, considering the previous context if provided.
Input: {"country": "EG", "domain": "Agriculture and farming", "participants": ["Waterresourcemanager", "Farmer"], "context": [], "current_turn": {"speaker": "Waterresourcemanager", "gender_direction": "male -> male", "text": "Good morning. Is the water reaching your land properly according to the schedule?"}}
Guidelines:
- Return the result strictly in valid JSON.
- Translate to Egyptian Arabic (Cairene) Dialect using Arabic script.
- Do not add any code, explanations, comments, or any other extra text.
- Keep the meaning and tone and respect the gender direction.
- Consider the country, the domain, the participants, and the speaker in your translation.
- Only translate the "text" field of the "current_turn".
- If a context is provided, do not translate it, and use it to inform your translation.
Output scheme: { "translation": "translation of the text from the current turn" }
Licence
CC BY-NC 4.0 (non-commercial), because the training data (Alexandria) is CC BY-NC 4.0. The base model, Gemma-4-E4B-it is Apache-2.0 (Gemma 4 terms).
Citation
If you use this model, please cite the Alexandria dataset:
@inproceedings{elmekki2026alexandria,
title = {Alexandria: A Multi-Domain Dialectal Arabic Machine Translation Dataset for Culturally Inclusive and Linguistically Diverse LLMs},
author = {El Mekki, Abdellah and others},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (ACL)},
year = {2026},
url = {https://aclanthology.org/2026.acl-long.1503/}
}
- Downloads last month
- 160