File size: 9,856 Bytes
746f31f 9d3c51c 405c1ba 9d3c51c 746f31f 5c454fa 746f31f 7c1c429 5c454fa 746f31f 5c454fa 746f31f 5c454fa 746f31f 5c454fa 746f31f 3c9bdc8 746f31f 5c454fa 746f31f 3c9bdc8 746f31f | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 | ---
license: apache-2.0
language:
- ott
- tr
- ar
pipeline_tag: image-to-text
tags:
- ocr
- ottoman
- qwen
- vllm
- vision-language
metrics:
- accuracy
- cer
---
# Azra 1-Mini (0.8B) — Ottoman Turkish OCR Model
<p align="center">
<img src="banner.png" alt="Azra 1-Mini Banner" width="45%">
</p>
**Azra 1-Mini 0.8b** is a lightweight, high-performance Vision-Language OCR model specialized in Ottoman Turkish text transcription across both **Nesih** (printed/calligraphic) and **Rika/Riqa** (handwritten) scripts.
Despite having only **0.8 billion parameters**, Azra 1-Mini achieves state-of-the-art accuracy on Ottoman Turkish OCR tasks, outperforming significantly larger proprietary models.
> ⚠️ **Important Note on Input Resolution & Segmentation:**
> This model has been fine-tuned and optimized specifically for **line-level text images (satır bazlı görüntüler)**. It may not achieve optimal accuracy directly on full-page images without prior text line cropping/segmentation.
---
## 📊 Benchmark Results & Performance Comparison
The model was evaluated against leading proprietary Vision-Language models on standard Ottoman Turkish test sets using character accuracy (`100% - CER`).
### 1. Nesih Script Test Set (Printed / Calligraphic)
| Model | Success Rate (%) | Rank |
| :--- | :---: | :---: |
| **Gemini 3.1 Pro** | **82.82%** | 👑 1st |
| **Azra 1-Mini 0.8b** | **80.23%** | 🥈 2nd |
| **Qwen 3.8 Max** | 74.26% | 🥉 3rd |
### 2. Rika Script Test Set (Handwritten)
| Model | Success Rate (%) | Rank |
| :--- | :---: | :---: |
| **Azra 1-Mini 0.8b** | **67.08%** | 👑 **1st (Winner)** |
| **Gemini 3.1 Pro** | 58.46% | 🥈 2nd |
| **Qwen 3.8 Max** | 54.99% | 🥉 3rd |
> 🌟 **Key Highlight:** Azra 1-Mini 0.8b achieves **1st place on the handwritten Rika dataset (67.08%)**, significantly outperforming both Gemini 3.1 Pro and Qwen 3.8 Max while running efficiently at sub-billion parameter scale.
---
## 📷 Qualitative Results & Sample Transcriptions
Below are top qualitative predictions generated by **Azra 1-Mini 0.8b** from the evaluation test sets:
### 1. Nesih Script Samples (Printed / Calligraphic)
| Image | Ground Truth (GT) | Model Prediction (Azra 1-Mini) | CER |
| :---: | :--- | :--- | :---: |
| <img src="assets/samples/nesih_3263_759633_eSc_line_f5695a2d.png" height="35"> | `امّا اری وابدار ونازک اولور هر اعجک زمان غرسی` | `امّا اری وابدار ونازک اولور هر اعجک زمان غرسی` | **0.00%** |
| <img src="assets/samples/nesih_3263_759704_eSc_line_211ddfad.png" height="35"> | `هلاک ایدر ازایسه علاج ایله خلاص اولور` | `هلاک ایدر ازایسه علاج ایله خلاص اولور` | **0.00%** |
| <img src="assets/samples/nesih_3263_759800_eSc_line_83cb8596.png" height="35"> | `یافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردل` | `یافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردل` | **0.00%** |
| <img src="assets/samples/nesih_3263_759772_eSc_line_21130720.png" height="35"> | `اغزی محکم باغلنوب اول بوداق اکلوب یره کوملسه وقت` | `اغزی محکمه باغلنوب اول بوداق اکلوب یره کوملسه وقت` | **2.08%** |
| <img src="assets/samples/nesih_3263_759629_eSc_line_e9d59dc1.png" height="35"> | `دکمک زماندر دیمش یعنی آیک نقصانی زمانی که اوّل` | `دکک زماندر دیمش یعنی آیک نقصانی زمانی که اوّل` | **2.17%** |
### 2. Rika Script Samples (Handwritten)
| Image | Ground Truth (GT) | Model Prediction (Azra 1-Mini) | CER |
| :---: | :--- | :--- | :---: |
| <img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-001.png" height="35"> | `دیمک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر و بو هدایت ایکی` | `دیک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر۔ و بوهدایت ایکی` | **4.62%** |
| <img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-015.png" height="35"> | `ایتمک دون بنی تنویر ایدن کونشک یارین تنویر ایدهمیهجکنی ادعا ایتمک کبی قانون استقرایی انکاردر۔` | `ایتمک دوند بنی تنویر ایدن کونشک یارین تنویر ایدرمهجیکنی ادعا ایتمک کبی قانون استقرالی انکاردر۔` | **5.38%** |
| <img src="assets/samples/riqa_c25bfa03-dbf2-41e2-8164-d639b87fbede-015.png" height="35"> | `ایمانده نه قدر بیوک بر سعادت و نعمت؛ و نه قدر بیوک بر لذت و راحت بولوندیغنی اڭلامق` | `ایمانده نه قدر یوک بر سعادت ونعمت و نه قدر یوک بر لذت و راحت بولوندیغی اشلامم` | **8.54%** |
| <img src="assets/samples/riqa_e2da2c9f-5dc9-4ffa-8241-d13cffc1caef-020.png" height="35"> | `”الله تعالی ابراهیم علیه السلامه وحی ایدوب دیدی که: اسماعیل حقندهکی دعاکی قبول ایتدم و اونی` | `"الله تعالی ابراهیم علمه السلام دحی ایدوب دیدی کی: اسماعیل حقندهکی دعاک قبول ایتدم واولی` | **8.79%** |
| <img src="assets/samples/riqa_81f5dd07-3ee3-4b26-a02a-a589e289f11c-003.png" height="35"> | `بوراده مطلوب اولمامق لازم کلیر، فی الواقع "الصراط المستقیم" نظم جلیلی بزه علی الاطلاق` | `بوراده مطلوب اولاسون لازم کلیر۔ فی الواقع "الصراط المستقیم" نظام جلیلی بزه علی الاطام` | **9.41%** |
---
## 🚀 Usage Guide (`transformers`)
Below is the standard, native PyTorch & Hugging Face `transformers` implementation using `AutoProcessor` and `Qwen3_5ForConditionalGeneration`:
```python
import os
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
from qwen_vl_utils import process_vision_info
# Device & dtype settings
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32
model_id = "OttomanNLP/Azra-1-Mini-0.8b"
print("[INFO] Loading model and processor...")
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=dtype,
device_map="auto" if device == "cuda" else None,
trust_remote_code=True
)
model.eval()
print("[INFO] Model loaded successfully!")
def extract_text(image_path: str, prompt: str = "Görseldeki Osmanlıca metni transkribe et:") -> str:
"""Extract Ottoman text from a line image"""
if not os.path.exists(image_path):
return f"File not found: {image_path}"
image = Image.open(image_path).convert("RGB")
# Adjust dimensions to multiples of 64
w, h = image.size
new_w = ((w + 63) // 64) * 64
new_h = ((h + 63) // 64) * 64
if (new_w, new_h) != (w, h):
image = image.resize((new_w, new_h), Image.Resampling.LANCZOS)
messages = [{
"role": "user",
"content": [
{"type": "image", "image": image},
{"type": "text", "text": prompt}
]
}]
text_input = processor.apply_chat_template(
messages, tokenize=False, add_generation_prompt=True
)
image_inputs, _ = process_vision_info(messages)
inputs = processor(
text=[text_input],
images=image_inputs,
padding=True,
return_tensors="pt"
).to(device)
with torch.inference_mode():
generated_ids = model.generate(
**inputs,
max_new_tokens=512,
do_sample=False,
repetition_penalty=1.2,
no_repeat_ngram_size=3,
pad_token_id=processor.tokenizer.pad_token_id,
eos_token_id=processor.tokenizer.eos_token_id,
)
input_len = inputs.input_ids.shape[1]
output_text = processor.batch_decode(
generated_ids[:, input_len:],
skip_special_tokens=True,
clean_up_tokenization_spaces=False
)[0]
return output_text.strip()
if __name__ == "__main__":
image_path = "sample_line.png" # Path to line-level image
text = extract_text(image_path)
print("📝 Transcribed Text:\n", text)
```
---
## 🏷️ Model Details
- **Developed by:** OttomanNLP
- **Authors:** Gökhan Usta, Oğuz Alpoğlu, Fatih Günaydın
- **Model Type:** Vision-Language Model (VLM) for OCR
- **Language(s):** Ottoman Turkish (Osmanlıca)
- **Base Architecture:** Qwen3.5-Vision
- **Parameters:** ~0.8B
- **License:** Apache-2.0
---
## 📚 Citation
If you use this model or dataset in your research, please cite our paper:
```bibtex
@article{usta2026cross,
title={Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition},
author={Usta, G{\"o}khan and Alpo{\u{g}}lu, O{\u{g}}uz and G{\"u}nayd{\i}n, Fatih},
journal={Research Square (Preprint)},
year={2026},
doi={10.21203/rs.3.rs-10418926/v1},
note={Under Review at International Journal on Document Analysis and Recognition (IJDAR)}
}
```
**APA:**
> Usta, G., Alpoğlu, O., & Günaydın, F. (2026). *Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition*. Research Square Preprint. DOI: [10.21203/rs.3.rs-10418926/v1](https://doi.org/10.21203/rs.3.rs-10418926/v1)
---
## 📄 License & Attribution
This model is released under the **Apache 2.0 License**. Free for commercial and research use.
|