File size: 9,856 Bytes
746f31f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9d3c51c
405c1ba
9d3c51c
 
 
746f31f
 
 
 
5c454fa
 
 
746f31f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
7c1c429
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5c454fa
746f31f
5c454fa
746f31f
 
5c454fa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
746f31f
5c454fa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
746f31f
 
 
 
 
 
 
3c9bdc8
746f31f
 
5c454fa
746f31f
 
 
 
 
3c9bdc8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
746f31f
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
---
license: apache-2.0
language:
- ott
- tr
- ar
pipeline_tag: image-to-text
tags:
- ocr
- ottoman
- qwen
- vllm
- vision-language
metrics:
- accuracy
- cer
---

# Azra 1-Mini (0.8B) — Ottoman Turkish OCR Model

<p align="center">
  <img src="banner.png" alt="Azra 1-Mini Banner" width="45%">
</p>


**Azra 1-Mini 0.8b** is a lightweight, high-performance Vision-Language OCR model specialized in Ottoman Turkish text transcription across both **Nesih** (printed/calligraphic) and **Rika/Riqa** (handwritten) scripts. 

Despite having only **0.8 billion parameters**, Azra 1-Mini achieves state-of-the-art accuracy on Ottoman Turkish OCR tasks, outperforming significantly larger proprietary models.

> ⚠️ **Important Note on Input Resolution & Segmentation:**
> This model has been fine-tuned and optimized specifically for **line-level text images (satır bazlı görüntüler)**. It may not achieve optimal accuracy directly on full-page images without prior text line cropping/segmentation.

---

## 📊 Benchmark Results & Performance Comparison

The model was evaluated against leading proprietary Vision-Language models on standard Ottoman Turkish test sets using character accuracy (`100% - CER`).

### 1. Nesih Script Test Set (Printed / Calligraphic)

| Model | Success Rate (%) | Rank |
| :--- | :---: | :---: |
| **Gemini 3.1 Pro** | **82.82%** | 👑 1st |
| **Azra 1-Mini 0.8b** | **80.23%** | 🥈 2nd |
| **Qwen 3.8 Max** | 74.26% | 🥉 3rd |

### 2. Rika Script Test Set (Handwritten)

| Model | Success Rate (%) | Rank |
| :--- | :---: | :---: |
| **Azra 1-Mini 0.8b** | **67.08%** | 👑 **1st (Winner)** |
| **Gemini 3.1 Pro** | 58.46% | 🥈 2nd |
| **Qwen 3.8 Max** | 54.99% | 🥉 3rd |

> 🌟 **Key Highlight:** Azra 1-Mini 0.8b achieves **1st place on the handwritten Rika dataset (67.08%)**, significantly outperforming both Gemini 3.1 Pro and Qwen 3.8 Max while running efficiently at sub-billion parameter scale.

---

## 📷 Qualitative Results & Sample Transcriptions

Below are top qualitative predictions generated by **Azra 1-Mini 0.8b** from the evaluation test sets:

### 1. Nesih Script Samples (Printed / Calligraphic)

| Image | Ground Truth (GT) | Model Prediction (Azra 1-Mini) | CER |
| :---: | :--- | :--- | :---: |
| <img src="assets/samples/nesih_3263_759633_eSc_line_f5695a2d.png" height="35"> | `امّا اری وابدار ونازک اولور هر اعجک زمان غرسی` | `امّا اری وابدار ونازک اولور هر اعجک زمان غرسی` | **0.00%** |
| <img src="assets/samples/nesih_3263_759704_eSc_line_211ddfad.png" height="35"> | `هلاک ایدر ازایسه علاج ایله خلاص اولور` | `هلاک ایدر ازایسه علاج ایله خلاص اولور` | **0.00%** |
| <img src="assets/samples/nesih_3263_759800_eSc_line_83cb8596.png" height="35"> | `یافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردل` | `یافوجی ایچنه دوشرلر اوّل التنه وافرد وکلمش خردل` | **0.00%** |
| <img src="assets/samples/nesih_3263_759772_eSc_line_21130720.png" height="35"> | `اغزی محکم باغلنوب اول بوداق اکلوب یره کوملسه وقت` | `اغزی محکمه باغلنوب اول بوداق اکلوب یره کوملسه وقت` | **2.08%** |
| <img src="assets/samples/nesih_3263_759629_eSc_line_e9d59dc1.png" height="35"> | `دکمک زماندر دیمش یعنی آیک نقصانی زمانی که اوّل` | `دکک زماندر دیمش یعنی آیک نقصانی زمانی که اوّل` | **2.17%** |

### 2. Rika Script Samples (Handwritten)

| Image | Ground Truth (GT) | Model Prediction (Azra 1-Mini) | CER |
| :---: | :--- | :--- | :---: |
| <img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-001.png" height="35"> | `دیمک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر و بو هدایت ایکی` | `دیک طلب و دیانت بزدن، دین، شریعت، هدایت اللهدندر۔ و بوهدایت ایکی` | **4.62%** |
| <img src="assets/samples/riqa_dfc88259-fa14-4ae1-b772-f5ea65db8b6c-015.png" height="35"> | `ایتمک دون بنی تنویر ایدن کونشک یارین تنویر ایدهمیهجکنی ادعا ایتمک کبی قانون استقرایی انکاردر۔` | `ایتمک دوند بنی تنویر ایدن کونشک یارین تنویر ایدرمهجیکنی ادعا ایتمک کبی قانون استقرالی انکاردر۔` | **5.38%** |
| <img src="assets/samples/riqa_c25bfa03-dbf2-41e2-8164-d639b87fbede-015.png" height="35"> | `ایمانده نه قدر بیوک بر سعادت و نعمت؛ و نه قدر بیوک بر لذت و راحت بولوندیغنی اڭلامق` | `ایمانده نه قدر یوک بر سعادت ونعمت و نه قدر یوک بر لذت و راحت بولوندیغی اشلامم` | **8.54%** |
| <img src="assets/samples/riqa_e2da2c9f-5dc9-4ffa-8241-d13cffc1caef-020.png" height="35"> | `”الله تعالی ابراهیم علیه السلامه وحی ایدوب دیدی که: اسماعیل حقندهکی دعاکی قبول ایتدم و اونی` | `"الله تعالی ابراهیم علمه السلام دحی ایدوب دیدی کی: اسماعیل حقندهکی دعاک قبول ایتدم واولی` | **8.79%** |
| <img src="assets/samples/riqa_81f5dd07-3ee3-4b26-a02a-a589e289f11c-003.png" height="35"> | `بوراده مطلوب اولمامق لازم کلیر، فی الواقع "الصراط المستقیم" نظم جلیلی بزه علی الاطلاق` | `بوراده مطلوب اولاسون لازم کلیر۔ فی الواقع "الصراط المستقیم" نظام جلیلی بزه علی الاطام` | **9.41%** |

---

## 🚀 Usage Guide (`transformers`)

Below is the standard, native PyTorch & Hugging Face `transformers` implementation using `AutoProcessor` and `Qwen3_5ForConditionalGeneration`:

```python
import os
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
from qwen_vl_utils import process_vision_info

# Device & dtype settings
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.float16 if device == "cuda" else torch.float32

model_id = "OttomanNLP/Azra-1-Mini-0.8b"

print("[INFO] Loading model and processor...")
processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
    model_id,
    torch_dtype=dtype,
    device_map="auto" if device == "cuda" else None,
    trust_remote_code=True
)
model.eval()
print("[INFO] Model loaded successfully!")

def extract_text(image_path: str, prompt: str = "Görseldeki Osmanlıca metni transkribe et:") -> str:
    """Extract Ottoman text from a line image"""
    if not os.path.exists(image_path):
        return f"File not found: {image_path}"
        
    image = Image.open(image_path).convert("RGB")
    
    # Adjust dimensions to multiples of 64
    w, h = image.size
    new_w = ((w + 63) // 64) * 64
    new_h = ((h + 63) // 64) * 64
    if (new_w, new_h) != (w, h):
        image = image.resize((new_w, new_h), Image.Resampling.LANCZOS)
    
    messages = [{
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": prompt}
        ]
    }]
    
    text_input = processor.apply_chat_template(
        messages, tokenize=False, add_generation_prompt=True
    )
    image_inputs, _ = process_vision_info(messages)
    
    inputs = processor(
        text=[text_input],
        images=image_inputs,
        padding=True,
        return_tensors="pt"
    ).to(device)
    
    with torch.inference_mode():
        generated_ids = model.generate(
            **inputs,
            max_new_tokens=512,
            do_sample=False,
            repetition_penalty=1.2,
            no_repeat_ngram_size=3,
            pad_token_id=processor.tokenizer.pad_token_id,
            eos_token_id=processor.tokenizer.eos_token_id,
        )
    
    input_len = inputs.input_ids.shape[1]
    output_text = processor.batch_decode(
        generated_ids[:, input_len:],
        skip_special_tokens=True,
        clean_up_tokenization_spaces=False
    )[0]
    
    return output_text.strip()

if __name__ == "__main__":
    image_path = "sample_line.png"  # Path to line-level image
    text = extract_text(image_path)
    print("📝 Transcribed Text:\n", text)
```

---

## 🏷️ Model Details

- **Developed by:** OttomanNLP
- **Authors:** Gökhan Usta, Oğuz Alpoğlu, Fatih Günaydın
- **Model Type:** Vision-Language Model (VLM) for OCR
- **Language(s):** Ottoman Turkish (Osmanlıca)
- **Base Architecture:** Qwen3.5-Vision
- **Parameters:** ~0.8B
- **License:** Apache-2.0

---

## 📚 Citation

If you use this model or dataset in your research, please cite our paper:

```bibtex
@article{usta2026cross,
  title={Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition},
  author={Usta, G{\"o}khan and Alpo{\u{g}}lu, O{\u{g}}uz and G{\"u}nayd{\i}n, Fatih},
  journal={Research Square (Preprint)},
  year={2026},
  doi={10.21203/rs.3.rs-10418926/v1},
  note={Under Review at International Journal on Document Analysis and Recognition (IJDAR)}
}
```

**APA:**
> Usta, G., Alpoğlu, O., & Günaydın, F. (2026). *Cross-Lingual Transfer Learning and Autonomous Data Bootstrapping for VLM-Based Ottoman Turkish Handwritten Text Recognition*. Research Square Preprint. DOI: [10.21203/rs.3.rs-10418926/v1](https://doi.org/10.21203/rs.3.rs-10418926/v1)

---

## 📄 License & Attribution

This model is released under the **Apache 2.0 License**. Free for commercial and research use.