| ---
|
| license: cc-by-nc-4.0
|
| language:
|
| - ckb
|
| - ar
|
| tags:
|
| - handwritten-text-recognition
|
| - paragraph-recognition
|
| - kurdish
|
| - arabic
|
| - densenet
|
| - transformer
|
| - pytorch
|
| - safetensors
|
| pipeline_tag: image-to-text
|
| ---
|
| # ETE-KHPR: End-to-End Kurdish Handwritten Paragraph Recognition
|
| ### A DenseNet121-Transformer Architecture with Synthetic Paragraph Generation
|
|
|
| This repository contains the source code, trained models, and vocabularies for end-to-end Kurdish handwritten paragraph recognition without explicit line segmentation, with cross-script evaluation on Arabic (KHATT) and cross-dataset transfer to an external Kurdish dataset (DASNUS).
|
|
|
| ---
|
|
|
| ## Repository Structure
|
|
|
| ```
|
| KHPR/
|
| ├── DASTNUS-Kurdish-ParagraphHTR/ # Best Kurdish paragraph model
|
| │ ├── model.safetensors # Model weights
|
| │ ├── config.json # Architecture configuration
|
| │ ├── vocab.json # Character vocabulary (char → index)
|
| │ ├── idx_to_char.json # Reverse vocabulary (index → char)
|
| │ └── README.md # Model card
|
| │
|
| ├── DASNUS-Kurdish-ParagraphHTR/ # Model fine-tuned on external Kurdish dataset
|
| │ ├── model.safetensors
|
| │ ├── config.json
|
| │ ├── vocab.json
|
| │ ├── idx_to_char.json
|
| │ └── README.md
|
| │
|
| ├── KHATT-Arabic-ParagraphHTR/ # Model fine-tuned on KHATT Arabic dataset
|
| │ ├── model.safetensors
|
| │ ├── config.json
|
| │ ├── vocab.json # KHATT Arabic vocabulary (143 tokens)
|
| │ ├── idx_to_char.json
|
| │ └── README.md
|
| │
|
| ├── Scripts/
|
| │ ├── pretrain.py # Pre-training on synthetic paragraphs
|
| │ ├── finetune.py # Fine-tuning on real handwritten paragraphs
|
| │ ├── inference.py # Single image and batch inference
|
| │ └── generate_paragraphs.py # Synthetic paragraph generation
|
| │
|
| ├── Sample/
|
| │ ├── sample_paragraph.tif # Example Kurdish handwritten paragraph
|
| │ └── sample_paragraph.txt # Corresponding ground truth
|
| │
|
| ├── requirements.txt
|
| └── README.md
|
| ```
|
|
|
| ---
|
|
|
| ## Architecture
|
|
|
| | Component | Details |
|
| |-----------|---------|
|
| | CNN Backbone | DenseNet-121 (ImageNet pre-trained) |
|
| | Encoder | 3 Transformer encoder layers |
|
| | Decoder | 6 Transformer decoder layers |
|
| | Attention Heads | 8 |
|
| | Hidden Size | 256 |
|
| | Feed-Forward Dim | 2048 |
|
| | Positional Encoding | 2D sinusoidal (encoder) + 1D sinusoidal (decoder) |
|
| | Total Parameters | 22.7M |
|
|
|
| The model processes full paragraph images end-to-end and outputs the complete multi-line text, including line break positions, without any explicit line segmentation.
|
|
|
| ---
|
|
|
| ## Performance
|
|
|
| ### Kurdish — DASTNUS Unique Handwritten Paragraphs
|
|
|
| | Decoding Strategy | CER | WER | CRR (%) | WRR (%) |
|
| |---|---|---|---|---|
|
| | Greedy | 0.0721 | 0.3624 | 92.79 | 63.76 |
|
| | Beam-10 | 0.0706 | 0.3580 | 92.94 | 64.20 |
|
| | Beam-10 + 8-gram LM (w=0.6) | 0.0676 | 0.3422 | 93.24 | 65.78 |
|
| | Beam-10 + RoBERTa (w=0.1) | 0.0680 | 0.3484 | 93.20 | 65.16 |
|
|
|
| ### Cross-Script Evaluation — KHATT Arabic Handwritten Paragraphs
|
|
|
| | Model | CER | WER | CRR (%) |
|
| |---|---|---|---|
|
| | Proposed | 0.1394 | 0.5075 | 86.06 |
|
| | MSdocTr-Lite (reimplemented, same conditions) | 0.1622 | 0.5227 | 83.78 |
|
|
|
| ### Cross-Dataset Transfer — DASNUS External Kurdish Dataset
|
|
|
| | Setting | Training Samples | CER | WER | CRR (%) |
|
| |---|---|---|---|---|
|
| | Zero-shot | 0 | 0.2257 | 0.6206 | 77.43 |
|
| | Few-shot 10% | 184 | 0.1535 | 0.4757 | 84.65 |
|
| | Few-shot 50% | 922 | 0.1034 | 0.3609 | 89.66 |
|
| | Full fine-tune | 1,843 | 0.0856 | 0.3148 | 91.44 |
|
|
|
| ---
|
|
|
| ## Installation
|
|
|
| ```bash
|
| git clone https://huggingface.co/karez/KHPR
|
| cd KHPR
|
| pip install -r requirements.txt
|
| ```
|
|
|
| ---
|
|
|
| ## Quick Start
|
|
|
| ### Inference
|
|
|
| ```bash
|
| # Single paragraph image (with config auto-load)
|
| python Scripts/inference.py \
|
| --image Sample/sample_paragraph.tif \
|
| --model_path DASTNUS-Kurdish-ParagraphHTR/model.safetensors \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --config_path DASTNUS-Kurdish-ParagraphHTR/config.json
|
|
|
| # Directory of images with timing
|
| python Scripts/inference.py \
|
| --image_dir ./test_paragraphs \
|
| --model_path DASTNUS-Kurdish-ParagraphHTR/model.safetensors \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --config_path DASTNUS-Kurdish-ParagraphHTR/config.json \
|
| --show_timing \
|
| --output_file predictions.txt
|
|
|
| # Arabic model (KHATT)
|
| python Scripts/inference.py \
|
| --image Sample/arabic_paragraph.tif \
|
| --model_path KHATT-Arabic-ParagraphHTR/model.safetensors \
|
| --vocab_path KHATT-Arabic-ParagraphHTR/vocab.json \
|
| --config_path KHATT-Arabic-ParagraphHTR/config.json
|
| ```
|
|
|
| ### Synthetic Paragraph Generation
|
|
|
| ```bash
|
| # Full three-source generation (best configuration)
|
| python Scripts/generate_paragraphs.py \
|
| --unique_train_dir ./data/UniqueLines/Training \
|
| --fixed_train_dir ./data/FixedLines/Training \
|
| --synthetic_train_dir ./data/SyntheticLines/Training \
|
| --unique_val_dir ./data/UniqueLines/Validation \
|
| --fixed_val_dir ./data/FixedLines/Validation \
|
| --synthetic_val_dir ./data/SyntheticLines/Validation \
|
| --output_dir ./SyntheticParagraphs_12000 \
|
| --dataset_size 12000
|
| ```
|
|
|
| ### Pre-training
|
|
|
| ```bash
|
| # Pre-train on synthetic paragraphs (Kurdish, default settings)
|
| python Scripts/pretrain.py \
|
| --data_dir ./SyntheticParagraphs_12000 \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --output_dir ./output \
|
| --model_name pretrained_kurdish
|
|
|
| # Pre-train without curriculum learning
|
| python Scripts/pretrain.py \
|
| --data_dir ./SyntheticParagraphs_12000 \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --no_curriculum
|
| ```
|
|
|
| ### Fine-tuning
|
|
|
| ```bash
|
| # Fine-tune on DASTNUS unique handwritten paragraphs
|
| python Scripts/finetune.py \
|
| --data_dir ./data/UniqueHandwrittenParagraphs \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --pretrained_path ./output/pretrained_kurdish.pth \
|
| --output_dir ./output \
|
| --model_name finetuned_dastnus
|
|
|
| # Fine-tune on DASNUS external Kurdish dataset
|
| python Scripts/finetune.py \
|
| --data_dir ./data/DASNUS-Paragraphs \
|
| --vocab_path DASTNUS-Kurdish-ParagraphHTR/vocab.json \
|
| --pretrained_path ./output/pretrained_kurdish.pth \
|
| --output_dir ./output \
|
| --model_name finetuned_dasnus
|
|
|
| # Fine-tune on KHATT Arabic dataset
|
| python Scripts/finetune.py \
|
| --data_dir ./data/KHATT-Paragraphs \
|
| --vocab_path KHATT-Arabic-ParagraphHTR/vocab.json \
|
| --pretrained_path ./output/pretrained_khatt.pth \
|
| --output_dir ./output \
|
| --model_name finetuned_khatt
|
| ```
|
| ---
|
|
|
| ## Training Data
|
|
|
| ### DASTNUS and DASNUS Models
|
|
|
| | Data Source | Training | Validation | Testing |
|
| |---|---|---|---|
|
| | Unique handwritten paragraphs | 710 | 144 | 144 |
|
| | Synthetic paragraphs (pre-training) | 10,200 | 1,800 | — |
|
|
|
| Synthetic paragraphs were generated from DASTNUS line sources using the `generate_paragraphs.py` script, combining unique handwritten lines, Fixed handwrwritten lines and recipe-based synthetic handwritten lines with single-writer consistency, zero duplicate text orderings, and source-level isolation between splits.
|
|
|
| ### KHATT Model
|
|
|
| | Data Source | Training | Validation | Testing |
|
| |---|---|---|---|
|
| | Reconstructed KHATT paragraphs | 1,193 | 144 | 150 |
|
| | Synthetic paragraphs (pre-training) | 10,201 | 1,199 | — |
|
|
|
| Synthetic paragraphs for KHATT pre-training were generated by combining KHATT handwritten lines with Kurdish line sources from DASTNUS to provide richer visual diversity across handwriting styles within the same Arabic script family.
|
|
|
| ---
|
|
|
| ## Hardware
|
|
|
| Experiments were conducted on a workstation equipped with an Intel Core i9-14900K processor, 128 GB RAM, and an NVIDIA GeForce RTX 5090 GPU with 32 GB VRAM.
|
|
|
| ---
|
|
|
| ## Citation
|
|
|
| ```bibtex
|
| []
|
| ```
|
|
|
| ---
|
|
|
| ## License
|
|
|
| This repository is released for non-commercial scientific research purposes only under the CC-BY-NC-4.0 license. The data used in this research is available upon request for non-commercial scientific research purposes only. |