File size: 4,507 Bytes
1d87586
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ad70274
 
1d87586
ad70274
1d87586
b47f5d9
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1d87586
b47f5d9
 
 
1d87586
ad70274
 
1d87586
b47f5d9
1d87586
ad70274
1d87586
 
 
 
 
 
 
 
ad70274
1d87586
 
 
 
 
 
 
 
ad70274
1d87586
ad70274
 
b47f5d9
ad70274
1d87586
ad70274
1d87586
 
 
b47f5d9
 
1d87586
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
---
language:
- bo
license: mit
library_name: transformers
pipeline_tag: image-to-text
base_model: microsoft/trocr-base-stage1
tags:
- ocr
- tibetan
- document-ai
- trocr
- vision-encoder-decoder
---

# PechaBridgeOCR

Tibetan line OCR model (323.24M parameters, VIT encoder + TROCR decoder in a VisionEncoderDecoder architecture) fine-tuned for traditional Tibetan pecha scans with PechaBridge.

> **Important:** This checkpoint expects PechaBridge's `gray` preprocessing followed by a fixed `256×1024` resize. The `gray` pipeline uses the minimum RGB channel as grayscale and does not binarize the image. Use the PechaBridge CLI for the supported end-to-end path; raw image input or a generic `AutoTokenizer` path will not reproduce the training pipeline.

## Recommended usage — PechaBridge CLI

```bash
git clone https://github.com/CodexAITeam/PechaBridge.git && cd PechaBridge
pip install -r requirements.txt
python cli.py download-models

python cli.py batch-ocr \
    --ocr-model     models/ocr/PechaBridgeOCR \
    --line-model    models/line_segmentation/PechaBridgeLineSegmentation.pt \
    --layout-engine yolo_line \
    --ocr-engine    donut \
    --input-dir     /path/to/pecha/images
```

Each page image produces a `.txt` transcript and an `*_overlay.jpg` preview.

## Advanced: standalone Python usage

The tokenizer is backed by BoSentencePiece plus PechaBridge-specific special tokens. Load it with PechaBridge's adapter rather than `AutoTokenizer`.

```python
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from PIL import Image
from transformers import AutoImageProcessor, VisionEncoderDecoderModel
from pechabridge.ocr.preprocess_bdrc import BDRCPreprocessConfig, preprocess_image_bdrc
from pechabridge.ocr.sentencepiece_tokenizer_adapter import load_sentencepiece_tokenizer

model_dir = Path(snapshot_download("TibetanCodexAITeam/PechaBridgeOCR"))
model = VisionEncoderDecoderModel.from_pretrained(model_dir).eval()
image_processor = AutoImageProcessor.from_pretrained(model_dir, use_fast=False)
tokenizer = load_sentencepiece_tokenizer(model_dir)
image = Image.open('line_crop.png').convert('RGB')
cfg = BDRCPreprocessConfig.vit_defaults()
cfg = BDRCPreprocessConfig.from_dict({
    **cfg.to_dict(), 'binarize': False, 'gray_mode': 'min_rgb',
})
prepared = preprocess_image_bdrc(image=image, config=cfg).convert('RGB')
pixel_values = image_processor(images=prepared, return_tensors='pt').pixel_values
with torch.inference_mode():
    generated_ids = model.generate(pixel_values)
text = tokenizer.batch_decode(generated_ids, skip_special_tokens=True)[0]
print(text)
```

## Model details

- **Checkpoint**: `checkpoint-184000` (step 184000)
- **Image preprocessing pipeline**: `gray`
- **Input geometry**: fixed `256×1024` resize, RGB tensor normalized with mean/std `0.5`
- **Architecture**: VIT encoder + TROCR decoder, 323.24M parameters
- **Repro configuration included**: yes — preprocessing, generation, normalization, and metric definitions are in `repro/`
- **Training framework**: [PechaBridge](https://github.com/CodexAITeam/PechaBridge)
- **Training data**: Tibetan pecha line images from OpenPecha and BDRC collections

## Evaluation

| Checkpoint | Internal validation CER | Valid samples |
|---|---:|---:|
| `checkpoint-184000` | 0.5751% | 241 |

CER is the global character edit distance divided by total normalized reference length. This is a training-time internal validation result, not an independent cross-collection benchmark.

### External multi-source evaluation

On a separate 3,306-sample evaluation set, mean CER was **0.80%**. **1.63%** of samples exceeded 10% CER and **0.57%** exceeded 20% CER.

| Source dataset | Samples | Mean CER |
|---|---:|---:|
| OCR-Norbuketaka | 2,350 | 0.43% |
| OCR-Google_Books | 812 | 1.59% |
| OCR-Lhasakanjur | 118 | 1.92% |
| OCR-Drutsa | 14 | 3.54% |
| OCR-Betsug | 12 | 3.86% |

This externally supplied result summary was not independently reproduced from artifacts included in this release.

## Intended use and limitations

- Intended for OCR of individual Tibetan pecha text-line crops.
- Page-level use requires a separate line-segmentation/layout stage.
- Performance may degrade on other scripts, modern book layouts, handwriting, severe blur, unusual scan colors, or collections not represented during training.
- Outputs should be reviewed before scholarly, archival, or other high-impact use.

## License

MIT. See the PechaBridge repository for the project license.