| --- |
| license: apache-2.0 |
| library_name: rfdetr |
| pipeline_tag: object-detection |
| language: |
| - bo |
| tags: |
| - tibetan |
| - document-layout-analysis |
| - rf-detr |
| - rfdetr |
| - object-detection |
| - bounding-box |
| - BDRC |
| datasets: |
| - BDRC/TDLA-Training-Dataset-v2 |
| metrics: |
| - name: canonical mean AP50 (test) |
| type: mAP |
| value: 0.960 |
| model-index: |
| - name: Tibetan-Modern-Book-Layout-Detection-RFDETR |
| results: |
| - task: |
| type: object-detection |
| dataset: |
| name: TDLA-Training-Dataset-v2 (test) |
| type: BDRC/TDLA-Training-Dataset-v2 |
| metrics: |
| - type: mAP |
| name: mAP@0.5 (native, test) |
| value: 0.996 |
| - type: mAP |
| name: mAP@0.5:0.95 (native, test) |
| value: 0.813 |
| --- |
| |
| # Tibetan Modern Book Layout Detection (RF-DETR-L) |
|
|
| An **RF-DETR-L** ([Roboflow](https://github.com/roboflow/rf-detr), DINOv2 backbone) |
| object detector that locates the four structural regions of a **modern Tibetan |
| book** page β **header**, **text-area**, **footnote**, **footer** β as a |
| preprocessing step for OCR and etext production. |
|
|
| - **Training code, recipes & write-up:** [buda-base/tibetan-book-layout-analysis](https://github.com/buda-base/tibetan-book-layout-analysis) |
| - **Dataset:** [BDRC/TDLA-Training-Dataset-v2](https://huggingface.co/datasets/BDRC/TDLA-Training-Dataset-v2) (gated, fair-use) |
|
|
| > This is one of several architectures BDRC fine-tuned on the same labels and |
| > recipe to test how much the choice of architecture matters (see the |
| > [blog post](https://github.com/buda-base/tibetan-book-layout-analysis/blob/main/BLOGPOST.md)). |
| > The primary, production release is the RT-DETR-l fine-tune at |
| > [BDRC/Tibetan-Modern-Book-Layout-Detection-RTDETR](https://huggingface.co/BDRC/Tibetan-Modern-Book-Layout-Detection-RTDETR). |
| > This RF-DETR-L checkpoint matches it almost exactly on every metric, and is |
| > published as a **permissively-licensed (Apache-2.0) alternative** for anyone |
| > who can't use the RT-DETR-l release's AGPL-derived training stack. |
|
|
| ## Model description |
|
|
| This is an RF-DETR-L fine-tuned with the **`tam2col`** labelling scheme: text-area |
| boxes merged into one envelope per page, except on genuine two-column pages, |
| where it keeps one box per column. Header and footer are kept as separate |
| classes (they can be combined losslessly downstream). |
|
|
| | Property | Value | |
| | --- | --- | |
| | Architecture | RF-DETR-L (via [`rfdetr`](https://github.com/roboflow/rf-detr), DINOv2-small windowed backbone) | |
| | Task | Object detection | |
| | Base checkpoint | `rf-detr-large-2026.pth` (Roboflow, Apache-2.0) | |
| | Resolution | 1008 Γ 1008 | |
| | Number of classes | 4 (+ background) | |
| | Framework | `rfdetr` (`RFDETRLarge`) | |
| | Weights file | `rfdetr_tibetan_book_layout.pth` | |
|
|
| ## Classes |
|
|
| Note: on the raw checkpoint, class id `0` is a reserved "none"/background |
| slot, so predicted class ids are shifted by one β see `infer.py`. |
|
|
| | Our class | Class name | Description | |
| | -- | --------- | --------------------- | |
| | 0 | header | running title / marginal text at top or side | |
| | 1 | text-area | main body text (one box per column) | |
| | 2 | footnote | notes below the text area | |
| | 3 | footer | folio numbers / marginal text at bottom or side | |
|
|
| ## Recommended usage β per-class confidence thresholds |
|
|
| Like the primary RT-DETR-l release, this detector is recall-happy, so the best |
| operating point differs by class. These are each class's own max-F1 confidence |
| from a native per-class sweep on the test set: |
|
|
| | class | recommended conf | |
| | --- | --- | |
| | header (0) | **0.46** | |
| | text-area (1) | **0.32** | |
| | footnote (2) | **0.26** | |
| | footer (3) | **0.52** | |
|
|
| If you need a single global threshold, **0.30** is the best compromise (it is |
| also the operating point used for the cross-architecture comparison in the |
| blog post). |
|
|
| ### Inference |
|
|
| ```python |
| from rfdetr import RFDETRLarge |
| |
| model = RFDETRLarge.from_checkpoint("rfdetr_tibetan_book_layout.pth") |
| CLASS_CONF = {0: 0.46, 1: 0.32, 2: 0.26, 3: 0.52} # header, text-area, footnote, footer |
| names = {0: "header", 1: "text-area", 2: "footnote", 3: "footer"} |
| |
| det = model.predict("page.jpg", threshold=min(CLASS_CONF.values()), shape=(1024, 1024)) |
| for box, cls_id, score in zip(det.xyxy, det.class_id, det.confidence): |
| cls = int(cls_id) - 1 # class 0 on the checkpoint is background |
| if cls < 0 or cls > 3 or score < CLASS_CONF[cls]: |
| continue |
| print(names[cls], round(float(score), 3), box.tolist()) |
| ``` |
|
|
| A ready-made CLI (`infer.py`) with the thresholds baked in is included in this |
| repo. |
|
|
| ### Downloading the weights |
|
|
| ```python |
| from huggingface_hub import hf_hub_download |
| path = hf_hub_download("BDRC/Tibetan-Modern-Book-Layout-Detection-RFDETR", |
| "rfdetr_tibetan_book_layout.pth") |
| ``` |
|
|
| ## Performance |
|
|
| Evaluated on the **held-out test split** (860 images) of |
| `BDRC/TDLA-Training-Dataset-v2`. |
|
|
| ### Native 4-class metrics |
|
|
| | class | P | R | F1 | mAP@0.5 | mAP@0.5:0.95 | |
| | --- | --- | --- | --- | --- | --- | |
| | header | 0.977 | 0.960 | 0.968 | 0.986 | 0.757 | |
| | text-area | 0.986 | 0.994 | 0.990 | 0.996 | 0.982 | |
| | footnote | 0.913 | 0.933 | 0.923 | 0.973 | 0.816 | |
| | footer | 0.961 | 0.964 | 0.963 | 0.969 | 0.697 | |
| | **overall** | β | β | **0.961** | **0.981** | **0.813** | |
|
|
| The **mAP** columns are threshold-independent β they integrate over the full |
| precision/recall curve (every confidence), so they do not depend on any |
| operating threshold. The **P / R / F1** columns are reported at each class's |
| own **max-F1 confidence** (see the per-class thresholds above), *not* at a |
| fixed threshold. |
|
|
| ### Canonical 3-class metrics |
|
|
| Header + footer are combined into one `header-footer` class (matched |
| individually), text-area is compared as one merged envelope, and footnote is |
| left as-is β a fair space in which every fine-tuned architecture in the blog |
| post was compared. |
|
|
| | class | best-F1 | |
| | --- | --- | |
| | header-footer | 0.963 | |
| | text-area | 0.994 | |
| | footnote | 0.923 | |
| | **mean F1** | **0.960** (@ conf 0.30) | |
|
|
| This matches BDRC's primary RT-DETR-l fine-tune (mean F1 0.960) almost exactly. |
|
|
| ### Contamination (the metric that actually matters for OCR) |
|
|
| Of the ground-truth headers/footers and footnotes this model *misses*, the |
| share that get folded into its predicted text-area box (silently corrupting |
| downstream OCR) rather than dropped cleanly: |
|
|
| | region | detected | folded into text-area | |
| | --- | --- | --- | |
| | header/footer | 97% | 0.1% | |
| | footnote | 93% | 7% | |
|
|
| For comparison, off-the-shelf systems in the same evaluation ranged from 1.2% |
| to 56% on header/footer contamination alone β see the |
| [blog post](https://github.com/buda-base/tibetan-book-layout-analysis/blob/main/BLOGPOST.md) |
| for the full picture. |
|
|
| ## Training details |
|
|
| | Parameter | Value | |
| | --- | --- | |
| | Base checkpoint | `rf-detr-large-2026.pth` (Roboflow, Apache-2.0) | |
| | Resolution | 1008 | |
| | Epochs | 100 planned, early-stopped ~epoch 59 (patience 20) | |
| | Batch size | 8 | |
| | GPU | single NVIDIA A10G (24 GB) | |
|
|
| - **Dataset:** [BDRC/TDLA-Training-Dataset-v2](https://huggingface.co/datasets/BDRC/TDLA-Training-Dataset-v2) β 8,325 images (6,751 train / 714 val / 860 test), volume-level leakage-free splits, augmented images confined to train. |
| - **Label variant (`tam2col`):** text-area boxes merged per page except on two-column pages; built with `data/build_curricula.py` in the GitHub repo. |
|
|
| ## Intended use |
|
|
| Automatic layout detection of **modern Tibetan book** pages, as a |
| preprocessing step for OCR pipelines, document digitization, structured text |
| extraction, and digital-library indexing. |
|
|
| ## Limitations |
|
|
| - Trained on **modern Tibetan books**; performance on traditional pecha, |
| manuscripts, or woodblock prints is not characterized and may be poor. |
| - Optimized for 1008β1024 px input; very high-resolution scans may benefit |
| from a higher inference resolution. |
| - The footnote class is rare in the source material (β1.4% of boxes); recall |
| is strong on the test set but the class remains the least-represented, and |
| this checkpoint's footnote contamination (7%) is higher than BDRC's primary |
| RT-DETR-l release (2%) despite a comparable footnote F1 β see the blog |
| post's discussion of why F1 and contamination aren't interchangeable. |
| - Header/footer boxes are small and easy to over-predict β use the |
| recommended per-class thresholds above. |
|
|
| ## License |
|
|
| The model weights are released under the **Apache License 2.0**, matching the |
| license of the RF-DETR-L base checkpoint they were fine-tuned from. The **page |
| images used for training are not covered by any content license** β they are |
| BDRC library scans distributed on a fair-use basis. You are solely responsible |
| for your own copyright / rights analysis before use; BDRC accepts no liability |
| for misuse. See the |
| [dataset card](https://huggingface.co/datasets/BDRC/TDLA-Training-Dataset-v2) |
| for the full notice. |
|
|
| ## Acknowledgements |
|
|
| Developed by the [Buddhist Digital Resource Center (BDRC)](https://www.bdrc.io) |
| for the BDRC Etext Corpus, with annotations produced and consolidated on the |
| Ultralytics platform. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @software{bdrc_tibetan_book_layout_rfdetr_2026, |
| title = {Tibetan Modern Book Layout Detection (RF-DETR-L)}, |
| author = {Buddhist Digital Resource Center (BDRC)}, |
| year = {2026}, |
| url = {https://huggingface.co/BDRC/Tibetan-Modern-Book-Layout-Detection-RFDETR} |
| } |
| ``` |
|
|