--- license: apache-2.0 tags: - object-detection - document-layout-analysis - indic - yolo - doclayout-yolo - document-ai library_name: doclayout-yolo pipeline_tag: object-detection --- # DocLayout-YOLO-Indic **Real-time document layout detection for 12 Indic scripts**, built on DocLayout-YOLO (YOLOv10-m + GL-CRM). This repository accompanies an M.Tech dissertation (Vignesh P, BITS Pilani WILP). It contains four model checkpoints that together form a **controlled ablation study** — and its central result is an honest **negative finding**: the two proposed adaptation techniques did *not* improve over a simple baseline. That finding, proven and explained, is the main scientific contribution. --- ## 1. What problem does this solve? Document Layout Analysis (DLA) is the first step in digitising any document: before OCR can read the text, a model must find the *structure* — which regions are paragraphs, headings, tables, figures, lists, and so on. Every downstream stage inherits the errors of this step. Fast, modern DLA detectors are trained almost entirely on **English and Chinese** documents and generalise poorly to **Indic scripts**, because those scripts are *visually* different: - **Shirorekha** — the continuous horizontal headstroke of Devanagari, Bengali and Gurmukhi — looks like a table border to a Latin-trained model. - **Conjunct consonants and stacked vowel marks (matras)** break line-height assumptions. - **Urdu** is written right-to-left. - Real pages **mix scripts** and contain **rare region types** (sidebars, pull-quotes). Over 1.4 billion people use Indic languages, yet no public real-time detector had been demonstrably adapted *and* evaluated across the major scripts. This project builds one — and rigorously tests whether its own adaptation ideas actually help. --- ## 2. The model - **Architecture:** DocLayout-YOLO = YOLOv10-m backbone + **GL-CRM** (Global-to-Local Controllable Receptive Module) + PAN-FPN neck + dual (one-to-many / one-to-one) detection heads. - **Parameters:** ~19.97 M (20.01 M at 42 classes). - **Input:** 1024×1024 (also evaluated at 640 and 1280). - **Vision-only:** it does not read text or depend on OCR — it responds purely to visual structure. - **Speed:** >100 FPS on an NVIDIA A100 — real-time. --- ## 3. The checkpoints (a controlled ablation) The four files differ **only** in what training happened *before* the final fine-tuning on the IndicDLP benchmark. Because everything else is held identical, comparing them isolates the effect of each proposed contribution — this is what licenses causal, not correlational, conclusions. | File | Training pipeline | Classes | Val mAP@[.5:.95] | **Test mAP@[.5:.95]** | |------|-------------------|---------|------------------|------------------------| | **`config_A_reported_42cls.pt`** | Public checkpoint → fine-tune. **No** synthetic pretraining, **no** self-training. **← the reported model** | 42 | 0.379 | **0.364** | | `config_B_synthetic_42cls.pt` | Synthetic pretrain → fine-tune. *Isolates synthetic pretraining.* | 42 | 0.353 | 0.252 | | `config_C_fullpipeline_42cls.pt` | Synthetic → self-train → fine-tune (the full proposed pipeline). | 42 | 0.328 | — | | `selftrained_intermediate_9cls.pt` | Intermediate self-trained checkpoint, *before* fine-tuning. | 9 | — | 0.577 (in-domain, Bengali) | **Which to use:** for inference, use **`config_A_reported_42cls.pt`** — it is the best model and the one the dissertation reports. (It is a stripped, deployment-ready checkpoint, ~41 MB; the others retain optimizer state and are ~242 MB.) ### The two class ontologies - **9 classes** (intermediate stages): `text_body, headline, table, figure, caption, advertisement, sidebar, pull-quote, decorative-frame` — a coarse, script-agnostic set used for synthetic pretraining and self-training. - **42 classes** (final): the full IndicDLP ontology, grouped into seven families — body text, heading hierarchy, list hierarchy, figures/tables/formulae, Q&A structure, navigation/metadata, and publishing furniture. The detection head is **re-shaped** between stages (27 → 9 → 42 classes); the backbone and neck weights transfer across all stages, and only the final classifier changes width. --- ## 4. Headline results (reported model, `config_A`) - **0.525 mAP@0.5** and **0.364 mAP@[0.5:0.95]** on the **11,633-image** IndicDLP test set (180,039 instances). - **Consistent across all 12 scripts** — per-script mAP@[.5:.95] spans only **0.304–0.426**, with no script failing catastrophically. (Best: Assamese 0.426; lowest: English 0.299.) - **>100 FPS** on A100 — real-time preserved. - **Precision 0.590 / Recall 0.506** overall. ### Scalability (accuracy vs. speed) | Input size | Test mAP@[.5:.95] | Note | |------------|-------------------|------| | 640 px | 0.329 | fastest; ~90% of peak accuracy at a fraction of the latency | | 1024 px | **0.364** | peak accuracy (training resolution) | | 1280 px | 0.353 | strictly dominated — slower *and* less accurate (train-test resolution mismatch) | --- ## 5. The key finding (a rigorous negative result) **Neither proposed contribution improved over plainly fine-tuning the public checkpoint.** Config A (no synthetic, no self-training) is the best model — by **+11.2 mAP** on the test set over Config B. **This is statistically significant across all 12 scripts:** - Wilcoxon signed-rank: **p < 0.001** (p = 0.00049) - Paired t-test: **p < 10⁻¹⁰** - Cohen's d = **7.37** (a very large effect) - All 12/12 scripts improved under Config A; 95% CI on the gain [0.096, 0.114]. **Mechanism — catastrophic forgetting.** The validation→test gap *widens* from 2.6 to 11.2 mAP under synthetic pretraining — the signature of damaged generalisation, not mere no-gain. The synthetic corpus (23 templates, 9 classes) is narrower than the base model's original DocSynth-300K pretraining; 15 epochs to 0.985 synthetic-validation over-specialised the backbone, and the coarse 9-class intermediate ontology erased fine distinctions the 42-class task then had to relearn from only 12,082 fine-labelled images. > **In short:** the techniques failed, but the science succeeded — a falsified hypothesis, proven under > controlled conditions and mechanistically explained, that saves others from the same mistake. ### Note on `config_B` and `config_C` These are provided **for transparency and reproducibility of the ablation** — they are *not* recommended for deployment, as they under-perform Config A. They let others verify the negative result independently. --- ## 6. Known limitations - **Hierarchical-depth confusion is the dominant error.** The model localises headings/lists but cannot infer nesting *depth* — recall collapses with depth (section-title 0.66 → sub 0.23 → subsub 0.01). This is **architectural** (a single-shot detector classifies from local appearance; depth is relational) and **not Indic-specific** — it affects English identically. - **Urdu is untested.** Urdu was synthesised and trained on, but the IndicDLP benchmark contains no Urdu, so right-to-left performance is unevaluated. - **An auxiliary script-classification head** was designed but not implemented (future work). - Results use a **12,082-image training subset** (compute-bounded); the reported mAP is therefore a lower bound, not a ceiling. --- ## 7. Usage ```python from huggingface_hub import hf_hub_download from doclayout_yolo import YOLOv10 # download the reported model from this repo ckpt = hf_hub_download("VigneshPR/doclayout-yolo-indic", "config_A_reported_42cls.pt") model = YOLOv10(ckpt) # run on a document page result = model.predict("page.jpg", imgsz=1024, conf=0.25)[0] result.plot() # annotated image (numpy array, BGR) print(len(result.boxes), "regions") for c in result.boxes.cls.tolist(): print(result.names[int(c)]) ``` Install the runtime: ```bash pip install git+https://github.com/opendatalab/DocLayout-YOLO.git ``` --- ## 8. Training setup - **Platform:** Google Colab Pro+ (no institutional cluster). Training on **A100 40 GB**; evaluation on **Tesla T4** to conserve compute. - **Data engineering under a 15 GB Drive limit:** partial-tar extraction reduced the 42 GB IndicDLP download to ~5 GB; bulky image sets kept on ephemeral local disk, only labels/checkpoints persisted. - **Fine-tuning:** ~30 epochs per config, batch 16 @ 1024, AMP, identical settings across all three ablation configs. - **Self-training (CBST):** class-balanced pseudo-labelling on the unlabelled BaDLAD Bengali pool, with provably disjoint splits (leakage-audited: 20,365 − 10,000 = 10,365). --- ## 9. Related resources - **Synthetic corpus:** [VigneshPR/indicsynth-150k](https://huggingface.co/datasets/VigneshPR/indicsynth-150k) — the 150,000-page SA-MCB dataset used for pretraining. - **Benchmark (working mirror):** [VigneshPR/IndicDLP](https://huggingface.co/datasets/VigneshPR/IndicDLP) — mirror of the IndicDLP benchmark (original authors: IndicDLP, ICDAR 2025); credit and licence belong to them. --- ## 10. Citation If you use these models or the finding, please cite the dissertation: ```bibtex @mastersthesis{vignesh2026doclayoutindic, title = {DocLayout-YOLO-Indic: Cross-Script Document Layout Analysis for Indic Scripts using Synthetic Pretraining, Self-Training and Controlled Ablation}, author = {Vignesh P}, school = {BITS Pilani (WILP)}, year = {2026} } ``` *Built on [DocLayout-YOLO](https://github.com/opendatalab/DocLayout-YOLO) (YOLOv10-m + GL-CRM).*