File size: 9,607 Bytes
587d6aa
 
a0b8e0e
 
 
 
 
 
 
587d6aa
a0b8e0e
587d6aa
 
 
 
a0b8e0e
587d6aa
a0b8e0e
 
 
 
587d6aa
a0b8e0e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
587d6aa
a0b8e0e
 
587d6aa
a0b8e0e
 
 
 
 
587d6aa
a0b8e0e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
587d6aa
 
 
 
 
a0b8e0e
587d6aa
 
a0b8e0e
 
587d6aa
a0b8e0e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
587d6aa
 
a0b8e0e
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
---
license: apache-2.0
tags:
  - object-detection
  - document-layout-analysis
  - indic
  - yolo
  - doclayout-yolo
  - document-ai
library_name: doclayout-yolo
pipeline_tag: object-detection
---

# DocLayout-YOLO-Indic

**Real-time document layout detection for 12 Indic scripts**, built on DocLayout-YOLO (YOLOv10-m + GL-CRM).

This repository accompanies an M.Tech dissertation (Vignesh P, BITS Pilani WILP). It contains four model
checkpoints that together form a **controlled ablation study** β€” and its central result is an honest
**negative finding**: the two proposed adaptation techniques did *not* improve over a simple baseline.
That finding, proven and explained, is the main scientific contribution.

---

## 1. What problem does this solve?

Document Layout Analysis (DLA) is the first step in digitising any document: before OCR can read the text,
a model must find the *structure* β€” which regions are paragraphs, headings, tables, figures, lists, and so on.
Every downstream stage inherits the errors of this step.

Fast, modern DLA detectors are trained almost entirely on **English and Chinese** documents and generalise
poorly to **Indic scripts**, because those scripts are *visually* different:

- **Shirorekha** β€” the continuous horizontal headstroke of Devanagari, Bengali and Gurmukhi β€” looks like a table border to a Latin-trained model.
- **Conjunct consonants and stacked vowel marks (matras)** break line-height assumptions.
- **Urdu** is written right-to-left.
- Real pages **mix scripts** and contain **rare region types** (sidebars, pull-quotes).

Over 1.4 billion people use Indic languages, yet no public real-time detector had been demonstrably adapted
*and* evaluated across the major scripts. This project builds one β€” and rigorously tests whether its own
adaptation ideas actually help.

---

## 2. The model

- **Architecture:** DocLayout-YOLO = YOLOv10-m backbone + **GL-CRM** (Global-to-Local Controllable Receptive Module) + PAN-FPN neck + dual (one-to-many / one-to-one) detection heads.
- **Parameters:** ~19.97 M (20.01 M at 42 classes).
- **Input:** 1024Γ—1024 (also evaluated at 640 and 1280).
- **Vision-only:** it does not read text or depend on OCR β€” it responds purely to visual structure.
- **Speed:** >100 FPS on an NVIDIA A100 β€” real-time.

---

## 3. The checkpoints (a controlled ablation)

The four files differ **only** in what training happened *before* the final fine-tuning on the IndicDLP
benchmark. Because everything else is held identical, comparing them isolates the effect of each
proposed contribution β€” this is what licenses causal, not correlational, conclusions.

| File | Training pipeline | Classes | Val mAP@[.5:.95] | **Test mAP@[.5:.95]** |
|------|-------------------|---------|------------------|------------------------|
| **`config_A_reported_42cls.pt`** | Public checkpoint β†’ fine-tune. **No** synthetic pretraining, **no** self-training. **← the reported model** | 42 | 0.379 | **0.364** |
| `config_B_synthetic_42cls.pt` | Synthetic pretrain β†’ fine-tune. *Isolates synthetic pretraining.* | 42 | 0.353 | 0.252 |
| `config_C_fullpipeline_42cls.pt` | Synthetic β†’ self-train β†’ fine-tune (the full proposed pipeline). | 42 | 0.328 | β€” |
| `selftrained_intermediate_9cls.pt` | Intermediate self-trained checkpoint, *before* fine-tuning. | 9 | β€” | 0.577 (in-domain, Bengali) |

**Which to use:** for inference, use **`config_A_reported_42cls.pt`** β€” it is the best model and the one the
dissertation reports. (It is a stripped, deployment-ready checkpoint, ~41 MB; the others retain optimizer
state and are ~242 MB.)

### The two class ontologies
- **9 classes** (intermediate stages): `text_body, headline, table, figure, caption, advertisement, sidebar, pull-quote, decorative-frame` β€” a coarse, script-agnostic set used for synthetic pretraining and self-training.
- **42 classes** (final): the full IndicDLP ontology, grouped into seven families β€” body text, heading hierarchy, list hierarchy, figures/tables/formulae, Q&A structure, navigation/metadata, and publishing furniture.

The detection head is **re-shaped** between stages (27 β†’ 9 β†’ 42 classes); the backbone and neck weights
transfer across all stages, and only the final classifier changes width.

---

## 4. Headline results (reported model, `config_A`)

- **0.525 mAP@0.5** and **0.364 mAP@[0.5:0.95]** on the **11,633-image** IndicDLP test set (180,039 instances).
- **Consistent across all 12 scripts** β€” per-script mAP@[.5:.95] spans only **0.304–0.426**, with no script failing catastrophically. (Best: Assamese 0.426; lowest: English 0.299.)
- **>100 FPS** on A100 β€” real-time preserved.
- **Precision 0.590 / Recall 0.506** overall.

### Scalability (accuracy vs. speed)
| Input size | Test mAP@[.5:.95] | Note |
|------------|-------------------|------|
| 640 px  | 0.329 | fastest; ~90% of peak accuracy at a fraction of the latency |
| 1024 px | **0.364** | peak accuracy (training resolution) |
| 1280 px | 0.353 | strictly dominated β€” slower *and* less accurate (train-test resolution mismatch) |

---

## 5. The key finding (a rigorous negative result)

**Neither proposed contribution improved over plainly fine-tuning the public checkpoint.**
Config A (no synthetic, no self-training) is the best model β€” by **+11.2 mAP** on the test set over Config B.

**This is statistically significant across all 12 scripts:**
- Wilcoxon signed-rank: **p < 0.001** (p = 0.00049)
- Paired t-test: **p < 10⁻¹⁰**
- Cohen's d = **7.37** (a very large effect)
- All 12/12 scripts improved under Config A; 95% CI on the gain [0.096, 0.114].

**Mechanism — catastrophic forgetting.** The validation→test gap *widens* from 2.6 to 11.2 mAP under
synthetic pretraining β€” the signature of damaged generalisation, not mere no-gain. The synthetic corpus
(23 templates, 9 classes) is narrower than the base model's original DocSynth-300K pretraining; 15 epochs to
0.985 synthetic-validation over-specialised the backbone, and the coarse 9-class intermediate ontology
erased fine distinctions the 42-class task then had to relearn from only 12,082 fine-labelled images.

> **In short:** the techniques failed, but the science succeeded β€” a falsified hypothesis, proven under
> controlled conditions and mechanistically explained, that saves others from the same mistake.

### Note on `config_B` and `config_C`
These are provided **for transparency and reproducibility of the ablation** β€” they are *not* recommended for
deployment, as they under-perform Config A. They let others verify the negative result independently.

---

## 6. Known limitations

- **Hierarchical-depth confusion is the dominant error.** The model localises headings/lists but cannot infer nesting *depth* β€” recall collapses with depth (section-title 0.66 β†’ sub 0.23 β†’ subsub 0.01). This is **architectural** (a single-shot detector classifies from local appearance; depth is relational) and **not Indic-specific** β€” it affects English identically.
- **Urdu is untested.** Urdu was synthesised and trained on, but the IndicDLP benchmark contains no Urdu, so right-to-left performance is unevaluated.
- **An auxiliary script-classification head** was designed but not implemented (future work).
- Results use a **12,082-image training subset** (compute-bounded); the reported mAP is therefore a lower bound, not a ceiling.

---

## 7. Usage

```python
from huggingface_hub import hf_hub_download
from doclayout_yolo import YOLOv10

# download the reported model from this repo
ckpt = hf_hub_download("VigneshPR/doclayout-yolo-indic", "config_A_reported_42cls.pt")
model = YOLOv10(ckpt)

# run on a document page
result = model.predict("page.jpg", imgsz=1024, conf=0.25)[0]
result.plot()                      # annotated image (numpy array, BGR)
print(len(result.boxes), "regions")
for c in result.boxes.cls.tolist():
    print(result.names[int(c)])
```

Install the runtime:
```bash
pip install git+https://github.com/opendatalab/DocLayout-YOLO.git
```

---

## 8. Training setup

- **Platform:** Google Colab Pro+ (no institutional cluster). Training on **A100 40 GB**; evaluation on **Tesla T4** to conserve compute.
- **Data engineering under a 15 GB Drive limit:** partial-tar extraction reduced the 42 GB IndicDLP download to ~5 GB; bulky image sets kept on ephemeral local disk, only labels/checkpoints persisted.
- **Fine-tuning:** ~30 epochs per config, batch 16 @ 1024, AMP, identical settings across all three ablation configs.
- **Self-training (CBST):** class-balanced pseudo-labelling on the unlabelled BaDLAD Bengali pool, with provably disjoint splits (leakage-audited: 20,365 βˆ’ 10,000 = 10,365).

---

## 9. Related resources

- **Synthetic corpus:** [VigneshPR/indicsynth-150k](https://huggingface.co/datasets/VigneshPR/indicsynth-150k) β€” the 150,000-page SA-MCB dataset used for pretraining.
- **Benchmark (working mirror):** [VigneshPR/IndicDLP](https://huggingface.co/datasets/VigneshPR/IndicDLP) β€” mirror of the IndicDLP benchmark (original authors: IndicDLP, ICDAR 2025); credit and licence belong to them.

---

## 10. Citation

If you use these models or the finding, please cite the dissertation:

```bibtex
@mastersthesis{vignesh2026doclayoutindic,
  title  = {DocLayout-YOLO-Indic: Cross-Script Document Layout Analysis for Indic Scripts
            using Synthetic Pretraining, Self-Training and Controlled Ablation},
  author = {Vignesh P},
  school = {BITS Pilani (WILP)},
  year   = {2026}
}
```

*Built on [DocLayout-YOLO](https://github.com/opendatalab/DocLayout-YOLO) (YOLOv10-m + GL-CRM).*