torchocr weights
Pretrained checkpoints for torchocr, a PyTorch-native,
torchvision-style OCR library. You normally never download these files by hand: each one backs a
member of a *_Weights enum, and torch.hub fetches it on first use and checks the SHA-256 prefix
embedded in the file name.
from torchocr import DBPostProcessor
from torchocr.models import DBNet_MobileNetV3_Large_05_Weights, dbnet_mobilenet_v3_large_05
weights = DBNet_MobileNetV3_Large_05_Weights.ICDAR2015
model = dbnet_mobilenet_v3_large_05(weights=weights).eval()
preprocess = weights.transforms()
postprocess = DBPostProcessor(**weights.meta["postprocess"])
Checkpoints
| Weights enum | File | Params | ICDAR-2015 test |
|---|---|---|---|
DBNet_MobileNetV3_Large_05_Weights.ICDAR2015 |
dbnet_mobilenet_v3_large_05_ic15-d9d17ae7.pth |
0.60 M | P 0.785 · R 0.689 · hmean 0.734 |
DBNet_MobileNetV3_Large_05_Weights.PPOCR_V3_EN |
dbnet_mobilenet_v3_large_05_ppocr_v3_en-9bfd3b59.pth |
0.60 M | hmean 0.441 ¹ |
DBNet_MobileNetV3_Large_05_Weights.PPOCR_V3_CH |
dbnet_mobilenet_v3_large_05_ppocr_v3_ch-fc000d1e.pth |
0.60 M | hmean 0.422 ¹ |
DBNet_ResNet18_VD_Weights.PPOCR_SERVER_V2 |
dbnet_resnet18_vd_ppocr_server_v2-59d99b11.pth |
12.4 M | hmean 0.414 ¹ |
CRNN_ResNet34_VD_Weights.PPOCR_SERVER_V2 |
crnn_resnet34_vd_ppocr_server_v2-803ba4c9.pth |
27.9 M | word acc. 0.665 ² |
¹ Line-level detectors converted from PaddleOCR. ICDAR-2015 scores word boxes, so adjacent words
merged into one line count as misses: these numbers compare checkpoints, they are not the accuracy
you will see on documents.
² 2077 alphanumeric care words of ICDAR-2015 test cropped from the full images with
torchocr.ops.crop_quads on their ground-truth quads, case-insensitive (not the official Task 4.3
crops). Axis-aligned crops of the same words score 0.579.
Evaluation protocol
Detection: official ICDAR-2015 IoU protocol (one-to-one greedy matching at IoU > 0.5, ###
don't-care filtering) on rotated quads, implemented by torchocr.metrics.DetectionHmean and verified
identical to PaddleOCR's DetectionIoUEvaluator. Pre/post-processing is exactly
weights.transforms() and weights.meta["postprocess"]. Every number is reproducible with:
python references/detection/evaluate.py --root <icdar2015> --weights <enum>
python references/recognition/evaluate.py --root <icdar2015> --weights CRNN_ResNet34_VD_Weights.PPOCR_SERVER_V2
The ICDAR2015 detector
PPOCR_V3_EN fine-tuned for 300 epochs on the 1000 ICDAR-2015 training images
(references/detection/train.py, PaddleOCR's augmentation recipe, 640 px crops, AdamW 1e-3 with
warmup + cosine, ~80 min on one laptop GPU). The reported weights are the last epoch: no
checkpoint was selected on the test set. box_thresh = 0.45 maximizes hmean on the training
split (references/detection/calibrate.py) and was applied once to the test split; with
PaddleOCR's default box_thresh = 0.6 the same checkpoint scores 0.688.
Intended inputs: incidental scene text in street-level photos, Latin script, evaluated at 736 × 1280. Expect lower recall on dense documents and on non-Latin scripts.
Provenance and licenses
PPOCR_*checkpoints are converted from PaddleOCR releases (Apache-2.0) using the build-time converters in torchocr'sscripts/, whose recipes are adapted from PaddleOCR2Pytorch (Apache-2.0). The PP-OCRv3 files hold theStudentnetwork that PaddleOCR exports for inference.- The
ICDAR2015checkpoint derives fromPPOCR_V3_EN(Apache-2.0) and was fine-tuned on the ICDAR 2015 Incidental Scene Text training set, which its organizers release for research use. Check that your use complies with those terms.