MonkeyOCRv2 Detection

This repository provides the text detection experiments from the MonkeyOCRv2 paper. The visual encoder from MonkeyOCRv2-AS (ViTAEv2-S, 21M parameters) is integrated into DBNet and PSENet scene text detectors via MMOCR. The four-stage ViTAEv2 features (strides 4/8/16/32) are exposed as res2res5 and fed to the standard FPNC / FPNF necks, so no change to the detection heads is required.

Training and evaluation follow the official MMOCR protocols on Total-Text, CTW1500, and ICDAR2015.

Models and Results

For each detector, three visual backbones are compared under identical settings: the original ImageNet-pretrained encoder, the text-specific oCLIP encoder, and MonkeyOCRv2. MonkeyOCRv2 consistently improves F-score across all datasets and detector architectures.

Total-Text

Method P R F
DBNet (ResNet-50) 82.6 78.4 80.4
DBNet + oCLIP 85.1 81.7 83.4
DBNet + MonkeyOCRv2 87.7 80.1 83.7

CTW1500

Method P R F
PSENet (ResNet-50) 80.1 82.7 81.4
PSENet + oCLIP 82.1 85.5 83.8
PSENet + MonkeyOCRv2 88.3 82.0 85.1

ICDAR2015

Method P R F
PSENet (ResNet-50) 84.0 76.2 79.9
PSENet + oCLIP 87.3 82.6 84.9
PSENet + MonkeyOCRv2 90.4 80.3 85.0
DBNet (ResNet-50) 88.8 81.5 85.0
DBNet + oCLIP 90.9 84.1 87.4
DBNet + MonkeyOCRv2 91.2 86.0 88.5

Checkpoints

Download the checkpoints from HB16888/MonkeyOCRv2_det (HuggingFace) or WangXinhan/MonkeyOCRv2_det (ModelScope):

# HuggingFace
hf download HB16888/MonkeyOCRv2_det --include "*.pth" --local-dir ./model_weight
# ModelScope
modelscope download --model WangXinhan/MonkeyOCRv2_det --local_dir ./model_weight

Environment

The reproduced environment uses Python 3.11, PyTorch 2.9.0, CUDA 12.8, MMEngine 0.10.7, MMCV 2.0.1, MMOCR 1.0.1, and Transformers 4.57. Baselines and oCLIP models were trained on 2 GPUs; the MonkeyOCRv2 PSENet models on 4 GPUs (NVIDIA GeForce RTX 3090).

Installation

This directory is an add-on on top of the official MMOCR v1.0.1. Run:

bash install.sh            # clones MMOCR v1.0.1 into ./mmocr and patches it

Pretrained Backbone

Download the MonkeyOCRv2-AS visual encoder before training or evaluation:

hf download zenosai/MonkeyOCRv2-AS --local-dir ./pretrained/monkeyocrv2_as

Datasets

cd mmocr
bash tools/dataset_converters/prepare_all_datasets.sh

This prepares data/icdar2015, data/ctw1500, and data/totaltext in the MMOCR format. Note that some official CTW1500 download links are currently unreliable; we also provide the prepared CTW1500 in MMOCR format at HB16888/CTW1500 (HuggingFace) and WangXinhan/CTW1500 (ModelScope):

# HuggingFace
hf download HB16888/CTW1500 --repo-type dataset --local-dir ./ctw1500_dl
unzip ctw1500_dl/ctw1500_mmocr.zip -d data/
# ModelScope
modelscope download --dataset WangXinhan/CTW1500 --local_dir ./ctw1500_dl
unzip ctw1500_dl/ctw1500_mmocr.zip -d data/

Directory layout:

data/ctw1500/
├── textdet_imgs/
│   ├── train/    # 1000 images
│   └── test/     # 500 images
├── textdet_train.json
└── textdet_test.json

Training

All commands are run from the MMOCR root directory. Baselines and oCLIP models use 2 GPUs (batch size 16 in total); MonkeyOCRv2 PSENet models use 4 GPUs.

# ---------- DBNet on Total-Text ----------
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_resnet50_1200e_totaltext_2gpu.py 2
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_resnet50-oclip_1200e_totaltext_2gpu.py 2
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_mkv2vitae_1200e_totaltext_2gpu_adamw.py 2

# ---------- PSENet on CTW1500 ----------
bash tools/dist_train.sh configs/textdet/psenet/psenet_resnet50_fpnf_600e_ctw1500_2gpu.py 2
bash tools/dist_train.sh configs/textdet/psenet/psenet_resnet50-oclip_fpnf_600e_ctw1500_2gpu.py 2
bash tools/dist_train.sh configs/textdet/psenet/psenet_mkv2vitae_fpnf_600e_ctw1500_4gpu_adamw.py 4

# ---------- PSENet on ICDAR2015 ----------
bash tools/dist_train.sh configs/textdet/psenet/psenet_resnet50_fpnf_600e_icdar2015_2gpu.py 2
bash tools/dist_train.sh configs/textdet/psenet/psenet_resnet50-oclip_fpnf_600e_icdar2015_2gpu.py 2
bash tools/dist_train.sh configs/textdet/psenet/psenet_mkv2vitae_fpnf_600e_icdar2015_4gpu_adamw.py 4

# ---------- DBNet on ICDAR2015 ----------
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_resnet50_1200e_icdar2015_2gpu.py 2
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_resnet50-oclip_1200e_icdar2015_2gpu.py 2
bash tools/dist_train.sh configs/textdet/dbnet/dbnet_mkv2vitae_1200e_icdar2015_2gpu_adamw.py 2

Evaluation

# single GPU
python tools/test.py \
  configs/textdet/dbnet/dbnet_mkv2vitae_1200e_totaltext_2gpu_adamw.py \
  model_weight/dbnet_mkv2vitae_totaltext.pth

# multi GPU
bash tools/dist_test.sh \
  configs/textdet/psenet/psenet_mkv2vitae_fpnf_600e_ctw1500_4gpu_adamw.py \
  model_weight/psenet_mkv2vitae_ctw1500.pth 4

The evaluation prints precision / recall / hmean with the HmeanIOUMetric, matching the tables above.

Acknowledgements

This project builds on MMOCR, DBNet, PSENet, oCLIP, and MonkeyOCRv2.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for HB16888/MonkeyOCRv2_det