ONNX - Paddle
| Model | FP32 | Task |
|---|---|---|
| PP-DocLayout_plus-L | pp-docLayout_plus-l.onnx | layout detection |
| PP-OCRv6_medium_det | pp-ocrV6_medium_det.onnx | text line detection |
| PP-OCRv6_medium_rec | pp-ocrV6_medium_rec.onnx | text recognition |
| PP-LCNet_x1_0_table_cls | pp-lcNet_x1_0_table_cls.onnx | table classification (wired / wireless) |
| RT-DETR-L_wired_table_cell_det | rt-detr-l_wired_table_cell_det.onnx | table cell detection (wired) |
| RT-DETR-L_wireless_table_cell_det | rt-detr-l_wireless_table_cell_det.onnx | table cell detection (wireless) |
Usage
# onnxruntime-gpu for run it on GPU
pip install onnxruntime opencv-python numpy pyclipper
Every model folder has the same layout:
src/helper.py= onnx session builder (provider selection, threads, memory options)src/example.py= full pipeline: image preprocess, inference, postprocessonnx/= the model file
Run the commands from the repository root.
Test images used by the examples:
- jp_1.jpg = full page
- det_box_0294.jpg = single text line
- Table_1.jpg = single table
PP-DocLayout_plus-L
python3 PP-DocLayout_plus-L/src/example.py jp_1.jpg
0.936555 | header | [530, 62, 1213, 104]
0.910084 | table | [1066, 176, 1692, 260]
0.891657 | footer | [1066, 1157, 1198, 1175]
0.886924 | table | [54, 395, 1694, 799]
0.827543 | text | [55, 804, 1029, 959]
Note:
- Input = RGB image resized to 800x800, float32 0..1, CHW with batch dimension.
- Feed =
image,im_shape(800, 800),scale_factor(800 / height, 800 / width). - Output row =
[classId, score, x1, y1, x2, y2]with coordinates already in the original image space,
the box count is in the second output. - 20 label classes (header, doc_title, text, paragraph_title, image, table, chart, formula, ...):
full map in PP-DocLayout_plus-L/src/example.py. - Boxes overlap by design (no NMS in the model): filter by score (0.3 in the example) and handle
the containment on your side if needed.
PP-OCRv6_medium_det
python3 PP-OCRv6_medium_det/src/example.py jp_1.jpg
0.965487 | [[60, 1158], [850, 1158], [850, 1176], [60, 1176]]
0.827508 | [[1069, 1156], [1200, 1156], [1200, 1176], [1069, 1176]]
0.916008 | [[64, 1128], [837, 1128], [837, 1149], [64, 1149]]
Note:
- Input = BGR image (plain
cv2.imread, no channel swap) normalized with the ImageNet mean / std,
CHW with batch dimension. - Feed =
x, the only input, with a fully dynamic shape. - Resize = longest side capped to 960 (4000 hard limit), then both sides rounded to a multiple of 32.
- Output = a single probability map
[1, 1, height, width], the DB postprocess is on your side:
binarize at 0.2,cv2.findContours, minimum area box, score as the mean probability inside the box,
keep it over 0.45, expand it with pyclipper (unclip ratio 1.4) and rescale to the original image. - Output row =
[score, [4 corner points]]clockwise from the top-left corner, already in the original
image space. Boxes are quads, not axis aligned: a rotated line keeps its rotation. - Crop the quads and feed them to PP-OCRv6_medium_rec to get the text.
PP-OCRv6_medium_rec
python3 PP-OCRv6_medium_rec/src/example.py det_box_0294.jpg
0.953031 | (ε₯葨ε(δΈ)γ15γθ₯γγγ―ε₯葨ε(δΊ)γ10γεγ―ε₯葨ε(δΈ)γ16γθ₯γγγ―ε₯葨ε(δΊ)γ11γ)
Note:
- Input = a single cropped text line, not a full page: use the boxes from PP-OCRv6_medium_det.
- Input = BGR image (plain
cv2.imread, no channel swap) resized to height 48 keeping the aspect ratio,
scaled to 0..1 then normalized to -1..1, right padded with zeros, CHW with batch dimension. - Feed =
x, the only input, with a dynamic width (minimum 320, capped at 3200). - Output =
[1, timeStep, 18710]already softmaxed, decoded with a greedy CTC: drop the repeats first,
then the blank at index 0. - Character map = index 0 is the blank, 1..18708 are the lines of
PP-OCRv6_medium_rec/onnx/dictionary.txt, 18709 is the space. - Score = mean probability of the kept time steps.
PP-LCNet_x1_0_table_cls
python3 PP-LCNet_x1_0_table_cls/src/example.py Table_1.jpg
0.843585 | wired
Note:
- Input = a single table crop, not a full page: use the
tableboxes from PP-DocLayout_plus-L. - Input = RGB image resized so the short side is 256, center cropped to 224x224, scaled to 0..1 then
normalized with the ImageNet mean / std, CHW with batch dimension. - Feed =
x, the only input, fixed shape[1, 3, 224, 224]. - Output =
[1, 2]already softmaxed, index 0 iswired(ruled table), index 1 iswireless(no ruling lines). - Use the label to pick the cell detection model below.
RT-DETR-L_wired_table_cell_det
python3 RT-DETR-L_wired_table_cell_det/src/example.py Table_1.jpg
0.953868 | [467, 189, 1012, 216]
0.953357 | [467, 162, 1012, 189]
0.953275 | [467, 216, 1012, 243]
0.951546 | [467, 243, 1012, 270]
0.950149 | [467, 135, 1012, 162]
Note:
- Input = a single table crop classified as
wiredby PP-LCNet_x1_0_table_cls. - Input = RGB image resized to 640x640, float32 0..1, CHW with batch dimension.
- Feed =
image,im_shape(640, 640),scale_factor(640 / height, 640 / width). - Output row =
[classId, score, x1, y1, x2, y2]with coordinates already in the original image space,
the box count is in the second output. There is a single class (cell). - Output = 300 queries per image, no NMS in the model: filter by score (0.3 in the example), then the example
applies a NMS (IoU 0.5) and drops any box that contains 2 or more smaller boxes (containment 0.9). - Cells are returned as axis aligned boxes, merged cells come out as a single wider / taller box:
snap the edges to a grid on your side to get row / column index and span.
RT-DETR-L_wireless_table_cell_det
python3 RT-DETR-L_wireless_table_cell_det/src/example.py Table_1.jpg
0.951309 | [464, 162, 1009, 189]
0.950602 | [464, 189, 1010, 216]
0.950491 | [464, 216, 1010, 243]
0.950054 | [464, 243, 1010, 270]
0.949992 | [464, 297, 1010, 324]
Note:
- Same input, feed, output and postprocess as RT-DETR-L_wired_table_cell_det.
- Use it on a table crop classified as
wirelessby PP-LCNet_x1_0_table_cls.