ONNX - Paddle

Model FP32 Task
PP-DocLayout_plus-L pp-docLayout_plus-l.onnx layout detection
PP-OCRv6_medium_det pp-ocrV6_medium_det.onnx text line detection
PP-OCRv6_medium_rec pp-ocrV6_medium_rec.onnx text recognition

Usage

# onnxruntime-gpu for run it on GPU
pip install onnxruntime opencv-python numpy pyclipper

Every model folder has the same layout:

  • src/helper.py = onnx session builder (provider selection, threads, memory options)
  • src/example.py = full pipeline: image preprocess, inference, postprocess
  • onnx/ = the model file

Run the commands from the repository root.

Test images used by the examples:

PP-DocLayout_plus-L

python3 PP-DocLayout_plus-L/src/example.py jp_1.jpg
0.936555 | header | [530, 62, 1213, 104]
0.910084 | table | [1066, 176, 1692, 260]
0.891657 | footer | [1066, 1157, 1198, 1175]
0.886924 | table | [54, 395, 1694, 799]
0.827543 | text | [55, 804, 1029, 959]

Note:

  • Input = RGB image resized to 800x800, float32 0..1, CHW with batch dimension.
  • Feed = image, im_shape (800, 800), scale_factor (800 / height, 800 / width).
  • Output row = [classId, score, x1, y1, x2, y2] with coordinates already in the original image space,
    the box count is in the second output.
  • 20 label classes (header, doc_title, text, paragraph_title, image, table, chart, formula, ...):
    full map in PP-DocLayout_plus-L/src/example.py.
  • Boxes overlap by design (no NMS in the model): filter by score (0.3 in the example) and handle
    the containment on your side if needed.

PP-OCRv6_medium_det

python3 PP-OCRv6_medium_det/src/example.py jp_1.jpg
0.965487 | [[60, 1158], [850, 1158], [850, 1176], [60, 1176]]
0.827508 | [[1069, 1156], [1200, 1156], [1200, 1176], [1069, 1176]]
0.916008 | [[64, 1128], [837, 1128], [837, 1149], [64, 1149]]

Note:

  • Input = BGR image (plain cv2.imread, no channel swap) normalized with the ImageNet mean / std,
    CHW with batch dimension.
  • Feed = x, the only input, with a fully dynamic shape.
  • Resize = longest side capped to 960 (4000 hard limit), then both sides rounded to a multiple of 32.
  • Output = a single probability map [1, 1, height, width], the DB postprocess is on your side:
    binarize at 0.2, cv2.findContours, minimum area box, score as the mean probability inside the box,
    keep it over 0.45, expand it with pyclipper (unclip ratio 1.4) and rescale to the original image.
  • Output row = [score, [4 corner points]] clockwise from the top-left corner, already in the original
    image space. Boxes are quads, not axis aligned: a rotated line keeps its rotation.
  • Crop the quads and feed them to PP-OCRv6_medium_rec to get the text.

PP-OCRv6_medium_rec

python3 PP-OCRv6_medium_rec/src/example.py det_box_0294.jpg
0.953031 | (εˆ₯葨十(δΈ€)γ€Œ15」θ‹₯しくはεˆ₯葨十(二)γ€Œ10γ€εˆγ―εˆ₯葨十(δΈ€)γ€Œ16」θ‹₯しくはεˆ₯葨十(二)γ€Œ11」)

Note:

  • Input = a single cropped text line, not a full page: use the boxes from PP-OCRv6_medium_det.
  • Input = BGR image (plain cv2.imread, no channel swap) resized to height 48 keeping the aspect ratio,
    scaled to 0..1 then normalized to -1..1, right padded with zeros, CHW with batch dimension.
  • Feed = x, the only input, with a dynamic width (minimum 320, capped at 3200).
  • Output = [1, timeStep, 18710] already softmaxed, decoded with a greedy CTC: drop the repeats first,
    then the blank at index 0.
  • Character map = index 0 is the blank, 1..18708 are the lines of
    PP-OCRv6_medium_rec/onnx/dictionary.txt, 18709 is the space.
  • Score = mean probability of the kept time steps.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support