abhishek-gola's picture
east text detection (#13)
6aa92fc
|
Raw
History Blame Contribute Delete
2.87 kB

EAST Text Detection

Scene-text detection with the EAST (Efficient and Accurate Scene Text) detector. The model was originally distributed as a frozen TensorFlow graph (frozen_east_text_detection.pb) and converted to ONNX for use with OpenCV's DNN module.

Model Details

  • Architecture: EAST with a ResNet-50 backbone and a feature-fusion head
  • Input: RGB image, 320×320, raw 0–255 float, mean (123.68, 116.78, 103.94), swapRB, NCHW layout (input_images:0, shape [1, 3, 320, 320])
  • Outputs:
    • feature_fusion/Conv_7/Sigmoid:0 — score map, shape [1, 1, 80, 80]
    • feature_fusion/concat_3:0 — RBOX geometry, shape [1, 5, 80, 80]
  • Post-processing: OpenCV's TextDetectionModel_EAST decodes the score/geometry maps into rotated boxes (confidence threshold + rotated-NMS)
  • Framework: ONNX (converted from the TensorFlow frozen graph via tf2onnx, opset 15)
  • Original weights: https://github.com/argman/EAST

Both input and outputs are emitted in NCHW so OpenCV consumes them directly.

Usage

Python

python demo.py --model east_text_detection_2026jul.onnx --image example_outputs/input_image.png --output example_outputs/output_image.png

Or import directly:

import cv2

model = cv2.dnn.TextDetectionModel_EAST("east_text_detection_2026jul.onnx")
# see demo.py for the full inference pipeline

C++

The C++ demo runs inference with OpenCV's DNN module (default engine — no ONNX Runtime needed). Adjust the OpenCV paths to your setup:

OCV=/path/to/opencv                          # OpenCV source tree
OCVBUILD=/path/to/opencv/build               # OpenCV build directory (generated headers + libs)
g++ -std=c++17 demo.cpp -o demo \
  -I$OCV/include \
  -I$OCV/modules/core/include \
  -I$OCV/modules/dnn/include \
  -I$OCV/modules/imgproc/include \
  -I$OCV/modules/imgcodecs/include \
  -I$OCVBUILD \
  -L$OCVBUILD/lib -Wl,-rpath,$OCVBUILD/lib -lopencv_dnn -lopencv_imgcodecs -lopencv_imgproc -lopencv_core
./demo --model east_text_detection_2026jul.onnx --image example_outputs/input_image.png

Conversion

The ONNX model was exported from the frozen TensorFlow graph with tf2onnx (opset 15) via convert_to_onnx.py — input input_images:0, outputs feature_fusion/Conv_7/Sigmoid:0 and feature_fusion/concat_3:0. Both the input and the outputs are forced to NCHW (inputs_as_nchw / outputs_as_nchw) so the tensors match OpenCV's layout; without outputs_as_nchw the score/geometry maps come out as NHWC and the EAST decoder rejects them. Requires tensorflow, tf2onnx, and onnx.

python convert_to_onnx.py --pb ../pb/frozen_east_text_detection.pb

License

See LICENSE — the model originates from argman/EAST, released under the GNU General Public License v3.0.