vagheshpatel's picture
Sync ocr-text-recognition from metro-analytics-catalog
67ba03b verified
|
Raw
History Blame Contribute Delete
15.8 kB
---
license: mit
license_link: LICENSE
library_name: openvino
pipeline_tag: image-to-text
tags:
- openvino
- intel
- paddleocr
- ocr
- text-recognition
- edge-ai
- metro
- dlstreamer
language:
- en
---
# OCR for Text
| Property | Value |
|---|---|
| **Category** | Optical Character Recognition (Text Detection + Recognition) |
| **Base Model** | [PP-OCRv4](https://github.com/PaddlePaddle/PaddleOCR) (PaddlePaddle) |
| **Source Framework** | PaddlePaddle |
| **Supported Precisions** | FP32, FP16 |
| **Inference Engine** | OpenVINO |
| **Hardware** | CPU, GPU, NPU |
| **Detected Class(es)** | Text regions + recognized text strings |
---
## Overview
OCR for Text is a Metro Analytics use case that detects and reads text in
images and video streams using the PaddleOCR PP-OCRv4 pipeline.
It composes two models:
- **PP-OCRv4 Detection** (`ch_PP-OCRv4_det`) -- a lightweight DBNet-based
text detector that locates text regions in the frame.
- **PP-OCRv4 Recognition** (`ch_PP-OCRv4_rec_server`) -- the larger "server"
CRNN-CTC recognizer variant, which is more accurate than the lightweight
mobile variant on stylized or decorative fonts, and converts each cropped
text region into a character string.
Both models are converted to OpenVINO IR using the `ovc` (OpenVINO Model
Converter) tool which reads PaddlePaddle models directly.
This is the best supported end-to-end OCR stack for OpenVINO.
Typical Metro deployments include:
- **Signage Reading** -- read platform signs, departure boards, safety notices.
- **Document Scanning** -- extract text from forms, labels, and ID cards.
- **Label Verification** -- read package labels or barcodes in logistics.
- **Multilingual Support** -- PP-OCRv4 supports multiple scripts out of the box.
For license-plate-specific OCR, see the
[license-plate-recognition](../license-plate-recognition/) use case which
includes a specialized plate detector.
---
## Prerequisites
- Python 3.11+
- [Install OpenVINO](https://docs.openvino.ai/2026/get-started/install-openvino.html) (latest version)
- [Install Intel DLStreamer](https://docs.openedgeplatform.intel.com/2026.0/edge-ai-libraries/dlstreamer/get_started/install/install_guide_ubuntu.html) (latest version)
Create and activate a Python virtual environment before running the scripts:
```bash
python3 -m venv .venv --system-site-packages
source .venv/bin/activate
```
> **Note:** The `--system-site-packages` flag is required so the virtual
> environment can access the system-installed OpenVINO and DLStreamer Python
> packages.
---
## Getting Started
### Download and Convert Models
Run the provided script to download the PaddleOCR models and convert them to
OpenVINO IR:
```bash
chmod +x export_and_quantize.sh
./export_and_quantize.sh
```
The script performs the following steps:
1. Installs dependencies (`openvino`).
2. Downloads the PP-OCRv4 detection and recognition inference models.
3. Converts both to OpenVINO IR format using `ovc`.
4. Downloads a sample test image with text, a sample test video
(`test_video.mp4`, a close-up of street name and stop signs), and the
PP-OCRv4 character dictionary (`ppocr_keys_v1.txt`) used to CTC-decode the
recognizer's output into text.
Output files:
- `ch_PP-OCRv4_det_infer/` -- detection model (OpenVINO IR).
- `ch_PP-OCRv4_rec_server_infer/` -- recognition model, server variant (OpenVINO IR).
- `ppocr_keys_v1.txt` -- character dictionary for the recognizer's CTC decoder.
### OpenVINO Sample
The sample below runs the full PP-OCRv4 pipeline across every frame of a
video: the detector locates text regions (using an aspect-ratio-preserving
resize and a dilation step so a whole word is captured in one box instead of
fragments), then the recognizer reads each cropped region and CTC-decodes it
into a text string, which is drawn as a solid-background label directly over
its box so the highlighted region visibly shows what is written.
Change the `device` string to run on CPU, GPU, or NPU.
```python
import cv2
import numpy as np
import openvino as ov
DET_MODEL = "ch_PP-OCRv4_det_infer/inference.xml"
REC_MODEL = "ch_PP-OCRv4_rec_server_infer/inference.xml"
DICT_FILE = "ppocr_keys_v1.txt"
INPUT_VIDEO = "test_video.mp4"
DET_SIZE = 960
core = ov.Core()
# Change device to "GPU" or "NPU" to run on integrated GPU or NPU.
det_compiled = core.compile_model(core.read_model(DET_MODEL), "CPU")
rec_compiled = core.compile_model(core.read_model(REC_MODEL), "CPU")
# CTC label map: index 0 is the blank symbol, followed by every character in
# the dictionary file, followed by a trailing space character.
chars = open(DICT_FILE, encoding="utf-8").read().splitlines()
dict_character = ["blank"] + chars + [" "]
def detect_text_regions(frame, thresh=0.3, pad=4):
"""Return (x, y, w, h) boxes for words/lines of text in a frame.
Resizing preserves aspect ratio (letterboxed onto a square canvas) so
text isn't skewed, and dilating the detection map merges nearby
characters into one box per word instead of one per character.
"""
h0, w0 = frame.shape[:2]
scale = DET_SIZE / max(h0, w0)
resized = cv2.resize(frame, (int(w0 * scale), int(h0 * scale)))
canvas = np.zeros((DET_SIZE, DET_SIZE, 3), dtype=np.uint8)
canvas[:resized.shape[0], :resized.shape[1]] = resized
blob = canvas.astype(np.float32).transpose(2, 0, 1)[np.newaxis] / 255.0
det_map = det_compiled([blob])[det_compiled.output(0)][0, 0]
binary = (det_map > thresh).astype(np.uint8) * 255
dilated = cv2.dilate(binary, np.ones((9, 25), np.uint8))
contours, _ = cv2.findContours(dilated, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
boxes = []
for c in contours:
x, y, w, h = cv2.boundingRect(c)
if w < 10 or h < 5:
continue
x0, y0 = max(0, x / scale - pad), max(0, y / scale - pad)
x1, y1 = min(w0, (x + w) / scale + pad), min(h0, (y + h) / scale + pad)
boxes.append((int(x0), int(y0), int(x1 - x0), int(y1 - y0)))
return boxes
def recognize_text(crop, rec_h=48, max_w=320):
"""Resize a cropped text region to the recognizer's input shape and
CTC-decode the predicted character sequence into a string."""
h, w = crop.shape[:2]
if h == 0 or w == 0:
return "", 0.0
resized_w = max(1, min(max_w, round(rec_h * w / h)))
blob = cv2.resize(crop, (resized_w, rec_h)).astype(np.float32) / 255.0
blob = ((blob - 0.5) / 0.5).transpose(2, 0, 1)[np.newaxis, ...]
preds = rec_compiled([blob])[rec_compiled.output(0)][0]
idx = np.argmax(preds, axis=1)
conf = np.max(preds, axis=1)
text, scores, prev = [], [], -1
for i, c in zip(idx, conf):
if i != 0 and i != prev:
text.append(dict_character[i])
scores.append(c)
prev = i
confidence = float(np.mean(scores)) if scores else 0.0
return "".join(text), confidence
def annotate(frame, box, text, confidence):
"""Draw a bounding box and, if any text was recognized, a legible
label (solid background so it stays readable over any color) above it."""
x, y, w, h = box
cv2.rectangle(frame, (x, y), (x + w, y + h), (0, 255, 0), 2)
if not text:
return
label = f"{text} ({confidence:.2f})"
(tw, th), base = cv2.getTextSize(label, cv2.FONT_HERSHEY_SIMPLEX, 0.7, 2)
top = max(0, y - th - base - 6)
cv2.rectangle(frame, (x, top), (x + tw + 6, top + th + base + 6), (0, 255, 0), -1)
cv2.putText(frame, label, (x + 3, top + th + 2),
cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 0, 0), 2)
cap = cv2.VideoCapture(INPUT_VIDEO)
fps = cap.get(cv2.CAP_PROP_FPS) or 30.0
width = int(cap.get(cv2.CAP_PROP_FRAME_WIDTH))
height = int(cap.get(cv2.CAP_PROP_FRAME_HEIGHT))
writer = cv2.VideoWriter(
"output_openvino.mp4", cv2.VideoWriter_fourcc(*"mp4v"), fps, (width, height))
frame_idx = 0
total_regions = 0
while True:
ok, frame = cap.read()
if not ok:
break
frame_idx += 1
for box in detect_text_regions(frame):
x, y, w, h = box
text, confidence = recognize_text(frame[y:y + h, x:x + w])
annotate(frame, box, text, confidence)
total_regions += 1
print(f"Frame {frame_idx}: region=({x},{y},{w},{h}) text={text!r} "
f"confidence={confidence:.2f}", flush=True)
writer.write(frame)
cap.release()
writer.release()
print(f"Total text regions across all frames: {total_regions}", flush=True)
print("Saved: output_openvino.mp4")
```
**Device targets:**
- `"CPU"` -- default, works on all Intel platforms.
- `"GPU"` -- Intel integrated or discrete GPU.
- `"NPU"` -- Intel NPU; PP-OCRv4 FP16 models are NPU-compatible.
> **Note:** Recognition accuracy depends heavily on font, angle, and image
> quality. Plain block-lettered signage (as in the sample video) decodes
> reliably; stylized or decorative fonts are harder for a general-purpose
> OCR model and may not decode perfectly.
#### Expected Output
![OpenVINO expected output](expected_output_openvino.gif)
### DLStreamer Sample
The sample below decodes a video with the DLStreamer/GStreamer stack
(`decodebin3 ! videoconvert`), pulls BGR frames through `appsink`,
runs the PP-OCRv4 text detector on each frame (using an aspect-ratio-preserving
resize and a dilation step so a whole word is captured in one box instead of
fragments), then runs the PP-OCRv4 recognizer on each cropped region and
CTC-decodes the result into text drawn as a solid-background label directly
over its box before writing the annotated output to `output_dlstreamer.mp4`.
> **Notes on running this sample:**
>
> - Export `PYTHONPATH` so the DLStreamer Python module is importable:
>
> ```bash
> source /opt/intel/openvino_2026/setupvars.sh
> source /opt/intel/dlstreamer/scripts/setup_dls_env.sh
> export PYTHONPATH=/opt/intel/dlstreamer/python:\
> /opt/intel/dlstreamer/gstreamer/lib/python3/dist-packages:${PYTHONPATH:-}
> ```
```python
import gi
gi.require_version("Gst", "1.0")
from gi.repository import Gst
import numpy as np
import openvino as ov
Gst.init([])
# Import cv2 after Gst.init to avoid GStreamer re-initialization conflicts.
import cv2
INPUT_VIDEO = "test_video.mp4"
DET_MODEL = "ch_PP-OCRv4_det_infer/inference.xml"
REC_MODEL = "ch_PP-OCRv4_rec_server_infer/inference.xml"
DICT_FILE = "ppocr_keys_v1.txt"
DET_SIZE = 960
core = ov.Core()
det_compiled = core.compile_model(core.read_model(DET_MODEL), "CPU")
rec_compiled = core.compile_model(core.read_model(REC_MODEL), "CPU")
# CTC label map: index 0 is the blank symbol, followed by every character in
# the dictionary file, followed by a trailing space character.
chars = open(DICT_FILE, encoding="utf-8").read().splitlines()
dict_character = ["blank"] + chars + [" "]
def detect_text_regions(frame, thresh=0.3, pad=4):
"""Return (x, y, w, h) boxes for words/lines of text in a frame.
Resizing preserves aspect ratio (letterboxed onto a square canvas) so
text isn't skewed, and dilating the detection map merges nearby
characters into one box per word instead of one per character.
"""
h0, w0 = frame.shape[:2]
scale = DET_SIZE / max(h0, w0)
resized = cv2.resize(frame, (int(w0 * scale), int(h0 * scale)))
canvas = np.zeros((DET_SIZE, DET_SIZE, 3), dtype=np.uint8)
canvas[:resized.shape[0], :resized.shape[1]] = resized
blob = canvas.astype(np.float32).transpose(2, 0, 1)[np.newaxis] / 255.0
det_map = det_compiled([blob])[det_compiled.output(0)][0, 0]
binary = (det_map > thresh).astype(np.uint8) * 255
dilated = cv2.dilate(binary, np.ones((9, 25), np.uint8))
contours, _ = cv2.findContours(dilated, cv2.RETR_EXTERNAL, cv2.CHAIN_APPROX_SIMPLE)
boxes = []
for c in contours:
x, y, w, h = cv2.boundingRect(c)
if w < 10 or h < 5:
continue
x0, y0 = max(0, x / scale - pad), max(0, y / scale - pad)
x1, y1 = min(w0, (x + w) / scale + pad), min(h0, (y + h) / scale + pad)
boxes.append((int(x0), int(y0), int(x1 - x0), int(y1 - y0)))
return boxes
def recognize_text(crop, rec_h=48, max_w=320):
"""Resize a cropped text region to the recognizer's input shape and
CTC-decode the predicted character sequence into a string."""
h, w = crop.shape[:2]
if h == 0 or w == 0:
return "", 0.0
resized_w = max(1, min(max_w, round(rec_h * w / h)))
blob = cv2.resize(crop, (resized_w, rec_h)).astype(np.float32) / 255.0
blob = ((blob - 0.5) / 0.5).transpose(2, 0, 1)[np.newaxis, ...]
preds = rec_compiled([blob])[rec_compiled.output(0)][0]
idx = np.argmax(preds, axis=1)
conf = np.max(preds, axis=1)
text, scores, prev = [], [], -1
for i, c in zip(idx, conf):
if i != 0 and i != prev:
text.append(dict_character[i])
scores.append(c)
prev = i
confidence = float(np.mean(scores)) if scores else 0.0
return "".join(text), confidence
def annotate(frame, box, text, confidence):
"""Draw a bounding box and, if any text was recognized, a legible
label (solid background so it stays readable over any color) above it."""
x, y, w, h = box
cv2.rectangle(frame, (x, y), (x + w, y + h), (0, 255, 0), 2)
if not text:
return
label = f"{text} ({confidence:.2f})"
(tw, th), base = cv2.getTextSize(label, cv2.FONT_HERSHEY_SIMPLEX, 0.7, 2)
top = max(0, y - th - base - 6)
cv2.rectangle(frame, (x, top), (x + tw + 6, top + th + base + 6), (0, 255, 0), -1)
cv2.putText(frame, label, (x + 3, top + th + 2),
cv2.FONT_HERSHEY_SIMPLEX, 0.7, (0, 0, 0), 2)
pipeline_str = (
f"filesrc location={INPUT_VIDEO} ! decodebin3 ! videoconvert ! "
"video/x-raw,format=BGR ! "
"appsink name=sink emit-signals=false sync=false"
)
pipeline = Gst.parse_launch(pipeline_str)
sink = pipeline.get_by_name("sink")
pipeline.set_state(Gst.State.PLAYING)
writer = None
frame_idx = 0
total_regions = 0
while True:
sample = sink.emit("pull-sample")
if sample is None:
break
buf = sample.get_buffer()
caps = sample.get_caps().get_structure(0)
width = caps.get_value("width")
height = caps.get_value("height")
ok, mapinfo = buf.map(Gst.MapFlags.READ)
if not ok:
continue
frame = np.ndarray((height, width, 3), dtype=np.uint8,
buffer=mapinfo.data).copy()
buf.unmap(mapinfo)
frame_idx += 1
for box in detect_text_regions(frame):
x, y, w, h = box
text, confidence = recognize_text(frame[y:y + h, x:x + w])
annotate(frame, box, text, confidence)
total_regions += 1
print(f"Frame {frame_idx}: region=({x},{y},{w},{h}) text={text!r} "
f"confidence={confidence:.2f}", flush=True)
if writer is None:
writer = cv2.VideoWriter(
"output_dlstreamer.mp4", cv2.VideoWriter_fourcc(*"mp4v"),
30.0, (width, height))
writer.write(frame)
pipeline.set_state(Gst.State.NULL)
if writer:
writer.release()
print(f"Total text regions across all frames: {total_regions}", flush=True)
```
**Device targets:**
- `"CPU"` -- default for OpenVINO inference inside the appsink loop.
- `"GPU"` -- change `"CPU"` to `"GPU"` in `core.compile_model()`.
- `"NPU"` -- change `"CPU"` to `"NPU"` in `core.compile_model()`.
#### Expected Output
![DLStreamer expected output](expected_output_dlstreamer.gif)
---
## License
Licensed under the MIT License. See [LICENSE](LICENSE) for details.
## References
- [PaddleOCR PP-OCRv4](https://github.com/PaddlePaddle/PaddleOCR)
- [PaddleOCR OpenVINO Deployment](https://github.com/PaddlePaddle/PaddleOCR/blob/main/deploy/paddle2onnx/readme.md)
- [OpenVINO Documentation](https://docs.openvino.ai/)
- [Intel DLStreamer](https://docs.openedgeplatform.intel.com/2026.0/edge-ai-libraries/dlstreamer/index.html)