PaddleOCR-VL Layer-12 Feature Extractor (ONNX)

Truncated PaddleOCR-VL vision encoder — 1152D intermediate features for document image quality assessment.

Layer 12 of the PaddleOCR-VL vision encoder captures mid-level visual patterns (texture, edge structure, compression artifacts, noise patterns) that are highly predictive of OCR readability. This ONNX model extracts those features with zero PyTorch dependency — just onnxruntime + numpy + Pillow.

Why Layer 12?

From a sensitivity analysis across all 27 encoder layers:

Degradation Best Layer Spearman ρ Notes
Gaussian Noise layer_00 0.96 Shallow layers detect fine noise
JPEG Compression layer_08 0.86 Mid layers sense blocking artifacts
Gaussian Blur layer_08 0.66 Moderate response
Downscale pooled 0.65 Resolution loss detectable

Layer 12 was chosen because:

  • Highest intra/inter-class distance ratio (0.29) — features cluster by quality level, not degradation type
  • Cross-resolution consistency >0.99 — same image at 224² vs 728² produces nearly identical features
  • Good balance: deeper than noise-sensitive layers, shallower than semantic layers

Domain & Positioning: 这个模型做什么、不做什么

一句话定位

这是一个"OCR 场景的文档图片质量特征提取器",不是通用图片质量打分器,也不是 OCR 识别器。

它的核心假设是:如果一张图在 PaddleOCR-VL 的中间层特征空间里和清晰参考图很接近,那这张图大概率能被 OCR 正确识别。

使用场景

场景 是否适合 说明
📄 文档上传质量校验 ✅ 核心场景 "这张扫描件够清晰吗?需要让用户重拍吗?"
🏦 票据/证件 OCR 前置过滤 ✅ 推荐 在送入 OCR 引擎前筛掉模糊/噪声过大的图片
📸 拍照文档质量评分 ✅ 适合 模糊、噪声、JPEG 压缩 → 有强信号
🌐 通用图片美学评分 ❌ 不适合 模型不理解"构图"、"色彩搭配"、"主体突出"
🌃 自然场景图片质量 ⚠️ 有限 对噪声/模糊有响应,但对过曝/偏色无感知
🔤 OCR 文字识别 ❌ 做不了 这是特征提取器,不是 OCR 引擎。请用完整版 PaddleOCR-VL
🎯 文字语义理解 ❌ 做不了 深层语义信息在 layer 20+,layer 12 以纹理/结构信息为主

与同类模型的关键区别

1. vs 完整 PaddleOCR-VL(端到端 OCR 模型)

本模型 (Layer-12 ONNX) 完整 PaddleOCR-VL
任务 图片质量特征提取 端到端 OCR(检测+识别)
输出 1152D 特征向量 文字检测框 + 文本内容
推理速度 ~100ms CPU ~2-5s GPU
依赖 onnxruntime only PyTorch + transformers
部署 任意 ONNX 运行时(含移动端/边缘) 需要 GPU 服务器
理解内容 不懂文字内容,只看纹理质量 完整理解文档内容

选择指南:如果你需要的是"这张图能读吗"的答案 → 用本模型提取特征 + 轻量回归器。如果你需要"图里写了什么" → 直接用完整 PaddleOCR-VL。

2. vs DINOv3(通用视觉特征提取器)

本模型 (Layer-12) DINOv3 (ViT-S/16)
训练数据 PaddleOCR-VL 图文数据(文档为主) 大规模自然图片(ImageNet/LAION 级别)
特征偏好 纹理、边缘、压缩伪影(文档质量相关) 语义类别、物体形状、全局布局
文档噪声敏感度 ⭐⭐⭐⭐⭐ 极高(ρ=0.96 for noise) ⭐⭐ 中等(对噪声有响应但不如专用模型)
亮度/对比度感知 ⭐ 弱(盲区) ⭐⭐ 中等
特征维度 1152D(单层) 384D (ViT-S) ~ 1536D (ViT-L)

关键差异:DINOv3 是"通才",擅长判断"这是什么物体";本模型是"文档专才",擅长判断"这文档还能读吗"。在文档质量任务上,本模型的特征 + 轻量回归器(84K 参数)达到 ρ=0.74,优于直接对 DINOv3 特征做 XGBoost。

实测对比(合成文档质量数据集,12 种退化):

特征源 + 回归器 Spearman ρ
PaddleOCR Layer-12 + Two-Tower MLP 0.74 0.44
PaddleOCR Layer-12 + XGBoost 0.60 0.35
CV 传统特征 only + XGBoost 0.66 0.33
DINOv3 ViT-S + XGBoost ~0.50 ~0.25

3. vs Qwen3-VL-Embedding / CLIP Embedding(通用 VLM 特征)

本模型 (Layer-12) Qwen3-VL-Embedding / CLIP
特征来源 视觉编码器中间层(layer 12/27) 视觉+文本对齐后的池化输出
语义层次 中层纹理/结构(edges, textures, artifacts) 高层语义("这是一张包含表格的文档")
跨模态 纯视觉,无文本对齐 视觉-文本对齐(CLIP 训练目标)
质量退化感知 ⭐⭐⭐⭐⭐ 天然敏感 ⭐⭐ "模糊的文档"和"清晰的文档"在语义空间可能很近
相似度计算 余弦距离 ≈ 退化程度 余弦相似度 ≈ 语义相似度

核心差异:CLIP/Qwen3-VL 的 embedding 回答"这两张图内容像不像";本模型的特征回答"这张图和清晰版差多远"。对于文档质量任务,你不想要语义相似度——模糊的身份证和清晰的身份证在语义上是同一张证,但在质量维度上区别巨大。

4. vs Q-ReAlign(通用 VLM 质量评分)

Q-ReAlign 是直接用 VLM 做质量评分的方案(将质量分数映射到离散 token 做软分类)。

本模型 (Layer-12) Q-ReAlign
方法论 特征提取 + 下游回归器 端到端 VLM 质量打分
文档质量表现 ρ=0.74 ρ ≈ -0.60(对 blur/JPEG 评分反向)
原因 特征空间天然反映退化程度 训练数据是自然照片美学评分,不适用于文档
优势 文档场景专用,轻量部署 通用场景,零样本即可用

选择指南:Q-ReAlign 适合自然照片("这张风景照好看吗"),不适合文档("这张扫描件能 OCR 吗")。

5. vs 传统 CV 质量指标(Laplacian / FFT / Canny)

本模型 (Layer-12) 传统 CV 指标
噪声检测 ρ 0.96 ~0.30 (Laplacian 对噪声不敏感)
模糊检测 ρ 0.72 ~0.85 (Laplacian 天然适合)
压缩伪影 ρ 0.86 ~0.20 (无对应指标)
亮度异常 0.05 (盲区) ~0.80 (直方图直接捕获)
可解释性 黑盒向量 ✅ 每个指标有物理含义
部署成本 ONNX Runtime ~100ms 零成本,纯数学运算 <1ms

结论:两者互补。传统 CV 覆盖亮度/对比度(模型的盲区),模型覆盖噪声/压缩(CV 的盲区)。最佳方案是双塔融合:本模型 1152D + CV 6D → 联合回归器。

模型选型速查表

你的需求 推荐方案
文档上传前质量检查 本模型 (feature extractor) + 轻量 Two-Tower 回归器
通用图片质量打分(非文档) Q-ReAlign / CLIP-IQA / MUSIQ
端到端 OCR 识别 完整 PaddleOCR-VL / PaddleOCR Server
零依赖快速模糊检测 传统 CV (Laplacian variance)
图片语义搜索/相似度 CLIP / Qwen3-VL-Embedding / DINOv3
移动端文档质量 本模型 ONNX (支持 ONNX Runtime 全平台)

Model Architecture

Input Image (any size)
    │
    ▼
PaddleOCRVLImageProcessor (smart_resize + normalize + patchify)
    │
    ▼
Patch Embedding + Position Encoding
    │
    ▼
Transformer Encoder Layers 0 → 12  ← TRUNCATED HERE (original: 27 layers)
    │
    ▼
Mean Pooling across patches
    │
    ▼
1152-D Feature Vector
Property Value
Parameters 237M (905 MB fp32 weights)
Input pixel_values: (1, N, 3, 14, 14) float32, position_ids: (1, 1) int64
Output layer_12_pooled: (1, 1152) float32
Dynamic axes num_patches (varies with input resolution)
Opset 18
Files model.onnx (851 KB graph) + model.onnx_data (902 MB weights)

Installation

pip install onnxruntime numpy Pillow opencv-python

No PyTorch, no transformers, no CUDA required. CPU inference is ~50-200ms per image.

Quick Start

from inference.onnx_inference import Layer12ONNXExtractor
from PIL import Image, ImageFilter

# Load model
extractor = Layer12ONNXExtractor("model.onnx")

# Extract features
img = Image.open("document.jpg").convert("RGB")
features = extractor.extract(img)  # → np.ndarray shape (1152,)

# Quality via distance from pristine reference
pristine = Image.open("pristine.jpg").convert("RGB")
blurred = img.filter(ImageFilter.GaussianBlur(radius=5))
quality = extractor.quality_score(blurred, reference=pristine)
# → 0.0 (heavily degraded) ~ 1.0 (pristine)

See inference/example.py for a complete walkthrough.

Feature Quality Benchmarks

Run with: python benchmark/run_benchmark.py

1. Degradation Sensitivity (Spearman ρ)

Correlation between feature distance and degradation severity across 12 degradation types × 7 levels × 5 image types:

Degradation Spearman ρ Monotonicity Sensitivity
Gaussian Noise 0.96 0.98 Excellent
JPEG Compression 0.86 0.97 Excellent
Gaussian Blur 0.72 0.95 Good
Motion Blur 0.68 0.93 Good
Median Blur 0.65 0.92 Good
Downscale 0.65 0.90 Good
ISO Noise 0.60 0.88 Fair
Salt & Pepper 0.55 0.85 Fair
Rotation 0.30 0.60 Weak
Brightness 0.05 0.52 Blind spot
Contrast 0.05 0.50 Blind spot

Mean |ρ| = 0.55 across all degradations.

2. Paired Ranking Accuracy

Given a pristine reference and two degraded copies at different severity levels, can the model correctly rank which is worse?

Overall: 89% across 200 random pairs.

3. Cross-Resolution Consistency

Same content at different resolutions → same features? Cosine similarity between features extracted at 224² vs 728²:

Mean: 0.997 — near-perfect consistency thanks to PaddleOCR-VL's smart_resize.

Use Cases

  1. Document upload quality gate — reject blurry/noisy scans before OCR
  2. Image preprocessing quality monitor — detect when enhancement pipeline degrades
  3. Reference-based quality scoring — compare against a known-good template
  4. Feature backbone for quality regression — use 1152D features as input to a lightweight quality regressor (e.g., Two-Tower MLP, XGBoost)

Limitations

  • Blind to brightness/contrast issues — these don't affect the feature space. Complement with traditional CV metrics (Laplacian variance, histogram stats).
  • Geometric degradations (rotation) have weak signal — the model uses position encoding, but rotated text still looks like text to mid-layer features.
  • Not a standalone quality scorer — this is a feature extractor. You need a downstream head (trained regressor, reference comparison, or CV complement) for final quality scores.
  • Original PaddleOCR-VL base model — the vision encoder is from PaddleOCR-VL, which is optimized for Chinese + English document OCR.

Recommended Quality Pipeline

For a robust document quality system, use a Two-Tower fusion:

PaddleOCR-VL Layer 12 (this model)     OpenCV traditional metrics
         │                                      │
   PCA → 128D                             6D vector
   Deep Tower MLP                         CV Tower MLP
         │                                      │
         └──────── concat (80D) ───────────────┘
                         │
                   Shared Head
                         │
                  quality_score [0, 1]

This approach achieves ρ = 0.74 on held-out data (vs 0.60 for XGBoost on raw features).

Open-Source Benchmarks for Evaluation

If you want to benchmark your quality model against published work:

Direct: OCR Quality Assessment

Dataset Size Focus Link
OCR-Quality 1,000 images Human-annotated OCR quality (4-level) Aslan-mingye/OCR-Quality
DocPTBench 1,300+ photos Photographed document quality (geometric + photometric) Topdu/DocPTBench

General Image Quality Assessment (IQA)

Dataset Size Focus Link
LIVE 779 images 5 distortion types, DMOS labels UT Austin LIVE
CSIQ 866 images 6 distortion types, DMOS labels CSIQ
TID2013 3,000 images 24 distortion types, MOS labels TID2013
KADID-10k 10,125 images 25 distortion types × 5 levels KADID-10k
KonIQ-10k 10,073 images In-the-wild quality, MOS labels KonIQ-10k

OCR Recognition Benchmarks

Dataset Size Focus Link
GlotOCR Bench 158 scripts Multilingual OCR with degraded variants HuggingFace (eval license)
Misraj-DocOCR 400 images Arabic document OCR (WER/CER) Open-source
OCRGenBench 1,060 samples OCR generation quality PeirongZhang/OCRGenBench

Recommended starting point: OCR-Quality — it's the only dataset that directly measures OCR quality assessment, with human annotations and per-page MOS scores. Our benchmark script includes a loader for it.

Files

paddleocr-quality-onnx/
├── README.md                        # This file
├── .gitattributes                   # HF LFS config
├── requirements.txt                 # Python dependencies
├── model.onnx           # ONNX model (~0.8 MB)
├── inference/
│   ├── __init__.py
│   ├── preprocessing.py             # Standalone image preprocessing (no torch)
│   ├── onnx_inference.py            # Layer12ONNXExtractor class
│   └── example.py                   # Complete usage example
└── benchmark/
    ├── run_benchmark.py             # Degradation sensitivity, ranking, resolution
    └── results/                     # Benchmark output

Citation

If you use this model in your research:

@software{paddleocr-quality-onnx,
  title = {PaddleOCR-VL Layer-12 Feature Extractor (ONNX)},
  year = {2025},
  note = {Truncated PaddleOCR-VL vision encoder for document image quality assessment},
  url = {https://huggingface.co/[your-username]/paddleocr-quality-onnx},
}

The original PaddleOCR-VL model is from PaddlePaddle. This is a derived work using only the vision encoder (layers 0-12), exported to ONNX for lightweight deployment.

License

This model is derived from PaddleOCR-VL, which is released under Apache 2.0. This ONNX export and accompanying code are also Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train efwfe/paddleocr-quality-onnx