RT-DocLayout: Real-Time End-to-End Document Layout Analysis with Reading Order in the Wild Paper • 2606.23344 • Published Jun 22 • 2
PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks Paper • 2606.13108 • Published Jun 11 • 9
PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks Paper • 2603.24373 • Published Mar 25
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Paper • 2606.03264 • Published Jun 2 • 27
Real5-OmniDocBench: A Full-Scale Physical Reconstruction Benchmark for Robust Document Parsing in the Wild Paper • 2603.04205 • Published Mar 4 • 3
PP-DocBee: Improving Multimodal Document Understanding Through a Bag of Tricks Paper • 2503.04065 • Published Mar 6, 2025
PP-FormulaNet: Bridging Accuracy and Efficiency in Advanced Formula Recognition Paper • 2503.18382 • Published Mar 24, 2025
PP-DocBee2: Improved Baselines with Efficient Data for Multimodal Document Understanding Paper • 2506.18023 • Published Jun 22, 2025
Sortblock: Similarity-Aware Feature Reuse for Diffusion Model Paper • 2508.00412 • Published Aug 1, 2025
PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing Paper • 2601.21957 • Published Jan 29 • 23
Running Agents Featured 59 ERNIE-4.5-VL-28B-A3B-Thinking Demo 👐 59 Compact model, powerful multimodal reasoning.
PP-MobileSeg: Explore the Fast and Accurate Semantic Segmentation Model on Mobile Devices Paper • 2304.05152 • Published Apr 11, 2023
RT-DETRv2: Improved Baseline with Bag-of-Freebies for Real-Time Detection Transformer Paper • 2407.17140 • Published Jul 24, 2024 • 2
PP-DocLayout: A Unified Document Layout Detection Model to Accelerate Large-Scale Data Construction Paper • 2503.17213 • Published Mar 21, 2025 • 2