RECAP-Forcing: Retaining Content Appearances for Long Video Generation Paper • 2608.26671 • Published 8 days ago • 5
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning Paper • 2512.08135 • Published Dec 9, 2025
VideoNSA: Native Sparse Attention Scales Video Understanding Paper • 2510.02295 • Published Oct 2, 2025 • 10
OverLayBench: A Benchmark for Layout-to-Image Generation with Dense Overlaps Paper • 2509.19282 • Published Sep 23, 2025 • 8
YOLO-Count: Differentiable Object Counting for Text-to-Image Generation Paper • 2508.00728 • Published Aug 1, 2025
Bayesian Diffusion Models for 3D Shape Reconstruction Paper • 2403.06973 • Published Mar 11, 2024
DepR: Depth Guided Single-view Scene Reconstruction with Instance-level Diffusion Paper • 2507.22825 • Published Jul 30, 2025
Science-T2I: Addressing Scientific Illusions in Image Synthesis Paper • 2504.13129 • Published Apr 17, 2025 • 3
OmniControlNet: Dual-stage Integration for Conditional Image Generation Paper • 2406.05871 • Published Jun 9, 2024
RECAP-Forcing: Retaining Content Appearances for Long Video Generation Paper • 2608.26671 • Published 8 days ago • 5
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published Jul 27 • 80