VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published 23 days ago • 272
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 61
Flow-DPPO: Divergence Proximal Policy Optimization for Flow Matching Models Paper • 2606.11025 • Published Jun 9 • 42
Composition of Memory Experts for Diffusion World Models Paper • 2605.18813 • Published May 12 • 2
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control Paper • 2412.11198 • Published Dec 15, 2024 • 2
Rethinking Visual Intelligence: Insights from Video Pretraining Paper • 2510.24448 • Published Oct 28, 2025 • 7
World Model Self-Distillation: Training World Models to Solve General Tasks Paper • 2606.12072 • Published Jun 10 • 16
Communication-Inspired Tokenization for Structured Image Representations Paper • 2602.20731 • Published Feb 24 • 4
Advancing End-to-End Pixel Space Generative Modeling via Self-supervised Pre-training Paper • 2510.12586 • Published Oct 14, 2025 • 115
Diffusion Transformers with Representation Autoencoders Paper • 2510.11690 • Published Oct 13, 2025 • 171
Temporal In-Context Fine-Tuning for Versatile Control of Video Diffusion Models Paper • 2506.00996 • Published Jun 1, 2025 • 40