SGF+: Decoupling Gradient Flows for Autoregressive Video Generation Paper • 2610.10429 • Published 2 days ago • 53
VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations Paper • 2610.00994 • Published 8 days ago • 26
view article Article Live Human Feedback in the Training Loop: Aligning Diffusion Models with Real People Rapidata • 7 days ago • 10
GUI vs. CLI: Execution Bottlenecks in Screen-Only and Skill-Mediated Computer-Use Agents Paper • 2606.24551 • Published Jun 22 • 29
Semantic Browsing: Controllable Diversity for Image Generation Paper • 2606.23679 • Published Jun 22 • 20
Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models Paper • 2606.25041 • Published Jun 23 • 126
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer Paper • 2605.30409 • Published May 28 • 39
Representation Forcing for Bottleneck-Free Unified Multimodal Models Paper • 2605.31604 • Published May 29 • 62
LongTraceRL: Learning Long-Context Reasoning from Search Agent Trajectories with Rubric Rewards Paper • 2605.31584 • Published May 29 • 41
Meta-CoT: Enhancing Granularity and Generalization in Image Editing Paper • 2604.24625 • Published Apr 27 • 26
DeepGen 1.0: A Lightweight Unified Multimodal Model for Advancing Image Generation and Editing Paper • 2602.12205 • Published Feb 13 • 83 • 6
FrankenMotion: Part-level Human Motion Generation and Composition Paper • 2601.10909 • Published Jan 15 • 19