PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 9 days ago • 12
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 9 days ago • 30
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Paper • 2609.25165 • Published 8 days ago • 74
RULER: Instance-aware Rubric Rewards for SVG Generation Paper • 2609.25270 • Published 8 days ago • 101
Training-Adaptive Convolutional Sparse Coding via Information Bottleneck for Robust Visual Representation Paper • 2609.19122 • Published 12 days ago • 5
Srijika: OpenType-Layout-Reusing Font Restyling for Nine Indic Scripts Paper • 2609.05661 • Published 25 days ago • 41
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 12 days ago • 53
The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing Prediction Paper • 2609.18063 • Published 13 days ago • 19
LynnReal-Omni: Native multi-modal Video Generation for Agentic Visual Workflows Paper • 2609.15863 • Published 15 days ago • 56
Disentangling Representation Evolution in Transformers through Directional Decomposition Paper • 2609.15975 • Published 15 days ago • 9
BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender Paper • 2609.15478 • Published 15 days ago • 34
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 15 days ago • 214