KeyRec: Bounded Visual Memory for Streaming and Long-Video Understanding Paper • 2609.32182 • Published 11 days ago • 11
DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration Paper • 2609.33403 • Published 10 days ago • 15
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation Paper • 2609.11412 • Published 27 days ago • 48
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 408
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published Aug 31 • 316
SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object Manipulation Paper • 2608.18701 • Published Aug 19 • 15
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time Paper • 2608.11745 • Published Aug 13 • 25
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published Aug 12 • 43
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
Mendel Gödel Machine: Recursive Self-Improving Coding Agents via Comparative Evolution Paper • 2608.07645 • Published Aug 7 • 27
World-to-Wrist: Task-Conditioned Future Wrist Modeling for Fine-Grained Robot Manipulation Paper • 2608.05369 • Published Aug 5 • 27
GST-Bench: Can VLMs Develop Global Spatial Awareness from Video? Paper • 2608.05747 • Published Aug 6 • 48
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107