Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 6 days ago • 298
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 7 days ago • 138
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 14 days ago • 308
Computational Humor with Multimodal LLMs: Methods, Datasets, Evaluation, and Challenges Paper • 2607.19011 • Published 15 days ago • 2
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 16 days ago • 61
Demystifying On-Policy Distillation: Roles, Pathologies, and Regulations Paper • 2607.13399 • Published 21 days ago • 20
electricsheepafrica/africa-ghana-national-fire-outbreaks-2000-2012-a9f5cb20 Viewer • Updated 17 days ago • 255 • 52 • 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published 20 days ago • 206
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published 20 days ago • 171
AI translation of literary texts is "fine", but readers still prefer human translations Paper • 2606.26040 • Published Jun 24 • 9
MemSlides: A Hierarchical Memory Driven Agent Framework for Personalized Slide Generation with Multi-turn Local Revision Paper • 2606.17162 • Published Jun 15 • 177
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models Paper • 2606.11324 • Published Jun 9 • 172