HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models Paper • 2610.05739 • Published 3 days ago • 14
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 7 days ago • 61
CARDEA: Auditable Reasoning Grounded in Spatial Evidence for End-to-End Coronary Angiography Interpretation Paper • 2609.06931 • Published Sep 7 • 27
Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation Paper • 2609.08084 • Published about 1 month ago • 72