OmniEcho: Spatial Audio Understanding for Embodied Agents Paper • 2609.23407 • Published 6 days ago • 20
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach Paper • 2607.28175 • Published Jul 30 • 1
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 18 days ago • 165
EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing Paper • 2608.18063 • Published Aug 18 • 23
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation Paper • 2607.27278 • Published Jul 29 • 15
Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation Paper • 2605.29430 • Published May 28 • 3
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm Paper • 2511.04570 • Published Nov 6, 2025 • 242
Describe Anything Collection Multimodal Large Language Models for Detailed Localized Image and Video Captioning • 7 items • Updated Aug 11 • 64