Beyond Accuracy: ARIA-Rubrics for Evaluating Audio Reasoning in Large Audio Language Models
Paper • 2609.09681 • Published
None defined yet.
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
ContextBias: Controlled Evaluation of Bias Persistence Under Context Shift in Text-to-Image Models