Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 11 days ago • 58
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published Jul 9 • 6
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models Paper • 2606.29600 • Published Jun 28 • 6
Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents Paper • 2606.23085 • Published Jun 22 • 15
See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents Paper • 2606.13594 • Published Jun 11 • 6
AFUN: Towards an Affordance Foundation Model for Functionality Understanding Paper • 2606.02551 • Published Jun 1 • 7
TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment Paper • 2602.13579 • Published Feb 14 • 11
Next-Embedding Prediction Makes Strong Vision Learners Paper • 2512.16922 • Published Dec 18, 2025 • 91
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs Paper • 2505.23816 • Published May 27, 2025
Neural Generation Meets Real People: Building a Social, Informative Open-Domain Dialogue Agent Paper • 2207.12021 • Published Jul 25, 2022
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Paper • 2412.18619 • Published Dec 16, 2024 • 60
Alternating Recurrent Dialog Model with Large-scale Pre-trained Language Models Paper • 1910.03756 • Published Oct 9, 2019
Hierarchical Task Learning from Language Instructions with Unified Transformers and Self-Monitoring Paper • 2106.03427 • Published Jun 7, 2021
A Probabilistic End-To-End Task-Oriented Dialog Model with Latent Belief States towards Semi-Supervised Learning Paper • 2009.08115 • Published Sep 17, 2020
Task-Oriented Dialog Systems that Consider Multiple Appropriate Responses under the Same Context Paper • 1911.10484 • Published Nov 24, 2019
Grounding Visual Illusions in Language: Do Vision-Language Models Perceive Illusions Like Humans? Paper • 2311.00047 • Published Oct 31, 2023 • 10
GROUNDHOG: Grounding Large Language Models to Holistic Segmentation Paper • 2402.16846 • Published Feb 26, 2024
DANLI: Deliberative Agent for Following Natural Language Instructions Paper • 2210.12485 • Published Oct 22, 2022