Learning Foresight without Explicit Trajectories for 3D Diffusion Policies Paper • 2609.20669 • Published 8 days ago • 8
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 8 days ago • 41
Attention-DP3: Spatially Object-aware 3D Diffusion Policy via Geometry-aligned Attentional Conditioning Paper • 2609.13318 • Published 15 days ago • 7
Mono2Stereo: A Benchmark and Empirical Study for Stereo Conversion Paper • 2503.22262 • Published Mar 28, 2025 • 1
OASIS: Open Agent Social Interaction Simulations with One Million Agents Paper • 2411.11581 • Published Nov 18, 2024 • 1
GLDesigner: Leveraging Multi-Modal LLMs as Designer for Enhanced Aesthetic Text Glyph Layouts Paper • 2411.11435 • Published Nov 18, 2024
MA-VLA: Multi-Arm Vision-Language-Action Model for Collaboration and Compositional Generalization Paper • 2608.25864 • Published 30 days ago • 9
From Forecasting to Planning: Policy World Model for Collaborative State-Action Prediction Paper • 2510.19654 • Published Oct 22, 2025
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective Paper • 2509.18905 • Published Sep 23, 2025 • 31
PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety Paper • 2401.11880 • Published Aug 20, 2024
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 8 days ago • 41
Autonomous Character-Scene Interaction Synthesis from Text Instruction Paper • 2410.03187 • Published Oct 4, 2024 • 8
Build error Agents 65 ArxivCopilot 🏢 65 Generate personalized research profiles and chat with Arxiv Copilot
Isomer: Isomerous Transformer for Zero-shot Video Object Segmentation Paper • 2308.06693 • Published Aug 13, 2023
Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning Paper • 2307.14786 • Published Jul 27, 2023
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception Paper • 2403.02969 • Published Mar 5, 2024 • 1