X-Planner: Event-Structured Task Planning for Embodied Intelligence
Abstract
Task planning bridges high-level instructions and executable behavior in long-horizon manipulation, yet modern Vision-Language-Action (VLA) systems often leave this intermediate structure implicit. Existing chain-of-thought (CoT) planners also tend to rely on coarse task-level annotations or serialize long reasoning traces token by token. We present X-Planner, a planning front-end that addresses both the supervision and representation of embodied reasoning. Our planning data combine Ego, UMI, and teleoperation under a hierarchy granularity with source-dependent annotation depth. Takeover-time annotations and human-designed failures supervise ongoing error recognition. On the model side, a shared VLM backbone exposes two event-structured plan forms: a discrete interface that emits interpretable event states and a latent interface that relays continuous CoT states across staggered Transformer depths through Staircase Decoding. A frozen latent-to-text reconstruction objective provides a semantic anchor for the latent representation. Offline two-step planning evaluation places X-Planner second among four evaluated models on both BERTScore-F1 and a judge-based Overall score. In real-robot experiments, respectively, outperforming the evaluated baselines. These results characterize planning-text quality and downstream execution.
Community
To address long-horizon robotic manipulation tasks, we propose X-Planner, an embodied long-horizon task planner:
It decomposes high-level natural language instructions into event-level subtasks, and can also output continuous implicit Chain-of-Thought (CoT). It directly interfaces with downstream VLA or WAM models to operate robots.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models (2026)
- LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation (2026)
- XCoT-VLA: Executable Chain-of-Thought for Vision-Language-Action Driving (2026)
- EMPIRE: Explicit Manipulation Planning as a Learnable Intermediate Representation for Egocentric Hand-Motion Forecasting (2026)
- HINT: Human-Intent Inception for Long-Horizon Robot Manipulation (2026)
- MobileVLA-R1 2.0: RL-Enhanced Reasoning for Mobile Robot Control (2026)
- Explicit Language Memory for Long-Horizon Planning in Vision-Language-Action Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.25187 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 1
Datasets citing this paper 1
x-square-robot/xplanner-benchmark
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper