EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
Abstract
We study coding agents for long-horizon, dexterous robotics and ask whether their solutions can provide scalable supervision for learning general robot policies. To test this, we develop EMBODIEDSWE-BENCH, a simulation benchmark for coding agents spanning contact-rich manipulation, deformable objects, and long-horizon tasks requiring up to half an hour of continuous interaction. We find that frontier coding agents can solve complex long-horizon tasks and transfer prior solutions across both tasks and embodiments. We also design supporting tools that help agents more effectively solve these tasks. However, the resulting solutions require substantial iterative interaction and are typically specialized to individual task instances. We therefore introduce EMBODIEDSWE-GEN, which expands a single solution from coding agent into large diverse trajectories for training a VLA. VLA performance improves with more generated demonstrations, and agent-aided diversification improves generalization to held-out task variations. We also show that a VLA finetuned solely on coding-agent-generated simulation demonstrations completes a long-horizon task on real robot. Together, our framework uses coding agents to solve complex robotics tasks and turn verified solutions into scalable supervision for robot policies.
Community
EmbodiedSWE: Coding Agents for Long Horizon Dexterous Robotics
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- EgoWild2Dex: Learning Dexterous Robotic Manipulation from In-the-Wild Human Experience (2026)
- Trajectory Divergence Horizon Decision for Reliable Dual-Arm Surgical Subtask Manipulation (2026)
- TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation (2026)
- Beyond Flat Policies: Hierarchical Post-Training for Embodied Agents in Robotic Manipulation (2026)
- RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents (2026)
- RoboSynChallenge: Mastering Real-World Dexterity via Generalizing Synthesized Manipulation Skills (2026)
- DREAM: Deployment-Time Demonstration Generation via Real-to-Sim for Scalable Policy Adaptation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.27308 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper