AnyPos / README.md
nielsr's picture
nielsr HF Staff
Add model card for AnyPos
984a81c verified
|
Raw
History Blame
2.4 kB
metadata
pipeline_tag: robotics

AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation

AnyPos (also referred to as Vidar in the codebase) is a unified pipeline that integrates large-scale automated task-agnostic exploration with robust embodiment modeling through inverse dynamics learning. By focusing on exploring all feasible actions of the embodiment, AnyPos learns embodiment dynamics directly from task-agnostic action data and decouples them from high-level policy learning.

Introduction

AnyPos generates diverse yet safe trajectories at scale, then learns embodiment representations by decoupling arm and end-effector motions. It employs a direction-aware decoder to stabilize predictions under distribution shift, allowing it to be coupled with diverse high-level policy models. The system has demonstrated significant success rate improvements across tasks such as operating a microwave, folding clothes, and watering plants.

Environment Setup

The code has been tested with CUDA 12.4.

1. Create and Activate Conda Environment

conda create -n vidar python==3.11.9
conda activate vidar

2. Install Dependencies

conda install pytorch==2.4.0 torchvision==0.19.0 torchaudio==2.4.0 pytorch-cuda=12.4 -c pytorch -c nvidia
python -m pip install -r requirements.txt
python -m pip install ninja
python -m pip install git+https://github.com/Dao-AILab/flash-attention.git@v2.6.3
python -m pip install xfuser==0.4.0

Usage

The AnyPos pipeline consists of a Video Diffusion Model and a Masked Inverse Dynamic Model.

Video Diffusion Model (Vidar) Inference

To generate a video based on an initial frame and a text instruction:

scripts/vm/sample.sh

Masked Inverse Dynamic Model Inference

To evaluate the inverse dynamic model:

scripts/idm/eval.sh

Citation

@article{anypos2025,
  title={AnyPos: Automated Task-Agnostic Actions for Bimanual Manipulation},
  author={Yao, Feng and others},
  journal={arXiv preprint arXiv:2507.12768},
  year={2025}
}