SO-101 Cup Nesting Policy Report
Data screening and training report for an SO-101 ACT policy
Physical AI
This dataset is a curated collection of real-world teleoperation data captured on the SO-101 robotic arm (so_follower), built to support training and evaluation of modern robot-learning models — from imitation-learning policies to large-scale Vision-Language-Action (VLA) and world models.
Each episode is a human-teleoperated demonstration of a manipulation task, recorded synchronously across two camera viewpoints alongside the arm's full proprioceptive state and control signals. Data collection followed a rigorous, standardized protocol:
This combination of careful human demonstration and disciplined data engineering is intended to make the dataset a dependable foundation for downstream policy and representation learning, rather than a loosely collected video corpus.
The dataset follows the Hugging Face LeRobotDataset v3.0 format. Unlike the earlier v2.1 format (one file per episode), v3.0 packs many episodes into a smaller number of larger, chunked files, with episode boundaries resolved through relational metadata rather than filenames. This makes the dataset scalable, faster to load, and streaming-ready directly from the Hub.
physical-ai/
└── SO-101/
└── <task_name>/ # e.g. cup_nesting
├── data/
│ └── chunk-000/
│ └── file-000.parquet # joint states, actions, indices — many episodes per file
├── meta/
│ ├── info.json # schema, fps, robot type, chunking config
│ ├── stats.json # per-feature normalization statistics
│ ├── tasks.parquet # task_index -> natural-language task description
│ └── episodes/
│ └── chunk-000/
│ └── file-000.parquet # per-episode lengths, task refs, file/byte offsets
└── videos/
├── observation.images.cam_front/
│ └── chunk-000/
│ └── file-000.mp4 # front-view camera, many episodes per file
└── observation.images.cam_top/
└── chunk-000/
└── file-000.mp4 # top-down camera, many episodes per file
Key structural points:
data/ — Apache Parquet shards containing frame-level observation.state, action (both 6-DoF: shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper), plus timestamp, frame_index, episode_index, and task_index.videos/ — Two synchronized camera streams per episode:observation.images.cam_front — front-facing view of the workspaceobservation.images.cam_top — top-down view of the workspace
Both are AV1-encoded MP4 at 480×640, 30 fps, with no audio.meta/ — Self-describing metadata: schema/config (info.json), normalization stats (stats.json), the task-language mapping (tasks.parquet), and per-episode index records (episodes/).chunk-000, chunk-001, …). The exact location of any given episode — which chunk, which file, and its frame offset — is resolved by looking it up in meta/episodes/, not by filename. This is what allows the dataset to scale to many episodes and tasks without file-system overhead.cup_nesting) is organized as its own self-contained LeRobotDataset directory under SO-101/, with its own data/, meta/, and videos/ subfolders.Data collection was designed to produce demonstrations that are diverse and robust enough for policies trained on them to generalize beyond the exact conditions seen during capture:
This dataset is designed to support several stages of the physical-AI model development stack:
Researchers and engineers working on imitation learning, robot foundation models, and embodied AI more broadly are the primary intended audience.
If you use this dataset in your work, please cite it as follows:
@misc{decisionfacts_teleops_dataset,
credits = {Prabhu Raghav, Sreeram B Unni, Balamurugan Pandi, Sriram Gopalan},
year = {2026},
howpublished = {\url{https://huggingface.co/datasets/DecisionFacts/physical-ai}}
}