Spaces:
Running
Running
| title: README | |
| emoji: 📚 | |
| colorFrom: blue | |
| colorTo: blue | |
| sdk: static | |
| pinned: false | |
| short_description: Modalitynet World Compiler by Noitom Robotics | |
| # Noitom Robotics | |
| **Making the physical world learnable.** | |
| Noitom Robotics builds human-centric data infrastructure for Physical AI. Our thesis, laid out in *The World Compiler*, is that the bottleneck of Physical AI is not data scarcity but learnability scarcity: the physical world already produces vast embodied intelligence, yet almost none of it exists in a form machines can learn from — and accumulation alone will not close that gap. Volume and learnability must scale together. **ModalityNet** ([modalitynet.com](https://modalitynet.com)) is our implementation of that thesis: compact, fully structured corpora that compile far larger, weakly structured data into learnable form. The name deliberately echoes ImageNet — where ImageNet organized visual reality into a learnable substrate for vision, ModalityNet organizes physical reality, across modalities, into a learnable substrate for Physical AI. | |
| We build at the human layer because human physical intelligence is the one prior every embodiment, architecture, and paradigm shares. Captured once, it transfers to all of them. | |
| ## Three corpora, three priors | |
| - **High Precision Human Interaction, Motion with Object and Vision (HiPHI-MOV)** — the motion prior. Whole-body motion with interacted-object tracking and side-view vision, captured on hybrid optical–inertial systems and exported as BVH skeletons with end-effector 6-DOF poses. It answers: *how does the body move?* | |
| - **High Precision Human Interaction, Omni-Modality (HiPHI-OM)** — the interaction prior. Hand–object tracking held to millimeter error, synchronized with full-body and finger motion, object mesh tracking, tactile and pressure signals, and ego- and side-view RGB-D. It answers: *why do interactions succeed or fail?* | |
| - **In-The-Wild (ITW)** — the distribution prior. Stereo ego vision, sparse body sensing, and audio captured in unconstrained daily environments, preserving the true distribution of physical reality — long-tail cases included. It answers: *does learned behavior hold in the real world?* | |
| The three corpora are designed for joint use: the high-precision layers act as a compiler toolchain that raises the learnability of in-the-wild data at scale, narrowing the gap between demonstration and real-world deployment. | |
| ## Scale | |
| HiPHI production runs at 100,000+ hours per year, and we work with close to 100 companies across robotics, embodied AI, and world modeling. | |
| ## Work with us | |
| None of this is work we do alone. Sample data, modality definitions, and the ModalityNet Technical Specification are available at [modalitynet.com](https://modalitynet.com) — and we welcome researchers and teams building humanoid policies, world models, and vision-language-action (VLA) systems to build with us. | |
| - **Tech Blog** — Episode 1, *ModalityNet: The Art of Modalities in Human-Centric Data* · Episode 2, *The World Compiler*: [noitomrobotics.com/tech-blog](https://noitomrobotics.com/tech-blog/) | |
| - **Company**: [noitomrobotics.com](https://noitomrobotics.com) | |