Papers
arxiv:2609.35652

MM-ABC: Towards Generalist Mobile Manipulation via Seeing, Coordinating and Imagining

Published on Sep 28
· Submitted by
KolaKivy
on Oct 6
Authors:
,
,
,
,
,
,
,
,

Abstract

Mobile manipulation extends robot interaction beyond a fixed kinematic workspace by making the reachable region itself controllable. This flexibility introduces two central challenges: spatially grounded perception under continuous ego-motion and coordinated control of heterogeneous arm and base actions. Existing approaches strengthen geometry through explicit 3D representations or predictive world models, and often decouple mobility and manipulation into separate action streams. We argue that effective mobile manipulation requires not only decoupling, but also representations that support efficient cross-stream collaboration. We present MM-ABC, a foundation model built around Seeing, Coordinating, and Imagining Arm-Base Collaboration. MM-ABC combines sparse multi-level VLM features for spatial perception; a training-only future branch that uses world imagination and geometric intent as extra supervision, strengthening perception and manipulation-intent prediction and improving the overall learning signal; and MM-APT, which coordinates separate manipulation and mobility streams through masked joint attention and clean-action x-prediction. In controlled ablations, replacing clean-action prediction with velocity prediction lowers success on RoboCasa365 composite-seen tasks from 32.8% to 29.2%, and removing future supervision or multilevel conditioning causes larger drops. We pretrain MM-ABC on 5,000+ hours of heterogeneous robot data spanning 400K+ episodes, 12 datasets, and 17 embodiments. Experiments cover EBench, RoboCasa365, ManiSkill-HAB, LIBERO, LIBERO-Plus, and real-world mobile manipulation. MM-ABC achieves 44.71% success on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, 82.8% on LIBERO-Plus without perturbation training, and 83% mean success on five real-world tasks.

Community

Paper submitter

We introduce MM-ABC, a generalist foundation model for mobile manipulation built around Seeing, Coordinating, and Imagining. It combines sparse multi-level VLM features, a two-stream arm–base action transformer with clean-action x-prediction, and future geometric supervision for jointly learning perception and coordinated whole-body control.

MM-ABC is pretrained on 5,000+ hours, 400K+ episodes, and 17 robot embodiments, and evaluated across five simulation benchmarks and real-world mobile manipulation. It achieves 44.71% on EBench, 61.2% on RoboCasa365, 99.1% on LIBERO, and 83% average success in real-world tasks.

Project page and the MM-30 real-world mobile manipulation dataset are publicly available. The code is coming soon

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.35652
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.35652 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.35652 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.