Papers
arxiv:2609.33748

AnyStep-WAM: Budget-Aligned Distillation and Adaptive Inference for World Action Models

Published on Sep 27
· Submitted by
Rui Wang
on Sep 30
Authors:
,
,
,
,
,
,

Abstract

World-action models (WAMs) couple predictive visual modeling with action generation, typically relying on iterative denoising with a fixed denoising steps. However, manipulation tasks contain actions chunks with varying sensitivity to generation errors: critical actions require precision, while less sensitive actions allow faster generation with fewer denoising steps. Here we introduce AnyStep World Action Model, a general framework for tunable-budget prediction and scene-dependent computation allocation. Our budget-aligned teacher-trajectory distillation trains interval-conditioned flow maps using explicit frozen-teacher transitions and shared low-rank adapters, supporting action generation from one-step prediction to multi-step refinement. Building on this capability, a lightweight risk-benefit scheduler predicts teacher-curvature-based difficulty and budget-specific student-teacher fidelity from a single one-step preview, selecting the smallest budget predicted to satisfy risk-adaptive fidelity requirements. We evaluate our framework on three widely used WAMs Motus, FastWAM, and LingBotVA using RoboTwin 2.0. Our method reduces average denoising steps by 60.2%, 49.8%, and 85.28%, respectively, while maintaining baseline task success rates. In particular, our AnyStep training substantially improves model performance under a one-step denoising budget, increasing task success rates by 7.07%, 12.08%, and 8.94% on Motus, FastWAM, and LingBotVA, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.

Community

Paper submitter

World-action models (WAMs) typically use fixed-step denoising, despite varying precision requirements across manipulation stages. We introduce AnyStep WAM, a framework for cross-budget prediction and scene-dependent computation allocation. Budget-aligned teacher-trajectory distillation with shared low-rank adapters enables action generation from one-step prediction to multi-step refinement. A lightweight risk-benefit scheduler predicts difficulty and budget-specific student–teacher fidelity from a single preview, selecting the smallest budget predicted to meet adaptive fidelity requirements. On RoboTwin 2.0, AnyStep reduces average denoising steps by 60.2%, 49.8%, and 85.28% on Motus, FastWAM, and LingBotVA, respectively, while maintaining baseline success rates. It also improves one-step success rates by 7.07, 12.08, and 8.94 percentage points, respectively. Experiments on six real-world manipulation tasks further validate its effectiveness.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.33748
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33748 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33748 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33748 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.