Papers
arxiv:2608.22591

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

Published on Aug 23
· Submitted by
ychunkai
on Aug 25
Authors:
,

Abstract

WorldToken fuses heterogeneous robot observations into per-timestep world tokens processed by a causal Transformer and diffusion action head, with scaling and temporal-context analyses on RoboCasa and RMBench.

Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token. A causal temporal Transformer models the resulting world-token sequence, and a diffusion action head generates action chunks. On 23 RoboCasa tasks, an 85.3M-parameter policy trained from scratch apart from a frozen pretrained CLIP text encoder achieves 59.45% mean closed-loop success using 2,900 generated demonstrations per task. A complete factorial sweep over five dataset sizes, five model sizes, and two training seeds shows consistent gains from additional target-domain data and diminishing returns beyond moderate model size. Under same-checkpoint history truncation, reducing visible history to one or two policy timesteps lowers closed-loop success for all 50 RoboCasa policies. On RMBench Blocks Ranking, reducing visible history from 146 to 8 seconds lowers evaluator success from 95% to 28%, while an exploratory extended rollout sustains the reference swap sequence for over 850 seconds. These results establish the empirical feasibility of the complete WorldToken instantiation and characterize its data-scaling and temporal-context behavior under the tested recipes. They do not establish superiority over alternative sequence organizations or isolate which components of the complete implementation drive the observed performance.

Community

Paper author Paper submitter

Can a robot read the physical world as a language model reads text? Inspired by this question, we introduce WorldToken, a time-first approach to robotic sequence modeling in which policy timesteps define the top-level temporal sequence. We study its scaling and temporal-context behavior across RoboCasa and RMBench, including controlled history truncation and extended rollouts that sustain ordered behavior for over 850 seconds.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.22591
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.22591 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.22591 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.22591 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.