File size: 10,402 Bytes
ff58e5b
 
 
2884ee7
 
ff58e5b
 
2884ee7
 
ff58e5b
 
2884ee7
ff58e5b
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
 
 
 
 
 
2884ee7
 
 
ff58e5b
 
 
 
 
 
 
 
 
 
2884ee7
 
 
ff58e5b
2884ee7
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
ff58e5b
2884ee7
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
 
 
 
 
 
 
 
 
 
 
 
2884ee7
ff58e5b
 
 
 
 
 
 
 
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
 
 
 
 
 
 
 
2884ee7
ff58e5b
 
 
 
 
 
 
 
 
2884ee7
ff58e5b
 
 
2884ee7
ff58e5b
 
 
 
 
 
 
 
2884ee7
ff58e5b
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
 
2884ee7
 
ff58e5b
2884ee7
ff58e5b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
2884ee7
ff58e5b
 
 
 
 
 
 
2884ee7
ff58e5b
 
 
 
 
2884ee7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ff58e5b
2884ee7
ff58e5b
 
 
 
 
2884ee7
ff58e5b
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
---
library_name: pytorch
pipeline_tag: robotics
language:
  - en
tags:
  - worlddit
  - world-action-model
  - world-models
  - libero
  - robot-learning
  - robotic-manipulation
  - imitation-learning
  - diffusion-transformer
  - diffusion-policy
  - flow-matching
inference: false
widget:
  - example_title: "LIBERO Spatial, task 5"
    text: "Successful rollout, front view."
    output:
      url: "https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_spatial_frontview_task05_episode01.mp4"
  - example_title: "LIBERO Object, task 8"
    text: "Successful rollout, agent view."
    output:
      url: "https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_object_agentview_task08_episode01.mp4"
  - example_title: "LIBERO Goal, task 10"
    text: "Successful rollout, side view."
    output:
      url: "https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_goal_sideview_task10_episode01.mp4"
  - example_title: "LIBERO Long, task 6"
    text: "Successful rollout, front view."
    output:
      url: "https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_10_frontview_task06_episode01.mp4"
---

<p align="center">
  <img src="https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/paris2/model-card/v1/bagel_labs_logo.png" alt="Bagel Labs">
</p>

# WorldDiT

## One diffusion backbone learns what to do and what comes next.

<p align="center">
  <a href="https://huggingface.co/bageldotcom/worlddit" target="_blank">
    <img src="https://img.shields.io/badge/πŸ€—_DOWNLOAD_WORLDDIT_WEIGHTS-FFD21E?style=for-the-badge&logoColor=000000" alt="Download WorldDiT Weights">
  </a>
  <a href="https://github.com/Lifelong-Robot-Learning/LIBERO" target="_blank">
    <img src="https://img.shields.io/badge/πŸ€–_LIBERO_BENCHMARK-FF6B6B?style=for-the-badge&logoColor=white" alt="LIBERO Benchmark">
  </a>
</p>

WorldDiT learns continuous robot action chunks and a future visual target
through one shared diffusion transformer. Deployment keeps only the action
path.

This release includes four LIBERO checkpoints, a self contained inference
runtime, and an evaluator for reproducing the reported suite results.

## See WorldDiT act

The four clips below show successful rollouts from the released checkpoints.
Each clip covers a different LIBERO suite and camera view.

<Gallery />

| Suite | View | Video |
|---|---|---|
| LIBERO Spatial | Front view | [Open MP4](https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_spatial_frontview_task05_episode01.mp4) |
| LIBERO Object | Agent view | [Open MP4](https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_object_agentview_task08_episode01.mp4) |
| LIBERO Goal | Side view | [Open MP4](https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_goal_sideview_task10_episode01.mp4) |
| LIBERO Long | Front view | [Open MP4](https://pub-2c09ae97630f4932a23e622b450076e0.r2.dev/worlddit/model-card/v1/worlddit_libero_10_frontview_task06_episode01.mp4) |

## What is in this release

| Release component | Included artifact |
|---|---|
| LIBERO Spatial policy | SafeTensors checkpoint |
| LIBERO Object policy | SafeTensors checkpoint |
| LIBERO Goal policy | SafeTensors checkpoint |
| LIBERO Long policy | SafeTensors checkpoint |
| Model runtime | `inference.py` |
| Evaluation runtime | `eval.py` |
| Frozen encoders | CLIP ViT B 32 and MAE ViT B |
| Configuration | `config.json` |
| Environment | Pinned Python requirements |

The repository is self contained for WorldDiT inference. LIBERO still provides
the benchmark environments, assets, task definitions, and initial states.

## Reported LIBERO results

Across the four released suite checkpoints, WorldDiT records 1,898 successful
episodes out of 2,000 under the selection aware evaluation protocol.

| Suite | Successful episodes | Success rate |
|---|---:|---:|
| LIBERO Spatial | 490 of 500 | 98.0 percent |
| LIBERO Object | 485 of 500 | 97.0 percent |
| LIBERO Goal | 464 of 500 | 92.8 percent |
| LIBERO Long | 459 of 500 | 91.8 percent |
| Selection aware mean | 1,898 of 2,000 | 94.9 percent |

The released runtime and checkpoints were revalidated from a clean installation
on eight RTX Pro 6000 Blackwell GPUs.

The result is selection aware because three hundred episodes per suite informed
staged checkpoint selection before the final five hundred episode score was
assembled.

## Model at a glance

| Property | Released configuration |
|---|---|
| Total parameters | 399.084 million |
| Trainable parameters | 135.107 million |
| Observation context | Three frames |
| Predicted action horizon | Seven actions |
| Executed before replanning | Three actions |
| Action dimension | Seven |
| Visual encoder | MAE ViT B |
| Language encoder | OpenAI CLIP ViT B 32 |
| Checkpoint format | SafeTensors |
| Evaluation environment | Headless LIBERO with EGL |

## Run a smoke test

Download the repository and create a clean Python 3.12 environment.

```bash
hf download bageldotcom/worlddit --local-dir worlddit
cd worlddit

python3.12 -m venv venv
source venv/bin/activate
python -m pip install -r requirements.txt
python -m pip install --no-deps robosuite==1.4.1
```

LIBERO supplies the benchmark definitions, assets, and initial states. Keep the
checkout at `~/LIBERO`, which is the evaluator's default.

```bash
git clone https://github.com/Lifelong-Robot-Learning/LIBERO.git ~/LIBERO
```

The released evaluation was validated with LIBERO commit
`8f1084e3132a39270c3a13ebe37270a43ece2a01`.

```bash
python eval.py \
  --suite libero_spatial \
  --gpus 1 \
  --tasks 1 \
  --episodes 1 \
  --max-steps 20 \
  --output-dir results/smoke
```

A successful smoke test confirms that the environment, checkpoint, visual
encoders, simulator, and rendering path load together. It is not a benchmark
result.

## How WorldDiT works

WorldDiT uses three recent observations, robot state, and language as context.
During training, one diffusion transformer learns a seven step action chunk and
an auxiliary future visual target. During deployment, the future visual path is
absent. The policy executes the first three predicted actions, observes again,
and replans.

> Future visual prediction is a training signal, not a deployment path.

| Training | Deployment |
|---|---|
| Action and future visual targets share one backbone | Only the action path remains |
| Seven action steps are supervised | Seven actions are predicted |
| Future visual supervision is present | No future visual output is requested |
| The complete training objective is active | Three actions execute before replanning |

## Evaluation

### One GPU

```bash
python eval.py \
  --suite libero_spatial \
  --gpus 1 \
  --output-dir results/libero_spatial
```

### Multiple GPUs

```bash
CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7 python eval.py \
  --suite libero_spatial \
  --gpus 8 \
  --output-dir results/libero_spatial_8gpu
```

Each GPU receives an independent progress bar. After all workers finish, rank 0
prints per task and overall success rates and writes a structured
`results.json`. Output directories must be new so an earlier evaluation is
never overwritten.

Supported suites.

```text
libero_spatial
libero_object
libero_goal
libero_10
```

## What this repository contains

```text
.
β”œβ”€β”€ checkpoints/
β”‚   β”œβ”€β”€ libero_10/model.safetensors
β”‚   β”œβ”€β”€ libero_goal/model.safetensors
β”‚   β”œβ”€β”€ libero_object/model.safetensors
β”‚   └── libero_spatial/model.safetensors
β”œβ”€β”€ dependencies/
β”‚   β”œβ”€β”€ ViT-B-32.pt
β”‚   └── mae_pretrain_vit_base.pth
β”œβ”€β”€ eval.py
β”œβ”€β”€ inference.py
β”œβ”€β”€ config.json
└── requirements.txt
```

`dependencies/` contains the frozen visual and language encoder weights needed
by the released policy. No additional model downloads are required.

## Inference API

```python
from inference import load_model

model = load_model(".", suite="libero_spatial", device="cuda")
actions = model(primary_images, wrist_images, robot_state, text_tokens)
```

| Input or output | Shape |
|---|---|
| Primary-camera images | `[B, 3, 3, 224, 224]` |
| Wrist-camera images | `[B, 3, 3, 224, 224]` |
| Robot state | `[B, 3, 8]` |
| OpenAI CLIP text tokens | `[B, 3, 77]` |
| Predicted action tensor | `[B, 3, 7, 7]` |

Evaluation uses the final temporal slot of the predicted action tensor.

## Architecture details

| Component | Specification |
|---|---|
| Policy | WorldDiT diffusion transformer |
| Observation context | 3 frames |
| Action horizon | 7 actions |
| Action dimension | 7 |
| Action aggregation | Temporal ensembling |
| Language encoder | OpenAI CLIP ViT-B/32 |
| Visual encoder | MAE ViT-B |
| Evaluation | Headless LIBERO with EGL |
| Checkpoint format | SafeTensors |

## Intended use

WorldDiT is intended for research on language conditioned robot manipulation in
the LIBERO simulator. The released checkpoints support reproduction,
evaluation, and architecture research across the four released suites.

## Scope of the release

The reported results describe LIBERO simulation under the released evaluation
protocol. They do not establish real robot reliability, safety, or transfer
across embodiments.

The present release does not isolate the causal contribution of the future
visual target. Total parameter count also does not measure training cost,
deployment latency, or runtime efficiency.

## Authors and contact

WorldDiT is developed by Sen Wang, Praveen Rajasekhar, Bidhan Roy, and Marcos
Villagra at Bagel Labs. Questions can be sent to research@bagel.com.

## Acknowledgments

This release builds on
[LIBERO](https://github.com/Lifelong-Robot-Learning/LIBERO),
[robosuite](https://github.com/ARISE-Initiative/robosuite),
[OpenAI CLIP](https://github.com/openai/CLIP), and
[Masked Autoencoders](https://github.com/facebookresearch/mae). Third party
components remain subject to their respective upstream terms.

---

<div style="display: flex; align-items: center; gap: 8px;">
  <span>Made with ❀️ by</span>
  <a href="https://twitter.com/bageldotcom" target="_blank">
    <img src="https://img.shields.io/badge/Bagel_Labs-1DA1F2?style=for-the-badge&logo=twitter&logoColor=white" alt="Follow Bagel Labs on Twitter" height="28">
  </a>
</div>