TBD-VLA single-arm base

Feature Shape Description
observation.images.image (3, 256, 256) RGB image, channels first
observation.images.image2 (3, 256, 256) Second RGB image, channels first
observation.state (8,) Robot state
action (7,) per timestep Robot action

The model uses one observation timestep and predicts a chunk of 16 actions, with output shape (batch_size, 16, 7).

Downloads last month
5
Safetensors
Model size
2B params
Tensor type
BF16
·
Video Preview
loading