Robotics
LeRobot
Safetensors
smolvla
stack-cups
so101
File size: 7,415 Bytes
dc9b5a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86e2409
dc9b5a7
 
 
 
 
 
86e2409
dc9b5a7
86e2409
 
dc9b5a7
 
 
 
 
 
 
 
 
 
 
86e2409
 
 
 
dc9b5a7
86e2409
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
dc9b5a7
 
 
 
 
 
 
86e2409
 
 
 
 
 
dc9b5a7
 
 
 
 
 
 
86e2409
dc9b5a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86e2409
 
dc9b5a7
 
 
86e2409
dc9b5a7
86e2409
 
 
 
 
dc9b5a7
86e2409
 
dc9b5a7
 
 
 
86e2409
dc9b5a7
86e2409
dc9b5a7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
86e2409
dc9b5a7
86e2409
 
 
dc9b5a7
 
 
 
 
 
 
 
 
 
 
 
 
 
86e2409
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
---
base_model: lerobot/smolvla_base
datasets: io-intelligence/so101_stack_cups
library_name: lerobot
license: apache-2.0
model_name: smolvla
pipeline_tag: robotics
tags:
- robotics
- smolvla
- stack-cups
- lerobot
- so101
---

# Model Card for smolvla

<!-- Provide a quick summary of what the model is/does. -->

[SmolVLA](https://huggingface.co/papers/2506.01844) is a compact, efficient vision-language-action model that achieves competitive performance at reduced computational costs and can be deployed on consumer-grade hardware.

Fine-tuned on SO-101 for the task **"Stack the cups"**.

<p align="center">
  <img src="https://cdn-uploads.huggingface.co/production/uploads/640e21ef3c82bd463ee5a76d/aooU0a3DMtYmy_1IWMaIM.png" alt="smolvla architecture" width="85%"/>
</p>

<p align="center">
  <video src="https://huggingface.co/io-intelligence/smolvla_so101_stack_cups/resolve/main/inference_demo.mp4" controls width="70%"></video>
</p>

<p align="center"><em>Real-robot inference demo (also available as <code>inference_demo.mp4</code> in this repo).</em></p>

This policy has been trained and pushed to the Hub using [LeRobot](https://github.com/huggingface/lerobot).

Learn how to train and run it in the [LeRobot smolvla guide](https://huggingface.co/docs/lerobot/main/en/smolvla), or browse the [full documentation](https://huggingface.co/docs/lerobot/index).

---

## Model Details

- **License:** apache-2.0
- **Fine-tuned from:** [lerobot/smolvla_base](https://huggingface.co/lerobot/smolvla_base)
- **Robot type:** `so101_follower` (SO-101)
- **Cameras (physical → policy):** see table below

### Camera mapping

The dataset records views as `front` / `top` / `wrist`. During SmolVLA training they are renamed to `camera1` / `camera2` / `camera3`. At inference you should keep the **physical** names on the robot and pass the same `rename_map`.

| Physical camera | Mount / role | Policy feature after rename |
| --- | --- | --- |
| `front` | Front-facing view of the workspace (Astra) | `observation.images.camera1` |
| `top` | Top-down / overhead view (Astra) | `observation.images.camera2` |
| `wrist` | Wrist / gripper camera | `observation.images.camera3` |

```json
{
  "observation.images.front": "observation.images.camera1",
  "observation.images.top": "observation.images.camera2",
  "observation.images.wrist": "observation.images.camera3"
}
```

---

## Inputs & Outputs

The policy consumes these observation features and produces these action features.

**Inputs**

| Feature | Physical source | Type | Shape |
| --- | --- | --- | --- |
| `observation.state` | Joint state | STATE | `(6,)` |
| `observation.images.camera1` | `front` | VISUAL | `(3, 256, 256)` |
| `observation.images.camera2` | `top` | VISUAL | `(3, 256, 256)` |
| `observation.images.camera3` | `wrist` | VISUAL | `(3, 256, 256)` |

**Outputs**

| Feature | Type | Shape |
| --- | --- | --- |
| `action` | ACTION | `(6,)` |

---

## Training Dataset

- **Repository:** [io-intelligence/so101_stack_cups](https://huggingface.co/datasets/io-intelligence/so101_stack_cups)
- **Episodes:** 660
- **Frames:** 188035
- **Frame rate:** 30 FPS
- **Task(s):** "Stack the cups"

<a class="flex" href="https://huggingface.co/spaces/lerobot/visualize_dataset?path=io-intelligence/so101_stack_cups">
<img class="block dark:hidden" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl.svg"/>
<img class="hidden dark:block" src="https://huggingface.co/datasets/huggingface/badges/resolve/main/visualize-this-dataset-xl-dark.svg"/>
</a>

## Training Configuration

| Setting | Value |
| --- | --- |
| Training steps | 30000 |
| Batch size | 8 |
| Optimizer | adamw |
| Learning rate | 0.0001 |
| Seed | 1000 |
| LeRobot version | 0.6.1 |

---

## How to Get Started with the Model

New to LeRobot? These guides cover the full workflow:

- **[Install LeRobot](https://huggingface.co/docs/lerobot/main/en/installation)** — set up the `lerobot` package.
- **[Hardware setup](https://huggingface.co/docs/lerobot/main/en/hardware_guide)** — assemble, wire, and calibrate your robot and cameras.
- **[Record data & train a policy](https://huggingface.co/docs/lerobot/en/il_robots)** — the end-to-end imitation-learning walkthrough.
- **[CLI cheat-sheet](https://huggingface.co/docs/lerobot/main/en/cheat-sheet)** — quick reference for the `lerobot-*` commands.

### Run the policy on your robot

Use physical camera keys `front` / `top` / `wrist`, then apply the rename map so they match the policy's `camera1` / `camera2` / `camera3` features. SmolVLA works best with RTC inference.

```bash
lerobot-rollout \
  --strategy.type=base \
  --robot.type=so101_follower \
  --robot.port=<your_robot_port> \
  --robot.cameras="{ \
    front: {type: opencv, index_or_path: <front_device>, width: 640, height: 480, fps: 30}, \
    top: {type: opencv, index_or_path: <top_device>, width: 640, height: 480, fps: 30}, \
    wrist: {type: opencv, index_or_path: <wrist_device>, width: 640, height: 480, fps: 30} \
  }" \
  --policy.path=io-intelligence/smolvla_so101_stack_cups \
  --rename_map='{"observation.images.front":"observation.images.camera1","observation.images.top":"observation.images.camera2","observation.images.wrist":"observation.images.camera3"}' \
  --inference.type=rtc \
  --task="Stack the cups" \
  --duration=60
```

Replace `<your_robot_port>` and the three camera device paths with your machine values. Camera **names** must stay `front` / `top` / `wrist` (not `camera1/2/3` on the robot side).

When `--strategy.type=base` is used the script doesn't record episodes. Set `--duration=0` (or omit duration depending on your CLI) to run until Ctrl+C. For more information see the [rollout / inference docs](https://huggingface.co/docs/lerobot/main/en/inference).

### Train your own policy

This policy type is usually fine-tuned from the pretrained base model [lerobot/smolvla_base](https://huggingface.co/lerobot/smolvla_base):

```bash
lerobot-train \
  --dataset.repo_id=${HF_USER}/<dataset> \
  --policy.path=lerobot/smolvla_base \
  --output_dir=outputs/train/<policy_repo_id> \
  --job_name=lerobot_training \
  --policy.device=cuda \
  --policy.repo_id=${HF_USER}/<policy_repo_id> \
  --wandb.enable=true
```

_Writes checkpoints to `outputs/train/<policy_repo_id>/checkpoints/`._

---

## Evaluation

Real-robot inference demo of stacking cups is included as [`inference_demo.mp4`](https://huggingface.co/io-intelligence/smolvla_so101_stack_cups/resolve/main/inference_demo.mp4) (copied from the [training dataset](https://huggingface.co/datasets/io-intelligence/so101_stack_cups) root).

| Task | Notes |
| ---- | ----- |
| Stack the cups | Qualitative success demo on SO-101 (see video above) |

---

## Citation

If you use this policy, please cite the method linked in the description above, along with LeRobot:

```bibtex
@misc{cadene2024lerobot,
    author = {Cadene, Remi and Alibert, Simon and Soare, Alexander and Gallouedec, Quentin and Zouitine, Adil and Palma, Steven and Kooijmans, Pepijn and Aractingi, Michel and Shukor, Mustafa and Aubakirova, Dana and Russi, Martino and Capuano, Francesco and Pascal, Caroline and Choghari, Jade and Moss, Jess and Wolf, Thomas},
    title = {LeRobot: State-of-the-art Machine Learning for Real-World Robotics in Pytorch},
    howpublished = "\url{https://github.com/huggingface/lerobot}",
    year = {2024}
}
```