Robotics
LeRobot
Safetensors
imitation-learning
act
b-spline
so101
bspline_act
File size: 6,975 Bytes
d2f77bd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
---
license: apache-2.0
library_name: lerobot
pipeline_tag: robotics
tags:
  - robotics
  - lerobot
  - imitation-learning
  - act
  - b-spline
  - so101
  - bspline_act
datasets:
  - aryankakad/CUPSTACKING
model_name: cupstack_bspline_act
---

# cupstack_bspline_act

An [ACT](https://huggingface.co/papers/2304.13705) policy that predicts **B-spline
trajectory segments** instead of discrete action chunks, trained on SO-101 cup
stacking. This is the "Reg.+BSP" variant from
[B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations](https://arxiv.org/abs/2607.09648).

The practical consequence: **you can run this checkpoint faster without retraining it.**
The policy outputs a curve, so executing `a(n·t)` replays the same trajectory geometry
n times faster. Speed-up is an inference flag.

> [!WARNING]
> **This checkpoint has never been evaluated on held-out data or on hardware.**
> It was trained on all 50 episodes with no validation split, so its numbers measure
> fit, not generalization. Treat first hardware runs as untested — keep the e-stop
> within reach. See [Limitations](#limitations).

## Usage

```bash
lerobot-rollout \
    --strategy.type=base \
    --policy.path=aryankakad/cupstack_bspline_act \
    --policy.speed_up=1.0 \
    --device=cuda \
    --robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=FOLLOWER \
    --robot.cameras='{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, gripper: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}}' \
    --task="stack the cups" \
    --fps=30 --duration=30 --display_data=true
```

Raise `--policy.speed_up` to 2.0 or 4.0 to execute faster. Start at 1.0 and step up
only once the task succeeds.

### Camera naming is not optional

The policy requires exactly these observation keys, so the `--robot.cameras` dict keys
must be **`front`** and **`gripper`**, lowercase:

| Key | Shape |
|---|---|
| `observation.images.front` | `(3, 480, 640)` |
| `observation.images.gripper` | `(3, 480, 640)` |
| `observation.state` | `(6,)` |

Joint order is `shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper`.
Swapping the two cameras produces confident but wrong behaviour rather than an error.

### Requirements

This policy type is not in upstream LeRobot. You need a checkout containing the
`bspline_act` policy, installed with the scipy extra:

```bash
uv pip install -e ".[feetech,bspline]"
```

`scipy>=1.15` is required for `interpolate.generate_knots`.

### No dataset needed at inference

Unnormalization statistics are stored as buffers inside `model.safetensors`, so the
checkpoint is self-contained. You do not need the training dataset on the robot machine.

## Training

| | |
|---|---|
| Steps | **100,000** |
| Batch size | 32 |
| Epochs | 133.09 |
| Dataset | [`aryankakad/CUPSTACKING`](https://huggingface.co/datasets/aryankakad/CUPSTACKING) → B-spline converted |
| Episodes / frames | 50 / 24,044 @ 30 fps |
| Optimizer | AdamW, lr 2e-5 (backbone 1e-5), wd 1e-4 |
| Precision | bf16 autocast + `channels_last` |
| Hardware | 1× RTX 4090, 5 h 24 min |
| Seed | 1000 |

Learning rate is 2e-5 rather than ACT's default 1e-5, sqrt-scaled for batch 32.

### Loss

| Step | loss | l1_loss | kld_loss |
|---|---|---|---|
| 500 | 3.027 | 0.390 | 0.264 |
| 10k | 0.166 | 0.104 | 0.007 |
| 50k | ~0.111 | 0.041 | 0.007 |
| **100k** | **0.095** | **0.026** | **0.007** |

Training loss only — there was no validation split.

## Architecture

Standard ACT with one change: the decoder emits a B-spline **parameter matrix** rather
than an action chunk.

```
(n_knots, 1 + action_dim) = (16, 7) = 112 values per prediction
   column 0   knot vector, in source-frame units, 0 = "now"
   columns 1: control points, one per joint
```

| | |
|---|---|
| Params | 52 M |
| Vision backbone | ResNet18 (ImageNet init) |
| dim_model / chunk_size | 512 / 16 |
| VAE | enabled, kl_weight 10.0 |
| B-spline degree | 3 (cubic, C² continuous) |
| bspline_chunk_size | 10 |
| Fitting tolerance ε | 0.2 (degrees) |

Knot spacing is fitted adaptively per episode, so a fixed 16 rows covers a **variable**
time horizon — 0.53 s to 1.47 s per segment on this dataset, 0.84 s mean. The network
only runs when a segment is exhausted, which decouples policy rate from control rate.

### On the fitting tolerance

ε = 0.2 is not the paper's value. The paper uses 0.002 for metre-scale end-effector
actions; SO-101 stores joint targets in **degrees** (~±120), roughly 100× larger. At
ε = 0.002 compression is 1.06× — one knot per frame, which defeats the representation.

| ε | Compression | p99 reconstruction | Segment span |
|---|---|---|---|
| 0.05 | 1.4× | 0.05° | 0.44 s |
| **0.2** | **2.7×** | **0.19°** | **0.84 s** |
| 0.5 | 4.1× | 0.48° | 1.24 s |

2.7× sits inside the paper's reported 1.12×–3.34× range.

## Evaluation

**Offline only. No hardware evaluation has been performed.**

Open-loop prediction error — decoded trajectory vs ground-truth actions, measured on
**training data**:

| Checkpoint | Mean | Median | p90 | Max |
|---|---|---|---|---|
| 20k | 5.99° | 4.12° | 8.78° | 45.2° |
| 50k | 3.34° | 2.28° | 6.53° | 30.8° |
| **100k** | **2.10°** | **1.61°** | **3.86°** | **18.6°** |

Error was still falling at 100k and 2.10° is far from zero, which argues against
outright memorization — but this is training data, so it is not evidence of
generalization.

The representation is not the bottleneck: B-spline fitting reconstructs to 0.19° p99,
so essentially all of the 2.10° is policy prediction error.

Temporal rescaling is **exact**. On a real predicted segment (0.76 s span), `a(2t)` and
`a(4t)` reproduce the 1× samples to 0.00e+00, consuming the segment in 22 / 11 / 5
control ticks.

## Limitations

- **No validation split.** All 50 episodes were used for training (`eval_steps: 0`).
  Generalization is unmeasured.
- **No hardware evaluation.** Task success rate is unknown.
- **Speed-up has a hardware ceiling.** The paper reaches 4× on cube picking but only 2×
  on speed stacking before the low-level controller loses tracking. Cup stacking is the
  same precise-placement regime, so expect degradation near 2×. Retiming does not make
  the arm faster — past the tracking limit it overshoots rather than stopping.
- **Single task, single scene.** 50 demonstrations of one cup-stacking setup; no
  robustness to lighting, camera placement, or cup position changes should be assumed.
- **Camera assignment is silent when wrong.** Swapped views degrade behaviour without
  raising an error.

## Citation

```bibtex
@article{han2026b,
  title={B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations},
  author={Han, Xiaoshen and Xiong, Haoyu and Chen, Haonan and Liu, Chaoqi and
          Torralba, Antonio and Zhu, Yuke and Du, Yilun},
  journal={arXiv preprint arXiv:2607.09648},
  year={2026}
}
```