Instructions to use aryankakad/cupstack_bspline_act with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use aryankakad/cupstack_bspline_act with LeRobot:
- Notebooks
- Google Colab
- Kaggle
File size: 6,975 Bytes
d2f77bd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 | ---
license: apache-2.0
library_name: lerobot
pipeline_tag: robotics
tags:
- robotics
- lerobot
- imitation-learning
- act
- b-spline
- so101
- bspline_act
datasets:
- aryankakad/CUPSTACKING
model_name: cupstack_bspline_act
---
# cupstack_bspline_act
An [ACT](https://huggingface.co/papers/2304.13705) policy that predicts **B-spline
trajectory segments** instead of discrete action chunks, trained on SO-101 cup
stacking. This is the "Reg.+BSP" variant from
[B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations](https://arxiv.org/abs/2607.09648).
The practical consequence: **you can run this checkpoint faster without retraining it.**
The policy outputs a curve, so executing `a(n·t)` replays the same trajectory geometry
n times faster. Speed-up is an inference flag.
> [!WARNING]
> **This checkpoint has never been evaluated on held-out data or on hardware.**
> It was trained on all 50 episodes with no validation split, so its numbers measure
> fit, not generalization. Treat first hardware runs as untested — keep the e-stop
> within reach. See [Limitations](#limitations).
## Usage
```bash
lerobot-rollout \
--strategy.type=base \
--policy.path=aryankakad/cupstack_bspline_act \
--policy.speed_up=1.0 \
--device=cuda \
--robot.type=so101_follower --robot.port=/dev/ttyACM0 --robot.id=FOLLOWER \
--robot.cameras='{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, gripper: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}}' \
--task="stack the cups" \
--fps=30 --duration=30 --display_data=true
```
Raise `--policy.speed_up` to 2.0 or 4.0 to execute faster. Start at 1.0 and step up
only once the task succeeds.
### Camera naming is not optional
The policy requires exactly these observation keys, so the `--robot.cameras` dict keys
must be **`front`** and **`gripper`**, lowercase:
| Key | Shape |
|---|---|
| `observation.images.front` | `(3, 480, 640)` |
| `observation.images.gripper` | `(3, 480, 640)` |
| `observation.state` | `(6,)` |
Joint order is `shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper`.
Swapping the two cameras produces confident but wrong behaviour rather than an error.
### Requirements
This policy type is not in upstream LeRobot. You need a checkout containing the
`bspline_act` policy, installed with the scipy extra:
```bash
uv pip install -e ".[feetech,bspline]"
```
`scipy>=1.15` is required for `interpolate.generate_knots`.
### No dataset needed at inference
Unnormalization statistics are stored as buffers inside `model.safetensors`, so the
checkpoint is self-contained. You do not need the training dataset on the robot machine.
## Training
| | |
|---|---|
| Steps | **100,000** |
| Batch size | 32 |
| Epochs | 133.09 |
| Dataset | [`aryankakad/CUPSTACKING`](https://huggingface.co/datasets/aryankakad/CUPSTACKING) → B-spline converted |
| Episodes / frames | 50 / 24,044 @ 30 fps |
| Optimizer | AdamW, lr 2e-5 (backbone 1e-5), wd 1e-4 |
| Precision | bf16 autocast + `channels_last` |
| Hardware | 1× RTX 4090, 5 h 24 min |
| Seed | 1000 |
Learning rate is 2e-5 rather than ACT's default 1e-5, sqrt-scaled for batch 32.
### Loss
| Step | loss | l1_loss | kld_loss |
|---|---|---|---|
| 500 | 3.027 | 0.390 | 0.264 |
| 10k | 0.166 | 0.104 | 0.007 |
| 50k | ~0.111 | 0.041 | 0.007 |
| **100k** | **0.095** | **0.026** | **0.007** |
Training loss only — there was no validation split.
## Architecture
Standard ACT with one change: the decoder emits a B-spline **parameter matrix** rather
than an action chunk.
```
(n_knots, 1 + action_dim) = (16, 7) = 112 values per prediction
column 0 knot vector, in source-frame units, 0 = "now"
columns 1: control points, one per joint
```
| | |
|---|---|
| Params | 52 M |
| Vision backbone | ResNet18 (ImageNet init) |
| dim_model / chunk_size | 512 / 16 |
| VAE | enabled, kl_weight 10.0 |
| B-spline degree | 3 (cubic, C² continuous) |
| bspline_chunk_size | 10 |
| Fitting tolerance ε | 0.2 (degrees) |
Knot spacing is fitted adaptively per episode, so a fixed 16 rows covers a **variable**
time horizon — 0.53 s to 1.47 s per segment on this dataset, 0.84 s mean. The network
only runs when a segment is exhausted, which decouples policy rate from control rate.
### On the fitting tolerance
ε = 0.2 is not the paper's value. The paper uses 0.002 for metre-scale end-effector
actions; SO-101 stores joint targets in **degrees** (~±120), roughly 100× larger. At
ε = 0.002 compression is 1.06× — one knot per frame, which defeats the representation.
| ε | Compression | p99 reconstruction | Segment span |
|---|---|---|---|
| 0.05 | 1.4× | 0.05° | 0.44 s |
| **0.2** | **2.7×** | **0.19°** | **0.84 s** |
| 0.5 | 4.1× | 0.48° | 1.24 s |
2.7× sits inside the paper's reported 1.12×–3.34× range.
## Evaluation
**Offline only. No hardware evaluation has been performed.**
Open-loop prediction error — decoded trajectory vs ground-truth actions, measured on
**training data**:
| Checkpoint | Mean | Median | p90 | Max |
|---|---|---|---|---|
| 20k | 5.99° | 4.12° | 8.78° | 45.2° |
| 50k | 3.34° | 2.28° | 6.53° | 30.8° |
| **100k** | **2.10°** | **1.61°** | **3.86°** | **18.6°** |
Error was still falling at 100k and 2.10° is far from zero, which argues against
outright memorization — but this is training data, so it is not evidence of
generalization.
The representation is not the bottleneck: B-spline fitting reconstructs to 0.19° p99,
so essentially all of the 2.10° is policy prediction error.
Temporal rescaling is **exact**. On a real predicted segment (0.76 s span), `a(2t)` and
`a(4t)` reproduce the 1× samples to 0.00e+00, consuming the segment in 22 / 11 / 5
control ticks.
## Limitations
- **No validation split.** All 50 episodes were used for training (`eval_steps: 0`).
Generalization is unmeasured.
- **No hardware evaluation.** Task success rate is unknown.
- **Speed-up has a hardware ceiling.** The paper reaches 4× on cube picking but only 2×
on speed stacking before the low-level controller loses tracking. Cup stacking is the
same precise-placement regime, so expect degradation near 2×. Retiming does not make
the arm faster — past the tracking limit it overshoots rather than stopping.
- **Single task, single scene.** 50 demonstrations of one cup-stacking setup; no
robustness to lighting, camera placement, or cup position changes should be assumed.
- **Camera assignment is silent when wrong.** Swapped views degrade behaviour without
raising an error.
## Citation
```bibtex
@article{han2026b,
title={B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations},
author={Han, Xiaoshen and Xiong, Haoyu and Chen, Haonan and Liu, Chaoqi and
Torralba, Antonio and Zhu, Yuke and Du, Yilun},
journal={arXiv preprint arXiv:2607.09648},
year={2026}
}
```
|