Instructions to use Dimios45/cupstack_bspline_act_v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Dimios45/cupstack_bspline_act_v2 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| library_name: lerobot | |
| pipeline_tag: robotics | |
| tags: | |
| - robotics | |
| - lerobot | |
| - imitation-learning | |
| - act | |
| - b-spline | |
| - so101 | |
| - bspline_act | |
| datasets: | |
| - aryankakad/CUPSTACKING | |
| model_name: cupstack_bspline_act_v2 | |
| # cupstack_bspline_act_v2 | |
| An [ACT](https://huggingface.co/papers/2304.13705) policy that predicts **B-spline | |
| trajectory segments** instead of discrete action chunks, trained on SO-101 cup stacking. | |
| This is the "Reg.+BSP" variant from | |
| [B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations](https://arxiv.org/abs/2607.09648). | |
| The practical consequence: **you can run this checkpoint faster without retraining it.** | |
| The policy outputs a curve, so executing `a(n·t)` replays the same trajectory geometry | |
| n times faster. Speed-up is an inference flag, and it is exact — verified to | |
| 0.000e+00 deviation at 2× and 4×. | |
| This supersedes `cupstack_bspline_act`. Two things changed: idle frames were trimmed | |
| from the demonstrations, and **5 episodes were held out for validation**, so for the | |
| first time there are held-out numbers for this task. | |
| > [!IMPORTANT] | |
| > **Read the [Evaluation](#evaluation) section before trusting any number here.** | |
| > Held-out error is ~19°, against 2.1° on training data. The policy is clearly | |
| > reading the scene — it beats the inter-demonstration spread by 42% — but it has | |
| > **never been run on hardware**, and task success rate is unknown. Keep the e-stop | |
| > within reach. | |
| ## Download | |
| ```bash | |
| hf download Dimios45/cupstack_bspline_act_v2 --local-dir ./cupstack_bspline_act_v2 | |
| ``` | |
| The repository root holds the recommended checkpoint, so `--policy.path` also accepts | |
| the repo id directly. Individual checkpoints live under `checkpoints/<step>/`. | |
| ## Run it | |
| ```bash | |
| lerobot-rollout \ | |
| --strategy.type=base \ | |
| --policy.path=./cupstack_bspline_act_v2/checkpoints/040000/pretrained_model \ | |
| --policy.speed_up=1.0 \ | |
| --device=cuda \ | |
| --robot.type=so101_follower \ | |
| --robot.port=/dev/ttyACM0 --robot.id=FOLLOWER \ | |
| --robot.cameras='{front: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30}, gripper: {type: opencv, index_or_path: 4, width: 640, height: 480, fps: 30}}' \ | |
| --task="Cup stacking" \ | |
| --fps=30 --duration=60 | |
| ``` | |
| `--policy.path` takes either a Hub repo id (`namespace/name`) or a **local directory | |
| containing `config.json`** — which is the `pretrained_model/` subdirectory, not the | |
| step directory above it. There is no `--policy.subfolder` flag. | |
| To try the speed-up, raise `--policy.speed_up` to 2.0 or 4.0. Start at 1.0 and step up | |
| only once the task succeeds. Retiming does not make the arm stronger — past the | |
| low-level controller's tracking limit it overshoots rather than stopping. The paper | |
| reaches 4× on cube picking but only 2× on speed stacking, and cup stacking is the same | |
| precise-placement regime. | |
| If the control loop can't keep up, drop `--display_data` (Rerun can saturate its | |
| channel at two 640×480 streams) rather than assuming the policy is too slow — a | |
| B-spline policy runs the network only when a segment is exhausted, roughly every | |
| 0.76 s, so most control ticks are just spline evaluation. | |
| ### Camera naming is not optional | |
| The `--robot.cameras` dict keys must be **`front`** and **`gripper`**, lowercase: | |
| | Key | Shape | | |
| |---|---| | |
| | `observation.images.front` | `(3, 480, 640)` | | |
| | `observation.images.gripper` | `(3, 480, 640)` | | |
| | `observation.state` | `(6,)` | | |
| Joint order is `shoulder_pan, shoulder_lift, elbow_flex, wrist_flex, wrist_roll, gripper`, | |
| in degrees. Swapping the two cameras produces confident but wrong behaviour rather | |
| than an error. | |
| ### Requirements | |
| This policy type is not in upstream LeRobot. You need a checkout containing the | |
| `bspline_act` policy, installed with the scipy extra: | |
| ```bash | |
| uv pip install -e ".[feetech,bspline]" | |
| ``` | |
| `scipy>=1.15` is required for `interpolate.generate_knots`. | |
| Unnormalization statistics are stored as buffers inside `model.safetensors`, so each | |
| checkpoint is self-contained — you do not need the training dataset on the robot machine. | |
| ## Checkpoints | |
| | Step | Epoch | Train error | **Held-out error** | Eval loss | | | |
| |---|---|---|---|---|---| | |
| | 10000 | 17 | 6.56° | 19.20° | 0.3074 | | | |
| | 20000 | 33 | 4.95° | 19.36° | 0.3072 | | | |
| | 30000 | 50 | 3.89° | 19.33° | 0.3060 | | | |
| | **40000** | 67 | 2.63° | **19.06°** | 0.3084 | **recommended** | | |
| | 45000 | 75 | 2.49° | **18.72°** | 0.3005 | best, within noise of 40k | | |
| | 50000 | 83 | 2.28° | 19.05° | 0.3088 | | | |
| | 60000 | 100 | 2.08° | 19.22° | 0.3121 | final | | |
| All twelve checkpoints of the run span 18.72–19.36° held-out (mean 19.03, sd 0.22), so | |
| **every checkpoint here performs the same on unseen data**. 45000 is nominally best and | |
| is included for that reason, but its margin over 40000 is a quarter of the run's own | |
| noise. Pick 40000 unless you have a reason not to. | |
| Error is mean absolute joint error in degrees, worst joint, over each predicted | |
| segment's horizon (~0.76 s), measured by decoding the spline and comparing against | |
| ground-truth actions. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | Steps | 60,000 | | |
| | Batch size | 32 | | |
| | Epochs | 100 | | |
| | Dataset | [`aryankakad/CUPSTACKING`](https://huggingface.co/datasets/aryankakad/CUPSTACKING) → idle-trimmed → B-spline converted | | |
| | Episodes | **45 train / 5 held out** | | |
| | Frames | 20,956 (from 24,044; 12.9% boundary idle removed) | | |
| | Optimizer | AdamW, lr 2e-5 (backbone 1e-5), wd 1e-4 | | |
| | Precision | bf16 autocast + `channels_last` | | |
| | Hardware | 1× RTX 4090, 3 h 16 min | | |
| | Seed | 1000 | | |
| | Augmentation | none | | |
| ### Preprocessing | |
| Leading and trailing motionless frames were removed from every episode — 77 s of | |
| lead-in and 26 s of trailing idle across 50 episodes. Interior pauses were kept | |
| (only 0.3% of frames, and they are real behaviour). | |
| Motion is measured as peak-to-peak joint range over a ±0.5 s window rather than | |
| frame-to-frame delta. A single jittery sample inside a dead stretch defeats a delta | |
| test, which silently leaves the idle in. | |
| Leading idle matters more than it looks: it maps the opening observation to | |
| "don't move", and since executing a stay-still segment doesn't change the scene, the | |
| policy can re-predict it indefinitely and never start. | |
| ## Architecture | |
| Standard ACT with one change: the decoder emits a B-spline **parameter matrix** rather | |
| than an action chunk. | |
| ``` | |
| (n_knots, 1 + action_dim) = (16, 7) = 112 values per prediction | |
| column 0 knot vector, in source-frame units, 0 = "now" | |
| columns 1: control points, one per joint | |
| ``` | |
| | | | | |
| |---|---| | |
| | Params | 51.6 M | | |
| | Vision backbone | ResNet18 (ImageNet init, fine-tuned) | | |
| | dim_model / chunk_size | 512 / 16 | | |
| | VAE | enabled, kl_weight 10.0 | | |
| | B-spline degree | 3 (cubic, C² continuous) | | |
| | bspline_chunk_size | 10 | | |
| | Fitting tolerance ε | 0.2 (degrees) | | |
| | Max knot span | 6 frames | | |
| Knot spacing is fitted adaptively, so a fixed 16 rows covers a **variable** time | |
| horizon — 0.50 s to 1.03 s per segment, 0.76 s mean. The network runs only when a | |
| segment is exhausted, which decouples policy rate from control rate. | |
| ### On the fitting tolerance | |
| ε = 0.2 is not the paper's value. The paper uses 0.002 for metre-scale end-effector | |
| actions; SO-101 stores joint targets in **degrees** (~±120), roughly 100× larger. | |
| At ε = 0.002 compression is 1.06× — one knot per frame, which defeats the | |
| representation. At 0.2 the fit compresses 2.41× and reconstructs to 0.187° p99, well | |
| below servo resolution. | |
| `max_knot_span` caps how wide a single knot interval may be, bounding open-loop | |
| exposure at `chunk_size × 6` frames. On this dataset it is nearly a no-op (uncapped | |
| p95 span is already 1.17 s); it matters on recordings with long pauses, where uncapped | |
| fitting produced 18-second segments. | |
| ## Evaluation | |
| **Offline only. No hardware evaluation has been performed.** | |
| | | Train (45 eps) | **Held-out (5 eps)** | | |
| |---|---|---| | |
| | Best checkpoint | 2.49° | **18.72°** | | |
| | Recommended (40k) | 2.63° | 19.06° | | |
| Two reference points for the held-out number: | |
| - **Inter-demonstration spread: 32.27°** — how far apart two different human | |
| demonstrations of this task are, phase-aligned. At 18.7° the policy predicts an | |
| unseen episode's trajectory substantially better than another demonstration would. | |
| It is reading the scene, not replaying an average. | |
| - **Hold-still baseline: 17.59°** — emitting the current action for the next 0.76 s. | |
| The policy scores near this, but the comparison flatters a do-nothing policy: over a | |
| short horizon on a slow task, "don't move" is a decent *predictor* while being a | |
| useless *policy*. Do not read this as "no better than nothing". | |
| **Held-out performance was fixed after step ~3,000.** Across twelve checkpoints and | |
| 55,000 steps it moved 0.6°, while training error fell from 8.00° to 2.08°. Everything | |
| after the first few thousand steps went into fitting the training episodes harder. | |
| > The predecessor model card reported "2.10° open-loop error at 100k" without a | |
| > validation split. This run reproduces that figure (2.08° at 60k) and shows the | |
| > held-out number underneath it is ~19°. That was a training-data measurement, not a | |
| > generalization result. | |
| Temporal rescaling is **exact**: on a real predicted segment, `a(2t)` and `a(4t)` | |
| reproduce the 1× samples to 0.000e+00, consuming the segment in 24 / 12 / 6 control | |
| ticks. | |
| ## Limitations | |
| - **No hardware evaluation.** Task success rate is unknown. | |
| - **Held-out split is the last 5 episodes**, not a random sample. Episodes were largely | |
| cut from a continuous teleoperation stream (40 of 49 boundaries are contiguous to | |
| under 1°), so adjacent episodes are correlated and an interleaved split would leak. | |
| The tail split avoids leakage at the cost of possible session drift. | |
| - **Single task, single scene.** No robustness to lighting, camera placement, or cup | |
| position changes should be assumed. | |
| - **Speed-up has a hardware ceiling.** Expect degradation near 2× on this task. | |
| - **Camera assignment is silent when wrong.** Swapped views degrade behaviour without | |
| raising an error. | |
| ## Citation | |
| ```bibtex | |
| @article{han2026b, | |
| title={B-spline Policy: Accelerating Manipulation Policies via B-spline Action Representations}, | |
| author={Han, Xiaoshen and Xiong, Haoyu and Chen, Haonan and Liu, Chaoqi and | |
| Torralba, Antonio and Zhu, Yuke and Du, Yilun}, | |
| journal={arXiv preprint arXiv:2607.09648}, | |
| year={2026} | |
| } | |
| ``` | |