File size: 6,831 Bytes
5ab1067
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b9c7e28
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5ab1067
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
---
license: mit
library_name: diffusion-policy
tags:
  - robotics
  - imitation-learning
  - diffusion-policy
  - manipulation
  - yam
datasets:
  - Dimios45/yam_duster_in_box
pipeline_tag: robotics
---

# YAM Duster-in-Box — Diffusion Policy (baseline)

A UNet [Diffusion Policy](https://diffusion-policy.cs.columbia.edu/) trained on
[`Dimios45/yam_duster_in_box`](https://huggingface.co/datasets/Dimios45/yam_duster_in_box)
for a single-arm I2RT **YAM** robot: *"pick up the duster and put the duster in the box."*

This is the **baseline** for the B-spline Policy comparison. It was trained on byte-identical
data with the same backbone, optimizer, schedule, and epoch count as
[`Dimios45/yam-duster-bspline-dp`](https://huggingface.co/Dimios45/yam-duster-bspline-dp) —
the only difference is the action representation. Use it to measure what the B-spline
parameterization actually buys in success rate and completion time.

## Action space

Joint-space targets, **not** end-effector poses:

```
[right_joint_1..6 (radians), right_gripper (0 = open, 1 = closed)]
```

No IK is involved at deploy. The policy predicts a plain `(16, 7)` trajectory of future joint
targets on a fixed 10 Hz grid and executes the first 8 (`n_action_steps: 8`) before replanning —
roughly 0.8 s per prediction, versus the ~1.1 s continuous, rescalable curve the B-spline
variant emits.

## Observations

| key | shape | notes |
| --- | --- | --- |
| `top_image` | `(3, 128, 128)` | RGB, resized from 640×480 (plain squash, no crop) |
| `wrist_image` | `(3, 128, 128)` | RGB, right wrist camera |
| `joint_pos` | `(7,)` | measured joints + gripper |

Two observation steps. Images normalized to [0,1], random crop to 116×116 in training.

> Two naming constraints in this codebase: the lowdim key must be named `joint_pos`
> (`get_normalizer` only accepts keys containing `pos`/`quat`/`qpos`), and the task config must
> set `abs_action: True`. At 7D→7D `abs_action` is a no-op for the *data*, but it also selects
> the action normalizer — `False` installs an identity normalizer, and joints 2/3 (up to 2.77 rad)
> then fall outside the sampler's `clip_sample` range of [-1, 1], silently breaking training.

## Files

| file | size | use |
| --- | --- | --- |
| `deploy_ema.ckpt` | 426 MB | **Inference.** EMA weights only. |
| `epoch0600_full.ckpt` | 1.5 GB | `model` + `ema_model` + `optimizer`, for resuming or fine-tuning. |

Both embed the Hydra config (pickled with `dill`); `diffusion_policy` must be importable on load.

## Training

| | |
| --- | --- |
| base | [B-spline-policy/bspline-policy](https://github.com/B-spline-policy/bspline-policy) UNet DP |
| data | 50 episodes, 37,986 frames @ 30 Hz → resampled to **10 Hz** (12,677 steps) |
| hardware | 1× RTX 4090, 2 h 00 m, ~16 it/s |
| epochs / batch | 601 / 64 |
| optimizer | AdamW, lr 1e-4, cosine, 500 warmup steps, EMA |
| params | 66.9M diffusion + 22.4M vision (ResNet18 ×2) |
| scheduler | DDIM, 100 train timesteps, 16 inference steps, epsilon prediction |
| horizon / obs / action steps | 16 / 2 / 8 |
| final train_loss | 0.001 |

## Measured behavior

Over 200 held-out samples, predicted chunks vs the recorded demos:

| metric | this model | B-spline variant |
| --- | --- | --- |
| open-loop arm error | median **0.48°**, p90 1.20° | median 0.81°, p90 1.68° |
| gripper predicted range | **[−0.012, 1.000]** | [−0.31, 1.31] |
| peak arm velocity | 46 °/s p95, 110 °/s max | 48 °/s p95, 112 °/s max |
| inference latency | 44 ms (RTX 4090) | 46 ms |

Velocity was sampled at this model's native 10 Hz action rate versus 100 Hz for the B-spline
curve, so those two rows are not strictly equivalent — the 10 Hz figure cannot see sub-100 ms
spikes.

Because DP predicts the trajectory directly rather than control points, its outputs stay inside
the demonstrated range, so it does **not** need the gripper clamp the B-spline variant requires.

## Deployment

Environment setup, camera configuration, arm bring-up, and the five repo patches for joint-space
actions are identical to the
[B-spline model card](https://huggingface.co/Dimios45/yam-duster-bspline-dp) — follow steps 1–4
there. Those patches are **applied and verified** in the working tree these models were trained
from, but are **not** in upstream `B-spline-policy/bspline-policy`; apply them yourself if you
start from upstream. The gripper clamp in `decode_action_vector` is unnecessary for this model
(see below) but harmless and worth keeping.

They are needed because the upstream rollout path decodes **end-effector** actions (its iPhone
teleop records EE poses), while this model predicts **joints** — a difference in what was
recorded, not in the LeRobot file format. Two things carry over unchanged: a missing `top_image`
does not raise (`policy_local_bspline.py:625-629` feeds a black frame, so a dead camera yields a
half-blind policy rather than an error), and image resolution needs no change because the rollout
resizes from the checkpoint's `shape_meta`. The top camera is an Intel RealSense; use
`TOP_CAMERA_TYPE = 'realsense'` so the RGB pipeline matches training.

The gripper-convention check in the B-spline card applies here too — verify it before the first
rollout, or the policy will open to grasp and close to release.

```bash
hf download Dimios45/yam-duster-dp deploy_ema.ckpt --local-dir ./ckpt

sudo ip link set can_follower_r up type can bitrate 1000000
python real_env/yam_teleop/yam_server.py --channel can_follower_r

python real_env/yam_teleop/rollout_local_policy.py \
  --env yam --policy dp \
  --ckpt-path ./ckpt/deploy_ema.ckpt \
  --diffusion-policy-dir $PWD/diffusion_policy \
  --control-freq 100 --data-freq 10 \
  --save --output-dir data/local_policy_rollouts
```

Note `--policy dp` (not `bspline`), and that `--origin-time-scale` / `--predict-before-end` /
`--speed-up-times` do not apply — this model emits a fixed-rate action grid with no temporal
rescaling. That is precisely the capability the B-spline variant adds.

## Reproducing

```bash
python tools/lerobot_v3_to_robomimic.py \
  --repo-id Dimios45/yam_duster_in_box \
  --output diffusion_policy/data/yam_duster_in_box.hdf5 \
  --target-fps 10 --image-size 128

cd diffusion_policy && python train.py \
  --config-name=yam_duster_unet_dp \
  hydra.run.dir=../outputs/yam_duster_dp_full \
  training.resume=false logging.mode=offline \
  checkpoint.topk.k=601 dataloader.persistent_workers=True
```

## Citation

```bibtex
@inproceedings{chi2023diffusionpolicy,
  title={Diffusion Policy: Visuomotor Policy Learning via Action Diffusion},
  author={Chi, Cheng and Feng, Siyuan and Du, Yilun and Xu, Zhenjia and
          Cousineau, Eric and Burchfiel, Benjamin and Song, Shuran},
  booktitle={Proceedings of Robotics: Science and Systems (RSS)},
  year={2023}
}
```