Instructions to use Bigenlight/diffusion_banana_in_pot_joint_fp16 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LeRobot
How to use Bigenlight/diffusion_banana_in_pot_joint_fp16 with LeRobot:
- Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| tags: | |
| - robotics | |
| - lerobot | |
| - diffusion-policy | |
| - imitation-learning | |
| - manipulation | |
| # Diffusion Policy β banana-in-pot (JOINT, fp16) | |
| Diffusion Policy trained on the **"put the right banana in the pot"** task (UR7e + GELLO | |
| teleoperation, 2 RGB cameras), in **JOINT** action space (6 joints + gripper), using | |
| **fp16 mixed-precision** training. | |
| - **Checkpoint:** step 80,000 (best open-loop MAE) | |
| - **Base library:** [LeRobot](https://github.com/huggingface/lerobot) 0.6.1 (pin `8a74e0a`) | |
| - **Dataset:** [`Bigenlight/banana_in_pot_lerobot_v3`](https://huggingface.co/datasets/Bigenlight/banana_in_pot_lerobot_v3) β 51 episodes / 21,524 frames / 30 fps | |
| - **fp32 sibling:** [`Bigenlight/diffusion_banana_in_pot_joint`](https://huggingface.co/Bigenlight/diffusion_banana_in_pot_joint) | |
| ## Architecture | |
| ResNet18 (ImageNet-pretrained, trained end-to-end) + SpatialSoftmax vision encoder β | |
| FiLM-conditioned 1D temporal U-Net denoiser (DDPM-100, epsilon prediction). ~278M params. | |
| Images resized to 360Γ640, crop off. `n_obs_steps=2`, `horizon=64`, `n_action_steps=32`. | |
| ## Training | |
| - **Precision:** fp16 via HF Accelerate `mixed_precision=fp16` (automatic GradScaler). | |
| - Requires a `dtype` field on `DiffusionConfig` (absent upstream at this pin); launched with | |
| `--policy.dtype=float16`. | |
| - Batch 8, 80k steps, AdamW, seed 1000, 45 train / 6 held-out episodes. | |
| - **Hardware:** single RTX A4000. **~4.48 step/s, ~4.96 h** (vs ~3.49 step/s / ~6.4 h fp32 β | |
| ~1.25Γ faster, ~22% less wall-clock). No NaN/instability. | |
| ## Open-loop evaluation (DDIM-10, held-out episodes 45β50) | |
| | | fp32 @80k | **fp16 @80k** | | |
| |---|---|---| | |
| | poseMAE (rad) | 0.0845 | **0.08268** | | |
| | gripper accuracy | 0.953 | **0.960** | | |
| **fp16 matches (marginally beats, within noise) fp32 accuracy at ~22% lower training cost.** | |
| Select the deploy checkpoint by open-loop MAE, **not** `eval_loss` (which rises during | |
| training for diffusion policies without indicating overfitting). | |
| ## Intended use & limitations | |
| Research artifact. Small single-task, single-scene, real-world (noisy) dataset of 51 | |
| success-only demonstrations; offline metrics only β no closed-loop hardware success rate | |
| measured yet. Not safety-validated for autonomous operation. | |