noahnowac's picture
README only
3baf60f verified
|
Raw
History Blame Contribute Delete
2.33 kB
# FARM UF850 Patch Policy β€” task_2 "stack the three cubes" (L1 head)
Reimplementation of **Patch Policy** (arXiv 2607.18236) as a single-task specialist for the
FARM UF850 arm: frozen **DINOv2-S/14** dense patch tokens (all 256/camera, base + wrist, no
pooling) -> 8-layer / d=512 transformer (**26.8M trainable**) -> L1-regressed 50-step action
chunk. No language conditioning (single task). Trained on 84 episodes, 10 held out.
**Training cost: ~0.5 GPU-hours** (30k steps, batch 64, one H200, ~24 it/s β€” episode subset
fully RAM-preloaded).
## Results (offline, held-out episodes β€” triage proxy ONLY, no arm rollouts yet)
| variant | held-out chunk-MAE (rad) | first-action MAE (rad) |
|---|---|---|
| **L1 head (this ckpt)** | **0.095** | **0.033 (~1.9Β°/joint)** |
| diffusion head (MLP denoiser) | 0.319 | 0.220 |
Open-loop prediction accuracy is a screen, not a verdict β€” closed-loop rollouts on the arm are
the only real test, and this model has not had them.
## Inference
```python
from infer_patch_policy import PatchPolicyRunner
r = PatchPolicyRunner("patch_policy_task2_l1.pt") # downloads DINOv2-S via torch.hub
chunk = r.predict(base_rgb, wrist_rgb, state7) # (50, 7)
```
- Inputs: two RGB frames (any resolution; resized to 224x224) + 7-d proprio
(6 joint positions rad + gripper). State is normalized internally (stats stored in ckpt).
- Output: 50-step chunk in the FARM dataset action space β€” **6 relative joint-position deltas
(rad) + 1 gripper position [0,1]**, recorded at 30 fps. Execute like the dataset semantics
(apply deltas to current joint positions; gripper is absolute). Re-plan by re-querying at
chunk end or earlier.
- Deps: torch >= 2.x, opencv, numpy, internet on first run (torch.hub DINOv2 weights, ~84MB).
## Files
- `patch_policy_task2_l1.pt` β€” EMA weights (incl. frozen ViT), config, normalization stats.
- `train_patch_policy.py` β€” full self-contained training script (data loader for lerobot-v2
FARM episodes, both heads, eval). Repro: see docstring; ~31 min on one H200.
- `infer_patch_policy.py` β€” the runner above.
## Caveats
Single-task specialist (no language, no multi-task); T=1 observation (no temporal context);
diffusion variant underperforms here (simplified MLP denoiser, not the paper's DP head);
offline-validated only.