--- license: apache-2.0 language: - en library_name: lerobot pipeline_tag: robotics tags: - lerobot - robotics - manipulation - pi0 - tactile - bimanual --- # pi05_bi — task2 (dish washing), step 8000 openpi `pi05_bi` checkpoint for a bimanual dish-and-sponge task with tactile inputs. Step 8000 (1.24 epochs), held-out validation loss **0.0535** (best through step 12000 is step 12000 at 0.0534). ## Task Single unified instruction for every episode: > Use the left hand to pick up the dish, and then use the right hand to pick up > the sponge to brush the dish. Finally, put all things back. Only `task2_01` shipped this string; `task2_02` and `task2_03` both carried the placeholder `"perform manipulation task"` in their `meta/tasks.jsonl`. Since `prompt_from_task=True` feeds that string straight to the model, all three sources were forced onto the instruction above at merge time -- otherwise 76% of the episodes (723 of 948) would have trained against an uninformative prompt. A trailing space in the original string was stripped. ## Data | source | episodes | frames | |---|---:|---:| | `KaiyueChen/task2_01` | 225 | 222,531 | | `KaiyueChen/task2_02` | 298 | 291,661 | | `KaiyueChen/task2_03` | 425 | 402,110 | | **merged** | **948** | **916,302** | LeRobot v2.1, 30 fps, `robot_type=bimanual`, images embedded in the parquet files (`total_videos=0`). Six camera streams: `camera0`, `camera1`, and four tactile sensors (`tactile_left_0/1`, `tactile_right_0/1`). Mean episode length ~966 frames. ### Split Episodes are held out **per source repo** (10%, seed 42) so the held-out set keeps the same source mix as train: | split | episodes | held out from | |---|---:|---| | `train` | 854 | | | `val_seen` (subset of train) | 94 | | | `val_unseen` (held out) | 94 | 22 / 30 / 42 from sources 01 / 02 / 03 | Normalization statistics (quantile q01/q99) are computed over the `train` split only. ## Training | | | |---|---| | config | `pi05_bi` | | hardware | 2 x A100-80GB, FSDP | | batch size | 128 | | this checkpoint | step 8000 (~1.24 epoch; 1 epoch = 6,444 steps) | | planned length | 20,000 steps | | lr | cosine decay, 1,000 warmup steps: peak 2.5e-5 -> 2.5e-6 over 30,000 steps | | | (`CosineDecaySchedule` defaults -- `pi05_bi` does not override `lr_schedule`; the `peak_lr=2e-4 / decay_steps=100000` block in `config.py` is referenced only by `pi05_single*`) | | LoRA | rank 16 on the LLM, rank 32 on the action expert | | vision tower | **fully fine-tuned** -- the freeze filter matches only `.*llm.*` | ## Validation curve Flow-matching loss, 20 batches per split, evaluated on the same leading batches each time so successive points are comparable. | step | train | val_seen | val_unseen | gap | |---:|---:|---:|---:|---:| | 0 | 0.6399 | 0.6621 | 0.6508 | -0.0113 | | 2000 | 0.0535 | 0.0671 | 0.0662 | -0.0009 | | 4000 | 0.0470 | 0.0574 | 0.0573 | -0.0001 | | 6000 | 0.0448 | 0.0565 | 0.0566 | +0.0001 | | **8000** | 0.0430 | 0.0513 | **0.0535** | +0.0022 | | 10000 | 0.0420 | 0.0514 | 0.0552 | +0.0038 | | 12000 | 0.0421 | 0.0493 | 0.0534 | +0.0041 | `val_unseen` fell steeply through step 8000 (0.0535), read 0.0552 at step 10000, then returned to 0.0534 at step 12000 -- the step-10000 reading was noise, not a turn. Improvement has nonetheless nearly stopped: the 4,000 steps from 8000 to 12000 bought 0.0001. Over that same span `val_seen` improved 4% (0.0513 -> 0.0493) and the gap widened from 0.0022 to 0.0041. The gap is still small in absolute terms. The sibling `two_tubes_0102` run had a gap of 0.0135 at its step 12000 -- more than 3x larger -- and its `val_unseen` had been flat for 6,000 steps by then. task2 is overfitting far less at the same step count, which is why training continued past this checkpoint. The early-step `gap` is slightly **negative** (-0.0009 at step 2000). That is not evidence of good generalization: each validation pass covers only ~2,560 frames, which at ~966 frames per episode is about 2-3 episodes per split, so the sign of the gap early on is dominated by which particular episodes landed in each split. Only the `val_unseen` absolute trend is reliable, since the same episodes and the same rng are used at every evaluation. Note that the training loss is measured on **augmented** images (random crop to 95%, +-5 deg rotation, colour jitter, applied to all six streams including the four tactile ones) while validation runs with `train=False`, which skips augmentation. The two numbers are therefore not directly comparable; on this run `val_seen` still sits above the training loss at every step, unlike the `two_tubes_0102` run where it dropped below after ~2.5 epochs of memorization. ## Contents ``` checkpoint/ params/ # inference weights train_state/ # optimizer state, for resuming assets/task2_all/ norm_stats.json # computed over the train split only ```