JackLiu0406's picture
Add evaluation guide to task-77 folder
7b10b03 verified
|
Raw
History Blame Contribute Delete
10.7 kB

Running evaluation β€” BEHAVIOR-1K PiBehavior / Ο€0.5 checkpoints

How to take a checkpoint from this repo and score it in the BEHAVIOR-1K simulator. Written for task85_putting_dirty_dishes_in_sink, but everything except --task_id and the stage count applies to any checkpoint here.

Provenance of this document. Every interface below was read directly from the solution source (scripts/serve_b1k.py, src/b1k/shared/eval_b1k_wrapper.py, src/b1k/policies/). The end-to-end loop has not been executed by the author of this file β€” the machine these checkpoints were trained on has no OmniGibson install and the BEHAVIOR-1K submodule was never checked out there. Treat the data contracts as accurate and the operational steps as unverified. Anything you find that differs from this, believe your own run.


1. Architecture of the loop

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”   websocket    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ OmniGibson evaluator     β”‚  ───────────>  β”‚ serve_b1k.py               β”‚
β”‚  - steps the simulator   β”‚   obs dict     β”‚  WebsocketPolicyServer     β”‚
β”‚  - scores task success   β”‚                β”‚    β”” B1KPolicyWrapper      β”‚
β”‚                          β”‚  <───────────  β”‚        β”” Policy (JAX)      β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜   action chunk β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

Two processes. serve_b1k.py holds the model and does nothing on its own β€” it waits for observations. The evaluator owns the simulator, the episode loop and the success metric.

The policy server does not need a display. The evaluator does (Isaac Sim), so in practice they often run on different machines; the server just needs a reachable port.


2. Prerequisites

Policy-server side

git clone https://github.com/IliaLarchenko/behavior-1k-solution
cd behavior-1k-solution
git submodule update --init --recursive     # openpi AND BEHAVIOR-1K
uv sync

These checkpoints were produced by a fork with the task-embedding table widened from 50 to 100 entries. On stock upstream the parameter shapes will not match β€” see Β§5.

serve_b1k.py imports from OmniGibson even though it only serves:

from omnigibson.learning.utils.network_utils import WebsocketPolicyServer
from omnigibson.learning.datas import BehaviorLerobotDatasetMetadata

so OmniGibson must be importable on the server host too. Checking git submodule status is worthwhile: a leading - means BEHAVIOR-1K was never checked out, and you will get ModuleNotFoundError: No module named 'omnigibson'.

Evaluator side

Isaac Sim + OmniGibson per the BEHAVIOR-1K challenge instructions. That stack is the authority on how episodes are launched and scored; this document does not attempt to replace it.

Download the checkpoint

huggingface-cli download JackLiu0406/b1k-checkpoints \
  --include "task85_putting_dirty_dishes_in_sink/29000/*" \
  --local-dir ./ckpts

3. Start the policy server

export B1K_TASK_SPACE=100                    # REQUIRED β€” 100-task embedding table
export XLA_PYTHON_CLIENT_MEM_FRACTION=0.9

uv run scripts/serve_b1k.py \
  --policy.config pi_behavior_b1k_fast \
  --policy.dir   ./ckpts/task85_putting_dirty_dishes_in_sink/29000 \
  --task_id 85 \
  --port 8000

--task_id selects the task embedding. Omit it only if your evaluator puts a task_id field in the observation, which the wrapper reads per-step and will hot-swap on (_handle_task_change).

assets_base_dir must resolve so that norm_stats.json and fast_tokenizer/ are found under <assets>/IliaLarchenko/behavior_224_rgb/. The bundled assets/ directory already has that layout.

Multi-task serving

--task_checkpoint_mapping <file.json> enables CheckpointSwitcher, which swaps checkpoints per task_id at runtime. Useful when evaluating several single-task models in one session.


4. The data contract

What the evaluator sends

Raw OmniGibson keys β€” the wrapper does the conversion, so send them unmodified:

key content
robot_r1::proprio full proprioception vector
robot_r1::robot_r1:zed_link:Camera:0::rgb head camera, RGB(A)
robot_r1::robot_r1:left_realsense_link:Camera:0::rgb left wrist
robot_r1::robot_r1:right_realsense_link:Camera:0::rgb right wrist
task_id optional, int; overrides --task_id per step

Only the first three channels of each image are used ([..., :3]), so RGBA is fine.

What the wrapper does with it

  1. Images β†’ resize_with_pad to the model's square input (aspect preserved, padded).

  2. State β†’ extract_state_from_proprio, which slices by PROPRIOCEPTION_INDICES["R1Pro"]:

    field dims
    base_qvel 3
    trunk_qpos 4
    arm_left_qpos 7
    arm_right_qpos 7
    left gripper (both fingers summed, normalised to [-1,1]) 1
    right gripper 1
    total 23

    The model's state vector is 32-wide; the remainder is padding.

  3. Conditioning β†’ the wrapper builds this itself, you do not supply it:

    batch["tokenized_prompt"]      = np.array([task_id, current_stage], np.int32)
    batch["tokenized_prompt_mask"] = np.array([True, True], bool)
    batch["subtask_state"]         = np.array(current_stage, np.int32)
    

    Any prompt string is deleted before inference. This model is not language-conditioned β€” conditioning is two integer embedding lookups.

What comes back

The model emits an action chunk [30, 32]. The wrapper slices it to actions[:, :23] β€” the same 23 controlled DOF as the state vector. The extra dims are padding and are discarded.


5. Task and stage indexing β€” the thing most likely to break

These are 2026-only activities, absent from the upstream 50-task table. The checkpoints were trained with the table expanded to 100 tasks.

B1K_TASK_SPACE=100 is mandatory. A 50-task model fails on a shape mismatch at load.

checkpoint task_id stages stage-embedding rows
task85_putting_dirty_dishes_in_sink 85 14 962–975
task77_installing_a_modem 77 5 888–892

Stage counts follow clip(ceil(avg_episode_length / 900), 5, 15). The wrapper clamps predictions to TASK_NUM_STAGES[task_id] - 1, so a wrong stage count silently caps progress rather than erroring β€” worth checking if a policy stalls late in an episode.


6. Execution behaviour you should understand before reading results

The wrapper is not a thin shim; it materially shapes rollout behaviour.

Chunk scheduling

parameter default meaning
--actions_to_execute 26 actions consumed per inference
--actions_to_keep 4 overlap retained between chunks
--execute_in_n_steps 20 env steps per executed chunk
--num_steps 20 flow-matching sampling steps

Re-inference happens when action_index >= execute_in_n_steps. Because execute_in_n_steps (20) < actions_to_execute (26), the chunk is compressed via cubic spline interpolation (_interpolate_actions) before execution. There is also rolling inpainting: the tail of the previous chunk is fed back as initial_actions.

Stage advancement by majority vote

parameter default
--history_len 3
--votes_to_promote 2

The stage head's logits are argmaxed each step and pushed into a 3-deep history. Once full, the wrapper counts votes for current_stage + 1; 2 of 3 promotes. It can also skip ahead (next_stage + 1) or fall back (current_stage - 1). Stage never exceeds TASK_NUM_STAGES[task_id] - 1.

Eval tricks β€” ⚠️ on by default

--apply_eval_tricks (default True) runs apply_correction_rules(task_id, stage, state, actions), which can rewrite actions and force a stage correction, plus gripper-variation checks.

For a clean measurement of the policy itself, set --apply_eval_tricks False. Leaving it on is legitimate for a leaderboard number, but it is not a measurement of the network alone, and results with and without are not comparable.


7. Sanity check without the simulator

If you only want to know whether a checkpoint loads and predicts sensibly, there is a forward-only gate that needs no Isaac Sim:

B1K_CKPT=./ckpts/task85_putting_dirty_dishes_in_sink/29000 \
B1K_TASK_SPACE=100 \
XLA_PYTHON_CLIENT_MEM_FRACTION=0.2 \
  uv run scripts/eval_gate_2026.py

It measures flow-matching action loss on recorded demos. A low value means the camera/state/task-index wiring is right; a high value means a pipeline bug. It is not a success rate and cannot substitute for a rollout.

For reference, training-set losses at the end of fine-tuning:

checkpoint action_loss
task 85 0.0342
task 77 0.0081

Do not compare across tasks β€” task 85 episodes are ~12,100 frames against task 77's ~2,413.


8. Troubleshooting

symptom cause
ModuleNotFoundError: omnigibson BEHAVIOR-1K submodule not checked out (git submodule status shows a leading -)
shape mismatch on task_embeddings B1K_TASK_SPACE=100 not set
norm_stats.json not found assets_base_dir doesn't resolve to <assets>/IliaLarchenko/behavior_224_rgb/
policy acts randomly wrong --task_id; only the trained task has a meaningful embedding
stalls near the end of an episode stage count wrong for the task β€” predictions are clamped, not errored
every episode identical prompt is ignored by design; conditioning is task_id + stage only

9. Known limitations of these checkpoints

  • Single-task. Only the named task was fine-tuned. Other task_id values are either inherited from the 50-task meta checkpoint (0–49) or still at random initialisation (50–99). Do not expect meaningful behaviour from them.
  • Inference only. train_state/ is not published β€” you cannot resume training.
  • Base-velocity frame. Trained on 2026 demos after the upstream fix (commit e6c9756) moving base_qvel from world to robot frame. A checkpoint trained on the older data saw a different distribution for state dims 0:3.
  • Reported accuracies are training-set numbers, not generalisation estimates.