VAM-Cross MimicVideo World2Action decoder

This private repository contains the World2Action decoder checkpoint from iteration 2374 of w2a_widowx_target_bridge_videolora_iter1060. The run stopped because of walltime. The newest complete model, optimizer, scheduler, and trainer checkpoint set was verified before selecting the uploaded model weight.

Required frozen inputs

  • MimicVideo commit: e3355dbc93132b576c02f920a59b4fc18a4f5906
  • Bridge checkpoint repository: jonpai/mimic-video
  • Bridge checkpoint revision: f28339034831e3c2374be075e622e1ff38ebe0f8
  • Bridge action initialization: action_decoder/w2a_bridge_v2w_bridge_lora_rank256_lr1.778e-04_bsz64_iter_000070043_fused_lr1.000e-04_layer20_bsz256_iter_000014112.pt
  • Frozen Video LoRA: dreamdifferent/vam-cross-target-widowx250-native-2cam-video-lora@0adf9bec12452cc113952b7d642153cf6e60cfb8

Action and data contract

  • Dataset: dreamdifferent/vam-cross-target-widowx250-native@44c28a3bada463f06845daee8a1dd87607d32956
  • Episodes/frames: 130 / 54301
  • Cameras: observation.images.corner_cam, observation.images.front_cam
  • Target: 15 achieved-EE/gripper actions at 5 Hz
  • Pose target: relative_to_current_achieved_pose in robot_base
  • Rotation: rotation_6d

The dataset and frozen inputs are not included. Use the pinned JSON and effective config.yaml included in this repository.

Downloads last month
1
Video Preview
loading