LeRobot Diffusion Policy fine-tuned on HARP VLA train37

This is the validated final LeRobot policy checkpoint from the joint 37-task HARP VLA fine-tuning run. It predicts absolute Franka joint-position targets plus gripper state (qpos_target_abs, action dimension 8).

Repository: zimplex/harp-vla-train37-diffusion
Final step: 6,250
Final training job: 131349
Final audit job: 131714

Dataset

Field Value
Public source zhouqh/harp_vla_data
Dataset card license MIT at the pinned revision
Pinned source revision 46691c4098ae2ca46a131745d1fe386d63a3fc33
Conversion LeRobot v3, qpos_target_abs, 20 Hz, PyAV video backend
Robot / renderer / schema Franka / NYX / vla_steps_v2_qpos_target
Coverage 37 source tasks, 36 language instructions, 300 episodes/task
Size 11,100 episodes; 4,373,192 frames
Camera payload head RGB + right-wrist RGB; recorded depth groups are empty by contract
Dataset completion marker SHA-256 f9d5121d9f7a210918e5d3dd394d915b00a413d1659538d01729e8f8f6b46db6
Converted data tree SHA-256 65dfc626682a26bfb64478d8b9da9b008aafe2d5536c97b52a21b2c320e8c055
Task manifest SHA-256 b0f46008dfe238c08be3c31f24d86ef80e155e7c5b9b8e75b99c8d73812aa8c6

Training protocol

All five policies used the same optimizer-step and example budget as the earlier HR-Bench train19 sweep. Every policy consumed exactly 6,400,000 training examples (6,400,000 for this model), approximately 1.46 passes over the 4,373,192-frame dataset. This is not equal-epoch training.

Field Value
Model / policy type diffusion
Initialization Diffusion Policy initialized from scratch with a TorchVision ResNet-18 ImageNet-1K backbone
Pinned base revision or digest f37072fd47e89c5e827621c5baffa7500819f7896bbacec160b1a16c560e07ec
Nodes × GPUs/node 4 × 8 NVIDIA B200
Batch per process 32
Global batch 1024
Optimizer steps 6,250
Save cadence every 5,000 steps
Total examples 6,400,000
Seed 1000
Image augmentation disabled
Mixed precision flag use_amp=false
Data workers 4 per process; prefetch factor 4
cuDNN deterministic false
W&B offline; project harp_vla_train37_gcp
Training budget policy same_optimizer_and_example_budget_as_hr_v2_train19

The Diffusion Policy was initialized from scratch apart from the pinned TorchVision ResNet-18 IMAGENET1K_V1 visual backbone.

Pinned software runtime

Component Version / revision
LeRobot source 26ff40ddd784280efc133a8e5af1a76e5ac731c2 plus the recorded HARP entrypoint
Python 3.12
PyTorch / TorchVision 2.8.0 / 0.23.0
Diffusers / Datasets 0.35.2 / 4.4.1
PyAV 15.1.0

Optimizer and scheduler

{
  "optimizer": {
    "betas": [
      0.95,
      0.999
    ],
    "eps": 1e-08,
    "grad_clip_norm": 10.0,
    "lr": 0.0001,
    "type": "adam",
    "weight_decay": 1e-06
  },
  "scheduler": {
    "name": "cosine",
    "num_warmup_steps": 500,
    "type": "diffuser"
  }
}

Policy architecture

Config field Value
n_obs_steps 2
horizon 64
n_action_steps 32
vision_backbone resnet18
pretrained_backbone_weights ResNet18_Weights.IMAGENET1K_V1
use_separate_rgb_encoder_per_camera true
down_dims [512, 1024, 2048]
kernel_size 5
n_groups 8
diffusion_step_embed_dim 128
noise_scheduler_type DDPM
num_train_timesteps 100
beta_schedule squaredcos_cap_v2
prediction_type epsilon
clip_sample true
normalization_mapping {"ACTION": "MIN_MAX", "STATE": "MIN_MAX", "VISUAL": "MEAN_STD"}

Inputs and outputs

Feature Direction Type Shape
observation.image input VISUAL 3 × 256 × 256
observation.state input STATE 8
observation.wrist_image input VISUAL 3 × 256 × 256
action output ACTION 8

The exact, machine-readable training and policy configurations are included as train_config.json and config.json. They are the files emitted by the final checkpoint; their hashes were checked against the immutable completion marker.

Files and loading

The repository root contains the complete pretrained_model/ payload from the LeRobot checkpoint (1.04 GiB), including policy weights, processor state, train_config.json, and config.json. checkpoint_manifest.json records every uploaded checkpoint file's byte size and SHA-256; SHA256SUMS provides the same digests in standard text form. training_provenance.json records the immutable run, dataset, source, job, and audit contract.

Optimizer and RNG files from training_state/ are intentionally not published; this public repository is an inference checkpoint, not a resume bundle.

from lerobot.policies.factory import make_policy_from_pretrained

policy = make_policy_from_pretrained("zimplex/harp-vla-train37-diffusion", device="cuda")
policy.eval()

Use a LeRobot checkout compatible with the configuration included here.

Reproducibility and validation

Field Value
Immutable run ID harp_vla_train37_full_deadline_20261001T151243Z_2747a40
Original protocol source commit 2747a401bd29a070452aeab815d99847d7b507bb
Final attempt source commit da1a0456fe16bab22bdf3759b6577b335c683680
HARP race-safe entrypoint SHA-256 bf6ac26d9d4487ffe3258fd482a4c3561d5effaab9fc501442367a159631e88e
Pinned LeRobot trainer SHA-256 5ae8b0b8f3312f14f054364b56d21f39e3843d650162443286e7342bc489d8c8
Training environment marker SHA-256 dd775f736fa809107dc69b653d5c0eef4679e6958fc9b8950fe6aaa3835c2bab
Completion marker SHA-256 598755aa7bf50578936bab4f523b28be175348c039b7e1f6650fe25f2787cda8
train_config.json SHA-256 36c90b7dcb16831cc72ff0fe3405c341bff4ea8b4e65473492f8e3df58c21eb1
config.json SHA-256 3306dca683170e16adfa3d61e8f46d2cfe5fc365bbcf7703e2821d2c426537fb

The final audit required Slurm COMPLETED/0:0, exact marker/config hashes, the pinned dataset and topology, and finite values in every final Safetensors tensor. This model passed with zero audit errors.

Limitations

  • This release contains training artifacts, not HARP benchmark evaluation scores. No downstream success rate is claimed by this model card.
  • The checkpoint is specific to the converted 20 Hz absolute-qpos convention and its feature/normalization schema.
  • No license is asserted by this card; users must comply with the licenses and terms of the base model/backbone, dataset, LeRobot, and other dependencies.
Downloads last month
17
Safetensors
Model size
0.3B params
Tensor type
F32
·
Video Preview
loading

Dataset used to train zimplex/harp-vla-train37-diffusion