SLIM for CALVIN

This repository contains the released SLIM Stage 2 epoch-15 policy checkpoint for the CALVIN ABC-D benchmark. The Stage 2 run was trained for 40 epochs, and epoch 15 was selected for the reported result. SLIM is a compact latent interaction policy for robot manipulation.

Checkpoint

  • Stage 1: action-grounded masked trajectory prediction, IDM:FDM = 0.125:1
  • Stage 1 duration: 3 epochs
  • Stage 2: flow-matching policy training for 40 epochs; this release uses the checkpoint saved at epoch 15
  • Action horizon and execution stride: 12
  • State/action dimensions: 15/7

The checkpoint is a plain PyTorch state_dict and loads directly with SLIM.

Results

The paper result on 1,000 CALVIN ABC-D long-horizon sequences is an average successful sequence length of 4.556. Success rates at sequence lengths 1 through 5 are 99.3%, 96.7%, 92.3%, 87.1%, and 80.2%.

A fresh evaluation of the released package produced an average successful sequence length of 4.578, with success rates of 99.9%, 96.4%, 92.2%, 88.0%, and 81.3%. Machine-readable results are included under evaluation/.

Usage

Install SLIM and configure the DINOv2 and T5 paths as described in the SLIM README. Then start a policy server from the SLIM repository root:

python -m slim.serving.server \
  --checkpoint /path/to/SLIM-CALVIN/checkpoints/epoch_15_pytorch_model.pt \
  --port 10093 \
  --bf16

Keep the included config.yaml and action_stats.json in the repository root. See checkpoint_manifest.json for hashes and the exact SLIM revision.

Limitations

This checkpoint is intended for research evaluation in CALVIN-compatible simulation environments. It should not be deployed on physical robots without task-specific safety validation and action-bound checks.

Downloads last month
7
Video Preview
loading