L-CAMP: Language-Conditioned Axis and Motion Prediction for Articulated Object Manipulation
Ikechukwu Daniel Adebi, Peter Stone, Mitchell Pryor
The University of Texas at Austin
Code: UTNuclearRobotics/l_camp
About this checkpoint
This repository hosts a pretrained L-CAMP checkpoint, lcamp_model.pt, available from the Files tab above. It predicts the screw axis and motion of an articulated object part from an RGB(-D) image conditioned on a language instruction.
To use it, download lcamp_model.pt from the Files tab and place it in a checkpoints/ directory inside a clone of the GitHub repo, then follow the setup and evaluation steps below.
Setup
- Clone the GitHub repo and create the conda environment from
environment_lcamp.yml:conda create -f environment_lcamp.yml conda activate lcamp - Build the LCAMP dataset (only needed once per dataset version), see
process_lcamp_dataset.py. Whatever path you pass as its--data_diris the dataset directory (containingtrain_dataset.json,test_dataset.json, andintrinsics.npy) referenced below as<DATASET_DIR>(--data_root/--data_dirintrain.py/eval.py). - Get the background images:
LCAMPDataset(inlcamp_dataset.py) unconditionally loads a background image list on init (used for the--randomize_background/--random_backgroundsaugmentation, but read regardless), sourced from the MIT Indoor 67 scene dataset. From that page, download the images archive and theTrainImages.txt/TestImages.txtsplit files, then lay them out as:
relative to wherever you runbackgrounds/mit_indoor_67/TrainImages.txt backgrounds/mit_indoor_67/TestImages.txt backgrounds/mit_indoor_67/raw/Images/<category>/<image>.jpgtrain.py/eval.pyfrom (or under$SCRATCH/backgrounds/mit_indoor_67/...if using--cluster_loc stampede3— see the--cluster_locnote below).
All commands below (train.py, eval.py) assume the lcamp environment is
active.
Training
If you'd rather train your own checkpoint instead of using this one, training is done with train.py via torchrun (DDP). The dataset must already be built (see process_lcamp_dataset.py) as a directory containing train_dataset.json, test_dataset.json, and intrinsics.npy.
Set --nproc_per_node to the number of GPUs on your machine (use 1 for a
single-GPU/local run):
torchrun --nproc_per_node=<NUM_GPUS> train.py \
--save_path checkpoints/my_lcamp_model \
--data_root <DATASET_DIR> \
--model_type resnet \
--epochs 50 \
--batch_size 64 \
--lr 1e-3 \
--min_lr 1e-5 \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--emphasize_part_mask \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--lambda_axis 8 \
--lambda_anchor 8 \
--lambda_joint 1
Notable flags (train.py --help for the complete list):
--model_type:resnet(default LCAMP backbone) ordino.--use_text_instructions: enables the language-conditioned model (LCAMPModelResnetLanguage); required if you plan to run thelcampbackend with instructions inlcamp_service_node.py.--use_depth: adds depth as an input channel.--bbox_loc: predict axis location as(u,v,z)decoded via the part bounding box instead of raw XYZ, must match--bbox_locineval.pyanduse_bbox_locinlcamp_service_node.py/lcamp.launch.pyfor the same checkpoint. Whenever this is active, pass--bbox_locexplicitly so the flag stays consistent acrosstrain.py/eval.py/deployment.--resume/--resume_from <path>: resume fromcheckpoint_best.ptin--save_path, or from an explicit checkpoint path.--save_freq,--eval_freq: checkpoint/eval cadence in epochs.--cluster_loc: the dataset JSON (fromprocess_lcamp_dataset.py) stores absolute file paths for the rgb/depth/mask files, so if you're loading the dataset on a different machine/path layout than where it was built, those paths won't resolve.nrg(default) uses the JSON paths as-is,stampede3is a legacy alias for$SCRATCH. For anything else, just pass your own dataset root directory as--cluster_loc /path/to/your/root—lcamp_ dataset.pywill swap it in for the legacy/storage/danieladebiprefix in the rgb/depth/mask paths, and root the background images (see above) under it too. No code changes needed.
Checkpoints are written to --save_path as checkpoint_best.pt,
checkpoint_last.pt, and periodic checkpoint_{epoch}.pt files.
Evaluation
Evaluation is done with eval.py (single process, no DDP). To evaluate this checkpoint:
python eval.py \
--model_path checkpoints/lcamp_model.pt \
--data_dir <DATASET_DIR> \
--split test \
--model_type resnet \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--visualize
This must use the same architecture flags used at training time
(--model_type, --use_depth, --use_text_instructions,
--exclude_object_mask, --emphasize_part_mask/--emphasize_object_mask,
--bbox_loc, etc.), mismatches will load state dict weights incorrectly or
error, and --use_camera_frame/--cluster_loc must match how the dataset
was built.
Notable flags:
--split:trainortest.--visualize: saves per-sample prediction images.--random_preds: baseline using random axis/location predictions.--free_motion_query_prob: fraction of eval samples using a free-motion (zero-vector) query instead of the GT axis, to measure reliance on visual cues vs. the GT axis itself.
Results (per-object/per-category success rates, axis/location error, joint
type accuracy) are written to results/<model_name>/, where <model_name>
is derived from --model_path.
Once you're happy with a checkpoint (this one or your own), point the
model_name launch argument in launch/lcamp.launch.py at it (see the
launch file for details on use_bbox_loc and other runtime flags).
Deploying on a Robot
Running predictions on a real Boston Dynamics Spot (via the ROS2 service node
in scripts/lcamp_service_node.py) is outside the scope of this model card,
see ros2_setup/ROS2_SETUP.md in the GitHub repo for that setup.