File size: 6,857 Bytes
8df83c1 98dd9e1 8df83c1 d00b170 8df83c1 d00b170 e67dc42 e9fa246 933e92a 9b4364f 933e92a 9b4364f 933e92a 9b4364f 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 98dd9e1 933e92a 9b4364f 933e92a 9b4364f 98dd9e1 933e92a 98dd9e1 933e92a 9b4364f 933e92a 9b4364f 98dd9e1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 | ---
license: apache-2.0
tags:
- computer-vision
- robotics
- articulated-objects
language:
- en
---
# L-CAMP: Language-Conditioned Axis and Motion Prediction for Articulated Object Manipulation
[Ikechukwu Daniel Adebi](https://danieladebi.github.io), [Peter Stone](https://www.cs.utexas.edu/~pstone/), [Mitchell Pryor](https://www.me.utexas.edu/people/faculty-directory/pryor)
The University of Texas at Austin
Code: [UTNuclearRobotics/l_camp](https://github.com/UTNuclearRobotics/l_camp)
## About this checkpoint
This repository hosts a pretrained L-CAMP checkpoint, `lcamp_model.pt`, available from the Files tab above. It predicts the screw axis and motion of an articulated object part from an RGB(-D) image conditioned on a language instruction.
To use it, download `lcamp_model.pt` from the Files tab and place it in a `checkpoints/` directory inside a clone of the [GitHub repo](https://github.com/UTNuclearRobotics/l_camp), then follow the setup and evaluation steps below.
## Setup
1. Clone the [GitHub repo](https://github.com/UTNuclearRobotics/l_camp) and create the conda environment from `environment_lcamp.yml`:
```bash
conda create -f environment_lcamp.yml
conda activate lcamp
```
2. Build the LCAMP dataset (only needed once per dataset version), see
`process_lcamp_dataset.py`. Whatever path you pass as its `--data_dir` is
the dataset directory (containing `train_dataset.json`, `test_dataset.json`,
and `intrinsics.npy`) referenced below as `<DATASET_DIR>`
(`--data_root`/`--data_dir` in `train.py`/`eval.py`).
3. Get the background images: `LCAMPDataset` (in `lcamp_dataset.py`)
unconditionally loads a background image list on init (used for the
`--randomize_background`/`--random_backgrounds` augmentation, but read
regardless), sourced from the
[MIT Indoor 67](http://web.mit.edu/torralba/www/indoor.html) scene dataset.
From that page, download the images archive and the `TrainImages.txt` /
`TestImages.txt` split files, then lay them out as:
```
backgrounds/mit_indoor_67/TrainImages.txt
backgrounds/mit_indoor_67/TestImages.txt
backgrounds/mit_indoor_67/raw/Images/<category>/<image>.jpg
```
relative to wherever you run `train.py`/`eval.py` from (or under
`$SCRATCH/backgrounds/mit_indoor_67/...` if using `--cluster_loc
stampede3` — see the `--cluster_loc` note below).
All commands below (`train.py`, `eval.py`) assume the `lcamp` environment is
active.
## Training
If you'd rather train your own checkpoint instead of using this one, training is done with `train.py` via `torchrun` (DDP). The dataset must already be built (see `process_lcamp_dataset.py`) as a directory containing `train_dataset.json`, `test_dataset.json`, and `intrinsics.npy`.
Set `--nproc_per_node` to the number of GPUs on your machine (use `1` for a
single-GPU/local run):
```bash
torchrun --nproc_per_node=<NUM_GPUS> train.py \
--save_path checkpoints/my_lcamp_model \
--data_root <DATASET_DIR> \
--model_type resnet \
--epochs 50 \
--batch_size 64 \
--lr 1e-3 \
--min_lr 1e-5 \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--emphasize_part_mask \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--lambda_axis 8 \
--lambda_anchor 8 \
--lambda_joint 1
```
Notable flags (`train.py --help` for the complete list):
- `--model_type`: `resnet` (default LCAMP backbone) or `dino`.
- `--use_text_instructions`: enables the language-conditioned model
(`LCAMPModelResnetLanguage`); required if you plan to run the `lcamp`
backend with instructions in `lcamp_service_node.py`.
- `--use_depth`: adds depth as an input channel.
- `--bbox_loc`: predict axis location as `(u,v,z)` decoded via the part
bounding box instead of raw XYZ, must match `--bbox_loc` in `eval.py` and
`use_bbox_loc` in `lcamp_service_node.py`/`lcamp.launch.py` for the same
checkpoint. Whenever this is active, pass `--bbox_loc` explicitly so the
flag stays consistent across `train.py`/`eval.py`/deployment.
- `--resume` / `--resume_from <path>`: resume from `checkpoint_best.pt` in
`--save_path`, or from an explicit checkpoint path.
- `--save_freq`, `--eval_freq`: checkpoint/eval cadence in epochs.
- `--cluster_loc`: the dataset JSON (from `process_lcamp_dataset.py`) stores
absolute file paths for the rgb/depth/mask files, so if you're loading the
dataset on a different machine/path layout than where it was built, those
paths won't resolve. `nrg` (default) uses the JSON paths as-is, `stampede3`
is a legacy alias for `$SCRATCH`. For anything else, just pass your own
dataset root directory as `--cluster_loc /path/to/your/root` — `lcamp_
dataset.py` will swap it in for the legacy `/storage/danieladebi` prefix in
the rgb/depth/mask paths, and root the background images (see above) under
it too. No code changes needed.
Checkpoints are written to `--save_path` as `checkpoint_best.pt`,
`checkpoint_last.pt`, and periodic `checkpoint_{epoch}.pt` files.
## Evaluation
Evaluation is done with `eval.py` (single process, no DDP). To evaluate this checkpoint:
```bash
python eval.py \
--model_path checkpoints/lcamp_model.pt \
--data_dir <DATASET_DIR> \
--split test \
--model_type resnet \
--use_camera_frame \
--use_depth \
--use_text_instructions \
--bbox_loc \
--cluster_loc <CLUSTER_LOC> \
--visualize
```
This must use the **same architecture flags used at training time**
(`--model_type`, `--use_depth`, `--use_text_instructions`,
`--exclude_object_mask`, `--emphasize_part_mask`/`--emphasize_object_mask`,
`--bbox_loc`, etc.), mismatches will load state dict weights incorrectly or
error, and `--use_camera_frame`/`--cluster_loc` must match how the dataset
was built.
Notable flags:
- `--split`: `train` or `test`.
- `--visualize`: saves per-sample prediction images.
- `--random_preds`: baseline using random axis/location predictions.
- `--free_motion_query_prob`: fraction of eval samples using a free-motion
(zero-vector) query instead of the GT axis, to measure reliance on visual
cues vs. the GT axis itself.
Results (per-object/per-category success rates, axis/location error, joint
type accuracy) are written to `results/<model_name>/`, where `<model_name>`
is derived from `--model_path`.
Once you're happy with a checkpoint (this one or your own), point the
`model_name` launch argument in `launch/lcamp.launch.py` at it (see the
launch file for details on `use_bbox_loc` and other runtime flags).
## Deploying on a Robot
Running predictions on a real Boston Dynamics Spot (via the ROS2 service node
in `scripts/lcamp_service_node.py`) is outside the scope of this model card,
see [ros2_setup/ROS2_SETUP.md](https://github.com/UTNuclearRobotics/l_camp/blob/main/ros2_setup/ROS2_SETUP.md) in the GitHub repo for that setup.
|