Spaces:
Running on Zero
Running on Zero
File size: 5,759 Bytes
49d36c0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 | # Demo
## Pipeline Overview
```
Input Video β Human Detection (YOLOX) β 2D Keypoints (VitPose) β Features (SAM3D) β GEM Model β 3D Pose (SOMA)
β β
2D Keypoint Overlay (Optional) Retarget β G1 Robot Motion
```
All demo scripts use **[YOLOX](https://github.com/Megvii-BaseDetection/YOLOX) + [ByteTrack](https://github.com/ifzhang/ByteTrack)** for person detection and tracking.
## Full 3D Pipeline (`demo_soma.py`)
Run full inference on a video:
```bash
python scripts/demo/demo_soma.py \
--video path/to/video.mp4 \
--output_root outputs \
--ckpt inputs/pretrained/gem_soma.ckpt
```
> **Note:** The `--ckpt` argument is optional. If omitted, the script will automatically download the pretrained checkpoint from [HuggingFace](https://huggingface.co/nvidia/GEM-X).
### Arguments
| Argument | Default | Description |
|---|---|---|
| `--video` | β | Input video path (required) |
| `--ckpt` | `null` | Pretrained checkpoint path |
| `-s` / `--static_cam` | off | Assume static camera (disables VO) |
| `--output_root` | `outputs/demo_soma` | Root directory for outputs |
| `--verbose` | off | Save debug overlays (bbox, pose) |
| `--render_mhr` | off | Render MHR identity model |
| `--retarget` | off | Retarget motion to Unitree G1 robot (requires soma-retargeter) |
### Outputs
Results are saved to `<output_root>/<video_name>/`:
| File | Description |
|---|---|
| `0_kp2d77_overlay.mp4` | 2D keypoint overlay on input video |
| `<video_name>_1_incam.mp4` | In-camera mesh overlay |
| `<video_name>_2_global.mp4` | Global-coordinate render |
| `<video_name>_3_incam_global_horiz.mp4` | Side-by-side (or 2x2 grid with `--retarget`) |
| `preprocess/bbx.pt` | Detected bounding boxes |
| `preprocess/vitpose.pt` | 2D keypoints (77 joints) |
| `preprocess/hpe_results.pt` | Full 3D pose prediction |
| `<video_name>_retarget_g1.bvh` | G1 robot motion in BVH format (with `--retarget`) |
| `<video_name>_retarget_g1.csv` | G1 robot joint angles (with `--retarget`) |
| `<video_name>_4_g1_retarget.mp4` | G1 robot motion video (with `--retarget`) |
### Preprocessing Fallbacks
- When no pre-computed `bbx.pt` exists, the demo runs human detection via YOLOX + ByteTrack.
- If VO modules are unavailable, the demo falls back to a static camera trajectory.
## Accelerated Pipeline (`demo_soma_onnx.py`)
ONNX/TensorRT-accelerated variant of `demo_soma.py`. Replaces PyTorch inference with ONNX Runtime for VitPose, SAM-3D-Body, and the GEM denoiser.
> **macOS:** This script supports Apple Silicon via ONNX Runtime with the CoreML Execution Provider. See [INSTALL_MACOS.md](INSTALL_MACOS.md) for setup instructions.
```bash
python scripts/demo/demo_soma_onnx.py \
--video path/to/video.mp4
```
### Prerequisites
ONNX models are automatically downloaded from [HuggingFace](https://huggingface.co/nvidia/GEM-X) on first run if not found locally. To export your own ONNX models instead:
```bash
python tools/export/export_vitpose_onnx.py
python tools/export/export_sam3db_onnx.py
python tools/export/export_denoiser_onnx.py --ckpt <path>
```
### Arguments
| Argument | Default | Description |
|---|---|---|
| `--video` | β | Input video path (required) |
| `--ckpt` | `null` | Pretrained checkpoint path |
| `-s` / `--static_cam` | off | Assume static camera (disables VO) |
| `--output_root` | `outputs/demo_soma_onnx` | Root directory for outputs |
| `--verbose` | off | Save debug overlays |
| `--force_pytorch` | off | Force PyTorch inference even if ONNX/TRT available |
| `--no-imgfeat` | off | Skip SAM3DB, use 2D keypoints only |
| `--ddim` | off | DDIM sampling (50 steps) instead of regression β slower but higher quality |
| `--retarget` | off | Retarget motion to Unitree G1 robot |
### Outputs
Same as `demo_soma.py`. When `--retarget` is used, the final composite is a 2x2 grid (kp2d, incam, global, retarget).
## Humanoid Robot Retargeting (`--retarget`)
Retarget the recovered SOMA motion to a Unitree G1 humanoid robot:
```bash
python scripts/demo/demo_soma.py \
--video path/to/video.mp4 \
--retarget
```
This requires the soma-retargeter package (see [Installation](INSTALL.md)). The output includes a G1 robot motion video and joint angle CSV. When `--retarget` is used, the final composite video shows a 2x2 grid: 2D keypoints, in-camera mesh, global mesh, and G1 robot motion.
## 2D Keypoint-Only Demo (`demo_2d_keypoints.py`)
A lightweight demo that runs only detection and 2D keypoint extraction β no GEM model, no 3D rendering, no Hydra config.
```bash
python scripts/demo/demo_2d_keypoints.py \
--video path/to/video.mp4
```
### Arguments
| Argument | Default | Description |
|---|---|---|
| `--video` | β | Input video path (required) |
| `--output_dir` | `outputs/demo_2d_kp/<video_name>/` | Output directory |
| `--detector_name` | `vitdet` | Human detector: `vitdet` or `sam3` |
| `--conf_thr` | `0.5` | Confidence threshold for visualization |
| `--save_raw` | off | Keep intermediate `.pt` files |
### Output
- `<video_name>_kp2d77_overlay.mp4` β 2D keypoint overlay video
## Accessing Results Programmatically
```python
import torch
# Load 2D keypoints
vitpose = torch.load("outputs/demo_soma/<video>/preprocess/vitpose.pt")
# vitpose shape: (num_frames, 77, 3) β x, y, confidence
# Load bounding boxes
bbx = torch.load("outputs/demo_soma/<video>/preprocess/bbx.pt")
bbx_xyxy = bbx["bbx_xyxy"] # (num_frames, 4)
bbx_xys = bbx["bbx_xys"] # (num_frames, 3) β center_x, center_y, scale
# Load 3D prediction
pred = torch.load("outputs/demo_soma/<video>/preprocess/hpe_results.pt")
```
|