Spaces:
Running on Zero
Running on Zero
File size: 1,834 Bytes
49d36c0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 | # Model Overview
## What is GEM?
GEM (Generalist Estimation of Human Motion) is a commercial-grade monocular video 3D human pose estimation model developed by NVIDIA. It recovers full-body 77-joint motion (body, hands, and face) from monocular video using the SOMA parametric body model.
## GEM vs GENMO
GEM is the commercial successor to the research project GENMO. Key differences:
| | GENMO (Research) | GEM (This Repo) |
|---|---|---|
| Body model | SMPL (24 joints) | SOMA (77 joints: body + hands + face) |
| Modalities | Video + text + audio + music | Video only |
| License | Research only | Apache 2.0 (commercial) |
| 2D Pose Model | COCO-17 keypoints | SOMA-77 keypoints |
| Training Data | Public + internal | NVIDIA-owned only |
## SOMA Body Model
GEM uses the **SOMA** parametric body model:
- **77 joints** covering full body, hands, and face
- **MHR identity model** for body shape representation
- Bundled as a submodule in `third_party/soma`
## Architecture
| Property | Value |
|---|---|
| Parameters | ~520M |
| Architecture | 16-layer Transformer encoder |
| Positional encoding | RoPE (Rotary Position Embedding) |
| Latent dimension | 1024 |
| Attention heads | 8 |
| Decoder | Regression-based denoising decoder |
| Optimizer | AdamW (lr=2e-4) |
| Precision | 16-bit mixed |
The model takes as input video features (from SAM-3D-Body), 2D keypoints, and camera intrinsics, and outputs per-frame SOMA body parameters in both camera-relative and global coordinate frames.
## Bundled 2D Pose Model
GEM includes a trained 2D pose estimation model:
- **Backbone:** DINOv3 (ViT-based)
- **Output:** 77 SOMA keypoints per frame (x, y, confidence)
- **Usage:** Automatically invoked during the demo preprocessing pipeline
The 2D pose model is accessed via `gem.utils.vitpose_extractor.VitPoseExtractor`.
|