open-rig3r / README.md
mca183's picture
Upload README.md with huggingface_hub
281cc2c verified
|
Raw
History Blame Contribute Delete
7.6 kB
---
license: mit
library_name: pytorch
tags:
- 3d-reconstruction
- camera-pose-estimation
- pointmap
- multi-camera-rig
- autonomous-driving
- waymo
- dust3r
datasets:
- waymo_open_dataset
metrics:
- pointmap_err
- pose_deg
- rig_deg
model-index:
- name: rig3r-waymo
results:
- task:
type: camera-pose-estimation
name: Multi-camera rig pose estimation
dataset:
type: waymo_open_dataset
name: Waymo Open Dataset (mini subset, held-out val split)
metrics:
- type: pose_deg
value: 1.4403
name: Pose rotation error (deg)
- type: rig_deg
value: 1.4075
name: Rig rotation error (deg)
- type: pointmap_err
value: 0.2159
name: Pointmap L2 error (scale-normalized)
---
# Open-Rig3R — Waymo (epoch 50)
Unofficial open reimplementation of **Rig3R**: a rig-aware transformer that takes multiple
camera views and predicts per-view pointmaps plus pose and rig raymaps, so a whole camera
rig is reconstructed jointly instead of one camera at a time.
Code: [engichang1467/Open-Rig3R](https://github.com/engichang1467/Open-Rig3R)
This checkpoint is epoch 50 of 50 from a single run on a Waymo mini subset. It is a
**research checkpoint from a reimplementation**, not a reproduction of the paper's
published numbers — see [Limitations](#limitations) before using it for anything.
## Results
Held-out Waymo val split, at the end of training:
| Metric | Value |
|---|---|
| Pointmap L2 error (scale-normalized) | **0.2159** |
| Pose rotation error | **1.440°** |
| Rig rotation error | **1.408°** |
| Pose centre error | 0.0156 |
| Rig centre error | 0.0161 |
### Read the validation curve carefully
`val/total` **rises** from 0.395 (epoch 6) to 0.961 (epoch 50). That is not overfitting,
and epoch 6 is not the better checkpoint.
The pointmap term is `C * err - alpha * log(C)`, where `C` is the model's own predicted
confidence. Over training `C` climbs from 1.2 to 9.98, saturating against the
`conf_max: 10.0` ceiling. The same geometric error therefore costs roughly 10x more at
epoch 50 than at epoch 6. Meanwhile the unweighted error stays flat or improves
(`pointmap_err` 0.279 → 0.216) and both angular errors fall monotonically
(14.10° → 1.44°, 14.16° → 1.41°).
| epoch | val/total | pointmap | **pointmap_err** | conf_mean | pose_deg | rig_deg |
|---|---|---|---|---|---|---|
| 6 | **0.395** | 0.253 | 0.2239 | 3.26 | 3.01 | 3.02 |
| 25 | 0.621 | 0.532 | 0.2187 | 6.86 | 1.63 | 1.59 |
| 50 | 0.961 | 0.880 | **0.2159** | 9.98 | **1.440** | **1.408** |
`val/total` is confounded by confidence saturation and should not be used for model
selection on this run. Select on `pointmap_err`, `pose_deg`, and `rig_deg` — all of which
plateau by roughly epoch 45 and are best at epoch 50.
## Usage
This is a plain `state_dict`, not a `transformers` `PreTrainedModel` — there is no
`from_pretrained`. Load it into the `Rig3R` class from the repo, with the **same
architecture arguments the checkpoint was trained with** (the class defaults differ and
will not load):
```python
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from models.rig3r import Rig3R
model = Rig3R(
encoder_ckpt=None, # weights come from the checkpoint below, not from DUSt3R
img_size=128,
patch_size=16,
embed_dim=1024,
num_decoder_layers=2, # NOT the class default of 6
num_heads=8,
mlp_dim=4096, # NOT the class default of 2048
metadata_dropout=0.5,
)
path = hf_hub_download("mca183/rig3r-waymo", "model.safetensors")
model.load_state_dict(load_file(path))
model.eval()
# images: (B, V, 3, 128, 128) — V = n_frames * num_cameras, e.g. 2 * 5 = 10 for the full Waymo rig
# metadata is optional; the decoder is trained with per-field dropout so it runs without it
outputs = model(images, metadata=None)
```
`forward` returns a dict:
| key | shape | notes |
|---|---|---|
| `pointmap` | `(B, V, H*W, 3)` | dense per-pixel 3D points |
| `pointmap_conf` | `(B, V, H*W, 1)` | per-pixel confidence — saturated, see [Limitations](#limitations) |
| `pose_raymap` | `(B, V, P, 6)` | centre + unit direction per patch |
| `rig_raymap` | `(B, V, P, 6)` | centre + unit direction per patch |
| `camera_center_pose` | `(B, V, 3)` | one centre per view |
| `camera_center_rig` | `(B, V, 3)` | one centre per view |
| `features` | `(B, V, P, C)` | decoder patch features, for downstream heads |
`P = (128 / 16)^2 = 64` patches per view.
## Architecture
| | |
|---|---|
| Total parameters | 340.2 M (373 tensors, fp32) |
| Encoder | 303.2 M — DUSt3R ViT-L/16, **frozen** during training |
| Rig-aware decoder | 37.0 M — 2 pre-norm transformer layers, 8 heads, MLP dim 4096 |
| Heads | `pointmap_head`, `pose_raymap_head`, `rig_raymap_head` |
| Embedding dim | 1024 |
| Input resolution | 128 x 128, patch size 16 |
Encoder initialized from `DUSt3R_ViTLarge_BaseDecoder_512_dpt` and kept frozen, so only
the 37.0 M decoder and heads were trained.
## Training
| | |
|---|---|
| Dataset | Waymo Open Dataset, mini subset, full 5-camera rig (FRONT, FRONT_LEFT, FRONT_RIGHT, SIDE_LEFT, SIDE_RIGHT) |
| Views per sample | 10 (`n_frames: 2` x 5 cameras) |
| Epochs | 50 |
| Batch size | 8 |
| Optimizer | AdamW, lr 1e-4, weight decay 0.01 |
| Scheduler | Cosine annealing, `eta_min` 1e-6 |
| Precision | bf16 autocast, no grad scaler |
| Loss weights | `w_point` 1.0, `w_pose` 1.0, `w_rig` 1.0 |
| Confidence regularizer | `alpha` 0.2, `beta` 1.0, `conf_max` 10.0 |
| Metadata dropout | 0.5 per field (frame index exempt) |
| Seed | 0 |
| Hardware | NVIDIA A100 80GB PCIe |
| Wall clock | ~3h 20m |
`conf_max: 10.0` is a deliberate deviation from the paper: Eq. 3 leaves `-alpha*log(C)`
unbounded below, so the ceiling floors the pointmap term at `-alpha*log(10) = -0.46`.
Setting it to `null` restores Eq. 3 exactly.
## Limitations
- **Reimplementation, not the paper.** Unofficial; numbers here are not comparable to published Rig3R results.
- **Mini subset.** Trained on a small Waymo subset, not the full dataset. Generalization is untested.
- **128 x 128 input.** Well below the paper's resolution; pointmap detail is correspondingly coarse.
- **Shallow decoder.** 2 layers rather than 6, and the encoder is frozen throughout.
- **Waymo domain only.** Driving scenes, one rig geometry. No indoor, handheld, or non-automotive evaluation.
- **Single run, no held-out test.** Metrics are val-split only, one seed, no ablations.
- **Confidence is saturated.** `conf_mean` sits at the `conf_max` ceiling of 10.0, so predicted confidence is not calibrated and should not be read as an uncertainty estimate.
## Dataset terms
Trained on the Waymo Open Dataset, which carries its own license and terms of use. Using
this model does not grant any rights to that data — obtain it from Waymo directly and
comply with their terms.
## License
MIT (see the [source repo](https://github.com/engichang1467/Open-Rig3R)). Copyright 2025
Michael Chang. The DUSt3R encoder initialization and the Waymo Open Dataset carry their
own separate licenses.
## Citation
Rig3R (original paper):
```bibtex
@article{rig3r,
title = {Rig3R: Rig-Aware Conditioning for Learned 3D Reconstruction},
year = {2025}
}
```
This reimplementation:
```bibtex
@software{open_rig3r,
author = {Chang, Michael},
title = {Open-Rig3R: An open reimplementation of Rig3R},
url = {https://github.com/engichang1467/Open-Rig3R},
year = {2025}
}
```