Buckets:
| # MoGe-2 monocular depth — qualitative previews | |
| Full-length monocular depth over the clips published in this bucket, rendered as a | |
| 2x2 panel video (2560x1440, 15 fps): **RGB | depth (global range) | the same depth at a | |
| near-field range | normal**, each panel at the clip's native 1280x720. | |
| These are previews for judging detail by eye. They are **not** a dataset tier and | |
| nothing in the pipeline consumes them — the pipeline's own depth comes from | |
| PointWorld `scene_flows` (sparse metric 3D tracks), not from MoGe. | |
| ## Model | |
| `Ruicheng/moge-2-vitl-normal`, `resolution_level=9`. | |
| `resolution_level=9` is the default and already maps to the **top** of MoGe-2's | |
| `num_tokens_range` (3600 tokens) — the finest setting the model offers. | |
| **MoGe-3 is not used, because it does not exist as runnable code.** The paper | |
| (arXiv 2607.17967, "Fine-Detail Monocular Geometry Estimation…") was announced in | |
| the upstream README on 2026-07-21, but at pinned commit `925b8ed` — still | |
| `origin/main` as of this run — there is no `moge/model/v3.py` and the pretrained | |
| table lists `moge-3-vitg`/`moge-3-vitl` as "coming soon" with no HF repo id. | |
| ## How to read the panels | |
| * **Depth** uses **one global normalisation for the whole clip** (robust 2/98 | |
| percentiles over every valid pixel of every frame), never per frame. Per-frame | |
| normalisation makes the panel breathe as the near/far extremes move, which by eye | |
| is indistinguishable from the model's own temporal instability — the exact thing | |
| these videos exist to let someone judge. The colour range in metres is burned into | |
| the panel label. | |
| * **Normal** is where fine detail actually shows. Depth is dominated by the | |
| metre-scale front-to-back ramp of the workbench; the normal map spends its whole | |
| colour budget on surface relief instead. | |
| * **No `fov_x` was passed**, so MoGe estimates its own FOV per frame even though this | |
| project knows the true DROID intrinsics. That is the out-of-the-box path — priming | |
| the model with the calibration would measure a best case that arbitrary video would | |
| not get. | |
| * Videos are muxed at **15 fps**, the DROID trajectory rate. The source mp4 | |
| containers advertise 60/1, but their frames are the 15 Hz trajectory samples (this | |
| repo indexes poses per video frame at `DEFAULT_TRAJECTORY_FPS = 15.0`). Playing at | |
| 60 would look like temporal jitter that is not there. | |
| ## Clips | |
| | file | camera / class | frames | depth range (m) | near mm/level | size | | |
| | --- | --- | --- | --- | --- | --- | | |
| | `IRIS+226b9f35+2023-07-11-12h-01m-54s.mp4` | IRIS / rigid_symmetric | 358 | 0.23–1.24 | 0.8 | 35.9 MB | | |
| | `RAD+62828209+2023-09-09-09h-01m-13s.mp4` | RAD / rigid_thin | 602 | 0.53–4.22 | 1.1 | 42.0 MB | | |
| | `RAIL+80edfcb1+2023-04-26-15h-45m-39s.mp4` | RAIL / deformable | 191 | 0.32–3.07 | 2.2 | 18.7 MB | | |
| | `RAIL+80edfcb1+2023-06-29-12h-44m-09s.mp4` | RAIL / articulated | 428 | 0.30–5.05 | 1.6 | 38.8 MB | | |
| The **near mm/level** column is the one to read before trusting a panel: it is the | |
| millimetres of depth per colour level in the near-field panel. Scenes with a deep | |
| background push the *global* panel to 15–19 mm/level, where nothing object-sized is | |
| visible at all; the near-field panel is what stays readable. | |
| These clips come from the diversity sample: different labs, different object classes, | |
| different scene depths. That is what makes them worth looking at — and also why the | |
| per-clip colour ranges differ so much. Do not compare colours between two clips; only | |
| the near-field range printed in each panel makes a panel's colours mean anything. | |
Xet Storage Details
- Size:
- 3.57 kB
- Xet hash:
- 37085305ae24667773cfb5153d4005200cc0c4e857d9577bee70f207f2fe3cb3
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.