M2H-MX ONNX weights β monocular depth + semantics
Exported ONNX weights for the two M2H-MX networks used by
vio_stack_jetson, a monocular
VIO stack running on a Jetson Orin with a ZED X One GS camera.
Both networks take a single RGB frame and produce metric depth and per-pixel semantic labels in one pass. Both publish the same two ROS 2 topics, so they are swappable at runtime:
/m2h/depth/image 32FC1 metric depth
/m2h/semantic/labels_argmax mono8 class IDs
Built with DINOv3.
Files
The layout mirrors the ROS packages that consume these files, so a manifest line maps to a path with no translation:
| File | Size | Notes |
|---|---|---|
m2h_mx_base_onnx_ros/models/scannet/m2h_mx_b_scannet_240x320.onnx |
421 MB | base network, default |
m2h_mx_base_onnx_ros/models/scannet/m2h_mx_b_scannet_480x640.onnx |
422 MB | base network, full res |
m2h_mx_base_onnx_ros/models/nyudv2/m2h_mx_b_nyudv2_240x320.onnx |
494 MB | NYUDv2-trained variant |
m2h_mx_base_onnx_ros/models/nyudv2/m2h_mx_b_nyudv2_480x640.onnx |
496 MB | NYUDv2-trained variant |
m2h_mx_large_onnx/onnx_models/scannet_depth_sem_192x256_trt_clean.onnx |
1.26 GB | large network, default |
m2h_mx_large_onnx/onnx_models/scannet_depth_sem_320x416.onnx |
1.26 GB | large network, higher res |
MANIFEST.sha256 in each package directory lists the expected SHA256 of every
file. Verify after downloading β a network running on the wrong weights produces
plausible-looking depth rather than an error, which is much harder to notice than
a refusal to start.
Usage
git clone https://github.com/BavanthaU/vio_stack_jetson
cd vio_stack_jetson
./tools/fetch_weights.sh # pulls these files and verifies every sum
Or directly:
from huggingface_hub import snapshot_download
snapshot_download(repo_id="Bavantha11/vio-stack-jetson-weights", repo_type="model")
Measured on the rig
Jetson Orin, JetPack r36.5, ONNX Runtime with the TensorRT execution provider:
| Network | Resolution | Latency | Rate |
|---|---|---|---|
| base | 240x320 | ~130 ms/frame | ~6.9 Hz |
| large | 192x256 | ~155 ms/frame | β |
The first run at any resolution builds a TensorRT engine β minutes, once, cached per machine. The engines are not portable across machines or driver versions and are deliberately not published here.
Depth is not fed back into the VIO filter by default; it runs alongside it.
Architecture
- Backbone: DINOv3 ViT-B/16 (
dinov3-vitb16-pretrain-lvd1689m) - Heads: two β semantic segmentation (40 classes) and metric depth
- Opset: 17, static shapes
- Training data: ScanNet, NYUDv2
Licensing β read before use
These weights combine inputs with different terms. The most restrictive one governs the result.
DINOv3 backbone β DINOv3 License,
included here as DINOv3_LICENSE.md. It permits commercial use, modification and
redistribution of derivative works, and asks in return that redistributions carry
a copy of the licence and display "Built with DINOv3". Both are satisfied above.
ScanNet β Terms of Use. Non-commercial research and educational use only, and those terms extend to derivative works.
NYUDv2 β research use.
Therefore these weights are for non-commercial research and educational use only. ScanNet is the binding constraint; DINOv3 permitting commercial use does not lift it. Anyone needing a commercial deployment must retrain on commercially licensed data β the ROS packages and the export pipeline are MIT and unaffected.
The MIT licence on the m2h_mx_* ROS packages covers the wrapper code only,
not these weights.
Citation
If you use these in published work, please acknowledge DINOv3 and cite the ScanNet and NYUDv2 datasets as their terms require.