mad-bot's picture
Publish verified DinoVision case-study artifacts
3ad454b verified
|
Raw
History Blame Contribute Delete
5.96 kB

DinoVision

DinoVision is a DINOv3 feature roundtrip for Meta Quest. A frozen DINOv3 encoder turns a passthrough frame into patch features; a decoder trained with Meganeura reconstructs RGB; and Blade renders the result in OpenXR. Both training and deployment use Meganeura—the decoder is not trained or exported through PyTorch.

This repository is the implementation and developer guide. The related material has deliberately separate jobs:

  • Hugging Face distributes the decoder weights, model card, machine-readable results, and reproducibility bundle.
  • REPORT.md is the audited, paper-facing case study and the only place in this repository that states experimental results.
  • experiments/README.md defines the frozen protocols; experiments/AUDIT.md records rejected evidence.

Get the weights

Download the selected decoder from the DinoVision artifact:

hf download mad-bot/dinovision decoder.bin --local-dir ref

The encoder checkpoint is gated and is not redistributed with DinoVision. Accept Meta's terms on the facebook/dinov3-vits16-pretrain-lvd1689m repository, then download it separately:

hf download facebook/dinov3-vits16-pretrain-lvd1689m model.safetensors \
  --local-dir ref

The Hugging Face model card documents the selected weight format, provenance, base-model hash, training data, evaluation records, and license obligations; those details are not duplicated here.

Run on a host

The windowed viewer uses screen capture but otherwise follows the same asynchronous inference path as the headset application:

cargo run --release --features capture --example viewer -- \
  ref/model.safetensors ref/decoder.bin

Useful development entry points:

# One static preview (synthetic weights are used when no model is supplied).
cargo run --release --example preview -- preview.png ref/model.safetensors

# Compute benchmark.
cargo run --release --example bench -- 20

# Train a decoder with Meganeura.
cargo run --release --example train_decoder -- \
  <dataset-directory-or-manifest> ref/model.safetensors \
  [steps] [images] [layers] [size] [seed] [output-directory]

Use the artifact protocol rather than ad-hoc commands when generating a result for citation.

Build and run on Quest

Install the Rust Android target and cargo-apk, and provide Android SDK API 34 with NDK r27. The checked build wrapper injects host-specific toolchain paths, builds the native benchmark and evaluator, builds the XR APK, inspects its manifest, and hashes every output:

.\tools\build_android_artifacts.ps1 -OutputDir C:\tmp\dinovision-android

The XR app loads model.safetensors and decoder.bin from /data/local/tmp/dinovision/. It falls back to synthetic weights when the encoder file is absent. A Quest must be both awake and worn before Horizon OS will actually start an OpenXR activity; SensorLockActivity means the launch is still blocked by the proximity lock.

For paper measurements, use the recovery-aware correctness and timing wrappers in experiments/README.md. They preserve pre-existing device files, retain raw records and device-state snapshots, and reject incomplete runs.

Code map

Path Responsibility
src/dinov3.rs DINOv3 graph, configuration, axial RoPE, and LayerScale layout
src/decoder.rs Trainable feature-to-RGB decoder shared by training and inference
src/preprocess.rs Image crop, resize, normalization, and patch layout
src/weights.rs Checkpoint-to-graph parameter binding
src/inference.rs Asynchronous inference worker and display modes
src/render.rs Blade rendering of the latest completed result
src/source.rs Test, desktop-capture, and headset-camera frame sources
src/bench.rs Host/device benchmark workloads and JSON records
android-xr/ OpenXR passthrough application
examples/ Training, evaluation, reference verification, and interactive tools
tools/ Audited dataset, build, device, artifact, and summary automation

Training caches features from the frozen encoder and optimizes a decoder-only batch graph. Deployment joins a batch-one encoder and decoder into one inference graph. The decoder definition, learned parameters, graph operators, compiler, and runtime are shared; calling the training and deployment graphs identical would be inaccurate.

On Quest, Meganeura receives the same blade_graphics::Context used by the renderer. Inference remains asynchronous, but compute and graphics share a Vulkan queue. Meganeura's submission-chunk control lets the application trade isolated throughput for opportunities to interleave rendering. The current camera path still includes CPU conversion and explicit transfers; it is not a zero-copy or capture-to-photon demonstration.

Validate changes

cargo test --workspace --all-targets
cargo clippy --workspace --all-targets

For an independent encoder comparison, install Torch and Transformers, then generate and verify the reference bundle:

python tools/dump_reference.py --out ref/reference
cargo run --release --example verify -- ref/reference ref/model.safetensors

DinoVision pins Blade to the exact revision selected by Meganeura so the two crates share one blade_graphics type identity. Bump both pins together. The paper artifact records immutable DinoVision, Meganeura, and Blade revisions; the neighboring working-tree locations are never treated as provenance.

License

DinoVision source is MIT-licensed. DINOv3 and the distributed decoder have additional terms described in the DinoVision model card.