You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
These weights are trained on NVIDIA PhysicalAI-AV data (TanitAD research program). Access is granted per request for research/evaluation use only; you agree not to redistribute.
Log in or Sign Up to review the conditions and access this model content.
TanitAD β REF-A 4-Brain, DINOv2 encoder (step 29,999, complete)
Reference arm A of the TanitAD 3-arm study: frozen-encoder world model β the "stability-by-construction" reference vs REF-B's trained encoder.
Architecture
- Frozen DINOv2-B/14 encoder (d_dino 768), 16Γ16 token grid β temporal grid adapter β shared operative predictor + tactical/strategic brains.
- speed-input (ego-speed v0, action_dim 3) + aux-egomotion (speed/yaw regressors) + aux-accel head; jerk 0.02; rollout_k 12; adapter=temporal.
Training
- Data: PhysicalAI-AV front-wide DINOv2 features, 2,376 episodes / 406,099 windows.
- Step 29,999 / 30,000 β complete.
Held-out gate (grounded rollout, 40 eps, 8 splits)
| metric | value |
|---|---|
| ADE@2s | 2.136 m |
| ADE@1s | 1.378 m |
| CV baseline ADE@2s | 0.825 m |
Key finding: frozen DINOv2 is metrically inert for ego-motion (decode-ladder speed RΒ²β0.29, yaw-rate negative) β the model drives as a learned action-integrator; vision is near-dead-weight. Motivates the I-JEPA variant (companion repo).
Evaluation
Requires the tanitad stack + refa_plus.py (RefAModelPlus, adapter=temporal, d_dino 768). See config.json.
Trained on PhysicalAI-AV derived features. Gated for research/eval use.
- Downloads last month
- 15