You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
These weights are trained on NVIDIA PhysicalAI-AV data (TanitAD research program). Access is granted per request for research/evaluation use only; you agree not to redistribute.
Log in or Sign Up to review the conditions and access this model content.
TanitAD β REF-A 4-Brain, I-JEPA encoder (in training)
I-JEPA-encoder variant of REF-A, testing whether a less-augmentation-invariant SSL encoder carries more ego-motion signal than DINOv2 (motivated by arXiv:2603.11417 + our encoder decode-ladder showing I-JEPA ~4-5Γ less metrically-inert).
Architecture
- Frozen I-JEPA ViT-H/14 encoder (d_dino 1280), 16Γ16 token grid β temporal grid adapter β operative predictor + tactical/strategic brains.
- Same recipe as the DINOv2 REF-A: speed-input + aux-egomotion + aux-accel; jerk 0.02; rollout_k 12; adapter=temporal. Only the encoder +
--d-dinodiffer.
Training
- Data: PhysicalAI-AV front-wide I-JEPA features, 320-episode rebuilt subset (matched-data comparison vs a DINOv2-320 baseline is planned).
- Checkpoint is a mid-training snapshot (step 14999/15k COMPLETE. Compare against the DINOv2 REF-A companion repo.
Held-out gate
Pending (training in progress). Early training-set grounded ADE dropping fast; positive aux_yaw_r2 early (vs negative for DINOv2) β consistent with I-JEPA carrying more ego-motion info.
Evaluation
Requires the tanitad stack + refa_plus.py (RefAModelPlus, adapter=temporal, d_dino 1280). See config.json.
Trained on PhysicalAI-AV derived features. Gated for research/eval use.
- Downloads last month
- 11