tanitad-refb-speed / README.md
Sayood's picture
refb-v2 30k FINAL: model card with full TanitEval results (ade@2s 0.592, beats CV, A/B vs 20k/v1/flagship, vision_use 39%)
055d0f0 verified
|
Raw
History Blame Contribute Delete
3.96 kB
metadata
license: other
license_name: physicalai-av-derived-research-only
tags:
  - tanitad
  - world-model
  - autonomous-driving
  - planner
extra_gated_prompt: >-
  These weights are trained on NVIDIA PhysicalAI-AV data (TanitAD research
  program). Access is granted per request for research/evaluation use only; you
  agree not to redistribute.
extra_gated_fields:
  Name: text
  Affiliation: text
  Intended use: text

TanitAD β€” REF-B v2 (arch-v2 --refbpatch, step 29,999 / 30,000 FINAL)

Reference arm B of the TanitAD 3-arm study: a hierarchical planner with a from-scratch trained encoder (the "trained-encoder" reference vs REF-A's frozen encoder). This is the v2 milestone-final checkpoint (271.6 M params).

Architecture (v2)

  • Trained ViT encoder: 9-channel, 256 px, patch 16, depth 25, d_model 768 (grad-checkpointed).
  • operative (d 768, depth 6, action_seq 5) + tactical + strategic (4 nav commands) + confidence head + OOD monitor.
  • v2 --refbpatch levers:
    • B1 β€” TIME-anchored tactical decoder (anchor space=time, 128 anchors from a 4096 pool, 4 layers / 8 heads) producing waypoints @ 5/10/15/20 steps (0.5/1/1.5/2 s).
    • B2 β€” proprioceptive ego conditioning [v0, yr0]: ego speed + backward-diff yaw-rate (raw rad/s), added to the nav embedding with ego_dropout 0.5 (anti-shortcut).
    • aux-yaw + aux-accel heads, path heads at fixed distances [2, 5, 10, 20] m, jerk penalty 0.02.

Training

  • Data: PhysicalAI-AV front-wide (phase-0 epcache).
  • Step 29,999 / 30,000 (FINAL). Optimizer AdamW (lr 3e-4, cosine), weight_decay 0.05, warmup 2000, AMP.

Evaluation (TanitEval, held-out PhysicalAI-AV val, n=881 windows; open-loop)

Planner-native trajectory = tactical waypoint heads (direct regression, nav=follow). Bootstrap 95% CI.

horizon ADE (m) CV baseline
0.5 s 0.103
1.0 s 0.217
1.5 s 0.379
2.0 s 0.592 Β± 0.069 0.825
  • fde@2s 1.23 (CV 1.71), miss@2m 0.203 (CV 0.313). Beats the constant-velocity baseline (0.592 vs 0.825, βˆ’28%).
  • By speed (model / CV, ADE@2s): high 0.43 / 0.65 Β· med 0.52 / 0.93 Β· low 0.82 / 0.93 β€” beats CV in every band.
  • By curvature (model / CV, ADE@2s): gentle 0.86 / 1.36 Β· sharp 0.90 / 2.38 (large margin) Β· straight 0.48 / 0.44 (CV is near-optimal on straights).
  • Error is ~79 % longitudinal at 2 s (along-track / speed-timing dominant); becomes lateral-dominant only on sharp curves.
  • Vision use: zeroing the camera frames worsens ADE@2s 0.591 β†’ 0.965 (β‰ˆ39 % of accuracy is vision-driven); the remainder rides the [v0, yr0] ego state.

Milestone / cross-arm A/B (paired, same windows)

  • vs REF-B v2 @20k (prior milestone): 0.591 vs 0.643 β€” v2-30k wins, significant (Ξ”CI [βˆ’0.077, βˆ’0.026]); largest gain on sharp curves (βˆ’0.135 m).
  • vs REF-B v1 (step-6k, ADE@2s 0.868): v2-30k wins decisively (Ξ”CI [βˆ’0.32, βˆ’0.22]); sharp-curve gain βˆ’1.19 m.
  • vs Flagship-4B v1 @30k (grounded world-model): flagship better on open-loop ADE (0.427 vs 0.591).

Generalization (cross-corpus, planner-direct)

corpus kind ADE@2s CV
PhysicalAI-AV in-dist 0.592 0.825
comma2k19 OOD (real highway) 0.515 0.646
Cosmos-Drive-Dreams OOD synthetic 0.495 0.451

The edge over CV shrinks OOD (in-dist βˆ’28 % β†’ comma βˆ’20 % β†’ cosmos +10 %, i.e. below CV). Cosmos numbers are PROVISIONAL (a subset of clips carry a ground-truth timing desync; a rebuilt cache is pending). Absolute OOD ADE is lower mainly because those corpora are geometrically easier (highway / synthetic, low curvature).

Imagination and hierarchy panels are N/A for REF-B (a direct planner with no grounded rollout head). Trained on PhysicalAI-AV derived data. Gated for research/eval use.