HOI-JEPA / protocols /DOWNSTREAM_PROTOCOLS.md
YunzeLiu's picture
Publish metadata official release files
022e2f6 verified
|
Raw
History Blame Contribute Delete
5.92 kB
# Downstream Decision Protocols
This document freezes the evaluation boundary for the three downstream tasks
used to test whether HOI-JEPA's latent forecasts are useful beyond latent L1.
The world model is never fine-tuned on a downstream label.
## 1. DROID interaction-success prediction
### Data boundary
- Source label: DROID's episode-level `is_episode_successful` field.
- Eligible source episodes: 95,636 of 95,658 declared episodes. The 22
excluded episodes have no usable transition.
- Probe train set: all 86,072 episodes in `train_full` (70,968 success,
15,104 failure).
- Probe test set: all 9,564 episodes in `heldout` (7,887 success, 1,677
failure).
- The split and label sidecar are hash-pinned in
`data/droid/heldout_v1/outcome_audit.json`.
### Inputs and probes
For each horizon H in {1, 2, 4}, a frozen model observes only the first K-H
visual bins and autoregressively predicts the last H bins. We fit the same
standardized, class-balanced, averaged SGD logistic classifier to five fixed
feature sets:
1. `context`: last and mean visible latent;
2. `action`: summary of the future command sequence;
3. `context_action`: their concatenation;
4. `forecast`: `context_action` plus the predicted endpoint, prediction
delta, and learned action gate;
5. `oracle`: `context_action` plus the true endpoint and true delta.
The `forecast` probe never receives a future visual latent. The `oracle` is an
upper bound and is not a deployable model. We report AUROC, AUPRC, balanced
accuracy, Brier score, NLL, and ECE. The primary comparison is forecast versus
context+action at H=4.
Because DROID success is imbalanced, we also construct a 1,677-pair hard test
set. Each failure is paired with a unique success while controlling language
availability and episode-length bin, then minimizing distance between initial
visual latents. Matching never uses a model prediction.
## 2. DROID image-goal candidate-action ranking
### Query construction
- Queries: every successful held-out episode (7,887 total).
- Context: the first four of eight causal visual bins.
- Goal: the episode's final image latent.
- Candidates: 16 four-step action plans. One is the recorded plan; 15 are
plans from other held-out episodes.
Hard negatives are drawn from the 512 nearest visible contexts, prioritize
the same language instruction when available, and are matched to the recorded
plan's standardized action statistics. Failure plans are selected first (up
to five before filling the remaining candidates). Crucially, candidate
construction never reads the query's goal latent. Candidate order is
deterministically randomized with seed 20260707, and every method must match
the published candidate-set SHA-256.
For every candidate, the frozen world model starts from the query context,
rolls out that candidate's commands, and scores
`L1(predicted endpoint latent, held-out image-goal latent)`.
Endpoint R@1, R@5, MRR, mean rank, and pairwise win rate are primary; full
trajectory distance is secondary. Pairwise energy ties receive half credit.
Context-only models are an empirical chance control because they cannot
distinguish the randomly ordered candidate plans. Residual MLP context/action
forecasters are input-matched learned baselines.
This is an offline goal-consistency test, not closed-loop control and not an
official DROID benchmark. Some hard negatives may be alternative valid plans;
that false-negative risk makes absolute recall conservative.
## 3. EPIC-KITCHENS future-action anticipation
This project-defined transfer protocol uses the existing EPIC-KITCHENS-100
validation subset of 1,000 annotated actions from ten videos, ten participants,
and 67 verb classes. It is not the full official EPIC anticipation benchmark.
For an action beginning at time t, four causal observation bins cover
`[t-6 s, t-1 s]`; the final one-second gap enforces anticipation rather than
recognition. Sixteen frames are sampled in each bin. The target action segment
is encoded only for an `oracle` upper bound and never enters a deployable
feature set.
The frozen EgoDex-trained context-only HOI-JEPA predicts a fifth latent without
an action label. A leave-one-participant-out linear classifier compares:
1. visible context;
2. persistence endpoint;
3. linear-extrapolation endpoint;
4. HOI-JEPA forecast endpoint;
5. true target endpoint (`oracle`).
Persistence, linear, forecast, and oracle features have identical dimension.
We report top-1/top-5 accuracy and mean top-1/top-5 recall, participant-cluster
bootstrap intervals, paired forecast deltas, and the fraction of test examples
whose verb appeared in that fold's training participants. Seen-class-only
metrics are reported separately.
The ten source videos are pinned to HuggingFace snapshot
`a7c3f50f6efbeb8c31ed03f2e699afa93f0615c0` and byte-audited. One upstream
video contains damaged H.264 NAL packets. If Decord's batch worker stops on a
damaged packet, extraction reopens a single-thread decoder and substitutes the
nearest decodable frame **within the same temporal bin**. It never crosses the
one-second observation gap or the action-segment boundary. Every affected
narration ID, substitution count, and maximum frame offset is stored in the
feature payload and propagated to the aggregate result; no action is silently
dropped because of decoder damage.
## Reproducibility rules
- Learned conditions use seeds 42, 43, and 44.
- Raw per-episode IDs, energies, and probabilities remain local and are
Git-ignored; aggregate JSON and paired-bootstrap outputs are published.
- All jobs are serialized. GPU stages request one eight-GPU node from
`batch_block1,batch_block3,batch_block4`; download is a one-GPU, one-worker
stage because those partitions reject CPU-only submissions.
- Every job is requeueable at the four-hour boundary. Candidate and downstream
outputs are atomically written and skipped on restart.