| --- |
| license: mit |
| pipeline_tag: image-feature-extraction |
| tags: |
| - self-supervised-learning |
| - vision-transformer |
| - computer-vision |
| - video |
| - temporal-difference |
| - optical-flow |
| - stereo-depth |
| - dense-prediction |
| arxiv: 2606.15956 |
| --- |
| |
| # TDV: You Don't Need Strong Assumptions — Visual Representation Learning via Temporal Differences |
|
|
| **[Ninad Daithankar](https://github.com/ninaddaithankar)\*, Alexi Gladstone\*, [Yann LeCun](https://yann.lecun.com), Heng Ji** |
| University of Illinois Urbana-Champaign · New York University |
| (\*Equal Contribution) |
| |
| [Paper](https://arxiv.org/abs/2606.15956) | [Project Page](https://temporal-difference-vision.github.io/) | [GitHub](https://github.com/ninaddaithankar/tdv) | [HF Paper](https://huggingface.co/papers/2606.15956) |
| |
| --- |
| |
| ## Overview |
| |
| TDV is a self-supervised video representation learning method built on a single causal assumption: **the past causes the future**. Rather than relying on strong inductive biases like augmentations, masking, or cropping, TDV jointly trains a frame encoder and a motion encoder such that: |
| |
| > *current frame representation + encoded motion = next frame representation* |
|
|
| The motion encoder processes RGB differences between consecutive frames, conditioned on the current frame representation via cross-attention. Predictions are supervised using MSE loss against an EMA teacher encoder, with DINO-style categorical cross-entropy to prevent representation collapse. |
|
|
| TDV matches DINO and iBOT on dense segmentation tasks while **surpassing both on optical flow and stereo depth** — tasks that directly reward temporally structured representations. |
|
|
| --- |
|
|
| ## Checkpoints |
|
|
| All checkpoints are in the `checkpoints/` folder. Optimizer states and training metadata have been stripped — these are inference-ready weights. |
|
|
| All models are pretrained on [Something-Something v2 (SSv2)](https://developer.qualcomm.com/software/ai-datasets/something-something). |
|
|
| | File | Model | Size | Notes | |
| |---|---|---|---| |
| | `checkpoints/tdv-base.pth` | TDV ViT-Base | ~1.2 GB | TDV frame encoder (our method) | |
| | `checkpoints/tdv-small.pth` | TDV ViT-Small | ~400 MB | TDV frame encoder (our method) | |
| | `checkpoints/dino-base.pth` | DINO ViT-Base | ~739 MB | DINO baseline (teacher weights) | |
| | `checkpoints/dino-small.pth` | DINO ViT-Small | ~220 MB | DINO baseline (teacher weights) | |
| | `checkpoints/ibot-base.pth` | iBOT ViT-Base | ~762 MB | iBOT baseline (teacher weights) | |
| | `checkpoints/ibot-small.pth` | iBOT ViT-Small | ~243 MB | iBOT baseline (teacher weights) | |
|
|
| DINO and iBOT checkpoints are our reproductions trained with online KNN monitoring, used as baselines in the paper. |
|
|
| --- |
|
|
| ## Usage |
|
|
| ### Loading TDV |
|
|
| TDV checkpoints are PyTorch Lightning checkpoints with the optimizer states stripped. Load the `state_dict` directly: |
|
|
| ```python |
| import torch |
| |
| ckpt = torch.load("checkpoints/tdv-base.ckpt", map_location="cpu") |
| state_dict = ckpt["state_dict"] |
| |
| # The frame encoder keys are prefixed with "model.frame_encoder." |
| # Strip the prefix to load into a standalone ViT: |
| encoder_state_dict = { |
| k.replace("model.frame_encoder.", ""): v |
| for k, v in state_dict.items() |
| if k.startswith("model.frame_encoder.") |
| } |
| ``` |
|
|
| For the full architecture and eval setup see [github.com/ninaddaithankar/tdv](https://github.com/ninaddaithankar/tdv). |
|
|
| ### Loading DINO / iBOT baselines |
|
|
| ```python |
| import torch |
| |
| ckpt = torch.load("checkpoints/dino-base.pth", map_location="cpu") |
| teacher_weights = ckpt["teacher"] # final model weights |
| # student weights also available: ckpt["student"] |
| ``` |
|
|
| --- |
|
|
| ## Results |
|
|
| TDV is evaluated on dense spatial tasks where temporal structure matters most: |
|
|
| - **Optical Flow (MPI-Sintel):** Lower EPE than DINO and iBOT baselines |
| - **Stereo Depth (SceneFlow):** Lower bad-pixel rates at 0.5px and 1px thresholds |
| - **Semantic Segmentation (ADE20K / Cityscapes):** Matches DINO and iBOT baselines |
|
|
| See the [paper](https://arxiv.org/abs/2606.15956) for full tables and ablations. |
|
|
| --- |
|
|
| ## Citation |
|
|
| ```bibtex |
| @misc{daithankar2026dontneedstrongassumptions, |
| title={You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences}, |
| author={Ninad Daithankar and Alexi Gladstone and Yann LeCun and Heng Ji}, |
| year={2026}, |
| eprint={2606.15956}, |
| archivePrefix={arXiv}, |
| primaryClass={cs.CV}, |
| url={https://arxiv.org/abs/2606.15956}, |
| } |
| ``` |
|
|