tdv / README.md
ninaddaithankar's picture
Update README.md
900b81f verified
|
Raw
History Blame Contribute Delete
4.41 kB
---
license: mit
pipeline_tag: image-feature-extraction
tags:
- self-supervised-learning
- vision-transformer
- computer-vision
- video
- temporal-difference
- optical-flow
- stereo-depth
- dense-prediction
arxiv: 2606.15956
---
# TDV: You Don't Need Strong Assumptions — Visual Representation Learning via Temporal Differences
**[Ninad Daithankar](https://github.com/ninaddaithankar)\*, Alexi Gladstone\*, [Yann LeCun](https://yann.lecun.com), Heng Ji**
University of Illinois Urbana-Champaign  ·  New York University
(\*Equal Contribution)
[Paper](https://arxiv.org/abs/2606.15956) | [Project Page](https://temporal-difference-vision.github.io/) | [GitHub](https://github.com/ninaddaithankar/tdv) | [HF Paper](https://huggingface.co/papers/2606.15956)
---
## Overview
TDV is a self-supervised video representation learning method built on a single causal assumption: **the past causes the future**. Rather than relying on strong inductive biases like augmentations, masking, or cropping, TDV jointly trains a frame encoder and a motion encoder such that:
> *current frame representation + encoded motion = next frame representation*
The motion encoder processes RGB differences between consecutive frames, conditioned on the current frame representation via cross-attention. Predictions are supervised using MSE loss against an EMA teacher encoder, with DINO-style categorical cross-entropy to prevent representation collapse.
TDV matches DINO and iBOT on dense segmentation tasks while **surpassing both on optical flow and stereo depth** — tasks that directly reward temporally structured representations.
---
## Checkpoints
All checkpoints are in the `checkpoints/` folder. Optimizer states and training metadata have been stripped — these are inference-ready weights.
All models are pretrained on [Something-Something v2 (SSv2)](https://developer.qualcomm.com/software/ai-datasets/something-something).
| File | Model | Size | Notes |
|---|---|---|---|
| `checkpoints/tdv-base.pth` | TDV ViT-Base | ~1.2 GB | TDV frame encoder (our method) |
| `checkpoints/tdv-small.pth` | TDV ViT-Small | ~400 MB | TDV frame encoder (our method) |
| `checkpoints/dino-base.pth` | DINO ViT-Base | ~739 MB | DINO baseline (teacher weights) |
| `checkpoints/dino-small.pth` | DINO ViT-Small | ~220 MB | DINO baseline (teacher weights) |
| `checkpoints/ibot-base.pth` | iBOT ViT-Base | ~762 MB | iBOT baseline (teacher weights) |
| `checkpoints/ibot-small.pth` | iBOT ViT-Small | ~243 MB | iBOT baseline (teacher weights) |
DINO and iBOT checkpoints are our reproductions trained with online KNN monitoring, used as baselines in the paper.
---
## Usage
### Loading TDV
TDV checkpoints are PyTorch Lightning checkpoints with the optimizer states stripped. Load the `state_dict` directly:
```python
import torch
ckpt = torch.load("checkpoints/tdv-base.ckpt", map_location="cpu")
state_dict = ckpt["state_dict"]
# The frame encoder keys are prefixed with "model.frame_encoder."
# Strip the prefix to load into a standalone ViT:
encoder_state_dict = {
k.replace("model.frame_encoder.", ""): v
for k, v in state_dict.items()
if k.startswith("model.frame_encoder.")
}
```
For the full architecture and eval setup see [github.com/ninaddaithankar/tdv](https://github.com/ninaddaithankar/tdv).
### Loading DINO / iBOT baselines
```python
import torch
ckpt = torch.load("checkpoints/dino-base.pth", map_location="cpu")
teacher_weights = ckpt["teacher"] # final model weights
# student weights also available: ckpt["student"]
```
---
## Results
TDV is evaluated on dense spatial tasks where temporal structure matters most:
- **Optical Flow (MPI-Sintel):** Lower EPE than DINO and iBOT baselines
- **Stereo Depth (SceneFlow):** Lower bad-pixel rates at 0.5px and 1px thresholds
- **Semantic Segmentation (ADE20K / Cityscapes):** Matches DINO and iBOT baselines
See the [paper](https://arxiv.org/abs/2606.15956) for full tables and ablations.
---
## Citation
```bibtex
@misc{daithankar2026dontneedstrongassumptions,
title={You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences},
author={Ninad Daithankar and Alexi Gladstone and Yann LeCun and Heng Ji},
year={2026},
eprint={2606.15956},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2606.15956},
}
```