File size: 4,412 Bytes
77aa298
 
900b81f
77aa298
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
383817a
 
77aa298
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
---
license: mit
pipeline_tag: image-feature-extraction
tags:
  - self-supervised-learning
  - vision-transformer
  - computer-vision
  - video
  - temporal-difference
  - optical-flow
  - stereo-depth
  - dense-prediction
arxiv: 2606.15956
---

# TDV: You Don't Need Strong Assumptions — Visual Representation Learning via Temporal Differences

**[Ninad Daithankar](https://github.com/ninaddaithankar)\*, Alexi Gladstone\*, [Yann LeCun](https://yann.lecun.com), Heng Ji**
University of Illinois Urbana-Champaign  ·  New York University
(\*Equal Contribution)

[Paper](https://arxiv.org/abs/2606.15956) | [Project Page](https://temporal-difference-vision.github.io/) | [GitHub](https://github.com/ninaddaithankar/tdv) | [HF Paper](https://huggingface.co/papers/2606.15956)

---

## Overview

TDV is a self-supervised video representation learning method built on a single causal assumption: **the past causes the future**. Rather than relying on strong inductive biases like augmentations, masking, or cropping, TDV jointly trains a frame encoder and a motion encoder such that:

> *current frame representation + encoded motion = next frame representation*

The motion encoder processes RGB differences between consecutive frames, conditioned on the current frame representation via cross-attention. Predictions are supervised using MSE loss against an EMA teacher encoder, with DINO-style categorical cross-entropy to prevent representation collapse.

TDV matches DINO and iBOT on dense segmentation tasks while **surpassing both on optical flow and stereo depth** — tasks that directly reward temporally structured representations.

---

## Checkpoints

All checkpoints are in the `checkpoints/` folder. Optimizer states and training metadata have been stripped — these are inference-ready weights.

All models are pretrained on [Something-Something v2 (SSv2)](https://developer.qualcomm.com/software/ai-datasets/something-something).

| File | Model | Size | Notes |
|---|---|---|---|
| `checkpoints/tdv-base.pth` | TDV ViT-Base | ~1.2 GB | TDV frame encoder (our method) |
| `checkpoints/tdv-small.pth` | TDV ViT-Small | ~400 MB | TDV frame encoder (our method) |
| `checkpoints/dino-base.pth` | DINO ViT-Base | ~739 MB | DINO baseline (teacher weights) |
| `checkpoints/dino-small.pth` | DINO ViT-Small | ~220 MB | DINO baseline (teacher weights) |
| `checkpoints/ibot-base.pth` | iBOT ViT-Base | ~762 MB | iBOT baseline (teacher weights) |
| `checkpoints/ibot-small.pth` | iBOT ViT-Small | ~243 MB | iBOT baseline (teacher weights) |

DINO and iBOT checkpoints are our reproductions trained with online KNN monitoring, used as baselines in the paper.

---

## Usage

### Loading TDV

TDV checkpoints are PyTorch Lightning checkpoints with the optimizer states stripped. Load the `state_dict` directly:

```python
import torch

ckpt = torch.load("checkpoints/tdv-base.ckpt", map_location="cpu")
state_dict = ckpt["state_dict"]

# The frame encoder keys are prefixed with "model.frame_encoder."
# Strip the prefix to load into a standalone ViT:
encoder_state_dict = {
    k.replace("model.frame_encoder.", ""): v
    for k, v in state_dict.items()
    if k.startswith("model.frame_encoder.")
}
```

For the full architecture and eval setup see [github.com/ninaddaithankar/tdv](https://github.com/ninaddaithankar/tdv).

### Loading DINO / iBOT baselines

```python
import torch

ckpt = torch.load("checkpoints/dino-base.pth", map_location="cpu")
teacher_weights = ckpt["teacher"]  # final model weights
# student weights also available: ckpt["student"]
```

---

## Results

TDV is evaluated on dense spatial tasks where temporal structure matters most:

- **Optical Flow (MPI-Sintel):** Lower EPE than DINO and iBOT baselines
- **Stereo Depth (SceneFlow):** Lower bad-pixel rates at 0.5px and 1px thresholds
- **Semantic Segmentation (ADE20K / Cityscapes):** Matches DINO and iBOT baselines

See the [paper](https://arxiv.org/abs/2606.15956) for full tables and ablations.

---

## Citation

```bibtex
@misc{daithankar2026dontneedstrongassumptions,
      title={You Don't Need Strong Assumptions: Visual Representation Learning via Temporal Differences},
      author={Ninad Daithankar and Alexi Gladstone and Yann LeCun and Heng Ji},
      year={2026},
      eprint={2606.15956},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2606.15956},
}
```