---
license: apache-2.0
library_name: transformers
pipeline_tag: video-classification
tags:
- video
- video-representation-learning
- self-supervised-learning
- motion
- temporal-modeling
- dinov3
- vision-transformer
- custom_code
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
base_model_relation: finetune
datasets:
- nkp37/OpenVid-1M
metrics:
- accuracy
model-index:
- name: TT-VidT (TT3D)
results:
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: hmdb51
name: HMDB51
metrics:
- type: accuracy
value: 25.2
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: arid
name: ARID
metrics:
- type: accuracy
value: 36.1
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: iard
name: IARD
metrics:
- type: accuracy
value: 86.0
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: jester
name: Jester
metrics:
- type: accuracy
value: 72.9
name: Top-1 accuracy (mean of 3 seeds)
- task:
type: video-classification
name: Frozen attentive probe
dataset:
type: something-something-v2
name: Something-Something v2
metrics:
- type: accuracy
value: 25.4
name: Top-1 accuracy (mean of 3 seeds)
---
# TT-VidT
### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining
**NeurIPS 2026 (main track)**
[](https://kohakublueleaf.github.io/TTVidT/)
[](https://huggingface.co/papers/2609.33419)
[](https://arxiv.org/abs/2609.33419)
[](https://github.com/KohakuBlueleaf/TTVidT)
[](https://huggingface.co/KBlueLeaf/TTVidT-decoders)
[](#license)

**TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16
processes every frame independently (the appearance path), while a compact
**Temporal Transfer** pathway turns each frame into a few motion tokens that
exchange information across time. This repository holds the pretrained **TT3D**
encoder (195.5M parameters).
## Model

- **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D)
layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion
tokens together with its 4x-downsampled spatial tokens, and writes the result back
to the spatial stream.
- **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every
later frame from the *first* frame's features plus that frame's motion tokens, so
the motion tokens carry what the first frame cannot explain.
| | |
|---|---|
| Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) |
| Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining |
| Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` |
| Weights | fp32 `safetensors`, encoder only |
## Quick start
**With `transformers` only** (the model code ships in this repository):
```python
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()
video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1]
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
motion = model(video).motion_output # [B, T, 1, 768]
```
**With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training,
evaluation, feature extraction):
```python
from ttvidt.hub import load_model
model = load_model("KBlueLeaf/TTVidT", device="cuda")
motion = model.encoder(video).motion_output
```
Loading needs no access to the (gated) DINOv3 base weights: every weight is in this
repository.
## Results
Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds,
8 frames (evaluation protocol of the paper):
| HMDB51 | ARID | IARD | Jester | SSv2 |
|:---:|:---:|:---:|:---:|:---:|
| 25.2 | 36.1 | 86.0 | 72.9 | 25.4 |
For the full study (24 architecture–objective pairs at matched scale, fine-tuning and
diagnostics), see the [paper](https://huggingface.co/papers/2609.33419).
## Possible downstream uses
The encoder gives a compact sequence of per-frame motion tokens alongside the
DINOv3 appearance features. Some directions we think are worth trying:
- **From image models to video models**: pair an existing image model with the
motion token sequence to get a video model for understanding in the broad sense:
classification, retrieval, captioning, question answering, or any other task.
- **Generation and motion transfer**: use the motion tokens as a conditioning
signal for video generation, or take them from one clip and apply them to another
subject or scene.
> [!TIP]
> Further exploration and feedback are very welcome, and so are attempts at larger
> scale (bigger backbones, more data, longer training). Please open an issue or a
> discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the
> Community tab.
## Related resources
| Resource | Link |
|---|---|
| Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) |
| Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) |
| Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) |
| Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) |
## Files
| File | Content |
|---|---|
| `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) |
| `model.safetensors` | fp32 encoder weights |
| `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) |
| `assets/` | figures of this card |
## License
Apache-2.0. The spatial path is initialised from
[DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is
released under its own license.