TTVidT / README.md
KBlueLeaf's picture
Update README.md
3c7a789 verified
|
Raw History Blame Contribute Delete
7.26 kB
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: video-classification
tags:
  - video
  - video-representation-learning
  - self-supervised-learning
  - motion
  - temporal-modeling
  - dinov3
  - vision-transformer
  - custom_code
base_model: facebook/dinov3-vitb16-pretrain-lvd1689m
base_model_relation: finetune
datasets:
  - nkp37/OpenVid-1M
metrics:
  - accuracy
model-index:
  - name: TT-VidT (TT3D)
    results:
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: hmdb51
          name: HMDB51
        metrics:
          - type: accuracy
            value: 25.2
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: arid
          name: ARID
        metrics:
          - type: accuracy
            value: 36.1
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: iard
          name: IARD
        metrics:
          - type: accuracy
            value: 86
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: jester
          name: Jester
        metrics:
          - type: accuracy
            value: 72.9
            name: Top-1 accuracy (mean of 3 seeds)
      - task:
          type: video-classification
          name: Frozen attentive probe
        dataset:
          type: something-something-v2
          name: Something-Something v2
        metrics:
          - type: accuracy
            value: 25.4
            name: Top-1 accuracy (mean of 3 seeds)

TT-VidT

Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining

NeurIPS 2026 (main track)

Project Page Paper arXiv Code Decoders License

image

TT-VidT is a self-supervised video encoder built for motion. A DINOv3 ViT-B/16 processes every frame independently (the appearance path), while a compact Temporal Transfer pathway turns each frame into a few motion tokens that exchange information across time. This repository holds the pretrained TT3D encoder (195.5M parameters).

Model

TT-VidT architecture

  • Encoder: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D) layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion tokens together with its 4x-downsampled spatial tokens, and writes the result back to the spatial stream.
  • Pretraining objective: Diff Compression. A DiT decoder reconstructs every later frame from the first frame's features plus that frame's motion tokens, so the motion tokens carry what the first frame cannot explain.
Parameters 195.5M (incl. the 85.1M DINOv3 spatial path)
Input 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining
Output motion_output: one 768-d motion embedding per frame, [B, T, 1, 768]
Weights fp32 safetensors, encoder only

Quick start

With transformers only (the model code ships in this repository):

import torch
from transformers import AutoModel

model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval()

video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1    # [B, T, C, H, W] in [-1, 1]
with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16):
    motion = model(video).motion_output                           # [B, T, 1, 768]

With the TT-VidT codebase (training, evaluation, feature extraction):

from ttvidt.hub import load_model

model = load_model("KBlueLeaf/TTVidT", device="cuda")
motion = model.encoder(video).motion_output

Loading needs no access to the (gated) DINOv3 base weights: every weight is in this repository.

Results

Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds, 8 frames (evaluation protocol of the paper):

HMDB51 ARID IARD Jester SSv2
25.2 36.1 86.0 72.9 25.4

For the full study (24 architecture–objective pairs at matched scale, fine-tuning and diagnostics), see the paper.

Possible downstream uses

The encoder gives a compact sequence of per-frame motion tokens alongside the DINOv3 appearance features. Some directions we think are worth trying:

  • From image models to video models: pair an existing image model with the motion token sequence to get a video model for understanding in the broad sense: classification, retrieval, captioning, question answering, or any other task.
  • Generation and motion transfer: use the motion tokens as a conditioning signal for video generation, or take them from one clip and apply them to another subject or scene.

Further exploration and feedback are very welcome, and so are attempts at larger scale (bigger backbones, more data, longer training). Please open an issue or a discussion on GitHub or in the Community tab.

Related resources

Resource Link
Project page kohakublueleaf.github.io/TTVidT
Paper huggingface.co/papers/2609.33419 · arXiv:2609.33419
Source code (training, evaluation, all paper configs) github.com/KohakuBlueleaf/TTVidT
Pretrained DiT decoders (Diff Compression and the other objectives) KBlueLeaf/TTVidT-decoders

Files

File Content
config.json architecture and loader configuration (transformers + TT-VidT codebase)
model.safetensors fp32 encoder weights
*.py encoder code for trust_remote_code (needs only torch and transformers)
assets/ figures of this card

License

Apache-2.0. The spatial path is initialised from DINOv3, which is released under its own license.