--- license: apache-2.0 library_name: transformers pipeline_tag: video-classification tags: - video - video-representation-learning - self-supervised-learning - motion - temporal-modeling - dinov3 - vision-transformer - custom_code base_model: facebook/dinov3-vitb16-pretrain-lvd1689m base_model_relation: finetune datasets: - nkp37/OpenVid-1M metrics: - accuracy model-index: - name: TT-VidT (TT3D) results: - task: type: video-classification name: Frozen attentive probe dataset: type: hmdb51 name: HMDB51 metrics: - type: accuracy value: 25.2 name: Top-1 accuracy (mean of 3 seeds) - task: type: video-classification name: Frozen attentive probe dataset: type: arid name: ARID metrics: - type: accuracy value: 36.1 name: Top-1 accuracy (mean of 3 seeds) - task: type: video-classification name: Frozen attentive probe dataset: type: iard name: IARD metrics: - type: accuracy value: 86.0 name: Top-1 accuracy (mean of 3 seeds) - task: type: video-classification name: Frozen attentive probe dataset: type: jester name: Jester metrics: - type: accuracy value: 72.9 name: Top-1 accuracy (mean of 3 seeds) - task: type: video-classification name: Frozen attentive probe dataset: type: something-something-v2 name: Something-Something v2 metrics: - type: accuracy value: 25.4 name: Top-1 accuracy (mean of 3 seeds) ---
# TT-VidT ### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining **NeurIPS 2026 (main track)** [![Project Page](https://img.shields.io/badge/Project-Page-green)](https://kohakublueleaf.github.io/TTVidT/) [![Paper](https://img.shields.io/badge/🤗%20Paper-2609.33419-yellow)](https://huggingface.co/papers/2609.33419) [![arXiv](https://img.shields.io/badge/arXiv-2609.33419-b31b1b)](https://arxiv.org/abs/2609.33419) [![Code](https://img.shields.io/badge/GitHub-KohakuBlueleaf%2FTTVidT-181717?logo=github)](https://github.com/KohakuBlueleaf/TTVidT) [![Decoders](https://img.shields.io/badge/🤗%20Decoders-TTVidT--decoders-orange)](https://huggingface.co/KBlueLeaf/TTVidT-decoders) [![License](https://img.shields.io/badge/License-Apache%202.0-blue)](#license)
![image](https://cdn-uploads.huggingface.co/production/uploads/630593e2fca1d8d92b81d2a1/JKhgNd-K2wnoaAp1fYwR2.png) **TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16 processes every frame independently (the appearance path), while a compact **Temporal Transfer** pathway turns each frame into a few motion tokens that exchange information across time. This repository holds the pretrained **TT3D** encoder (195.5M parameters). ## Model ![TT-VidT architecture](assets/architecture.png) - **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D) layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion tokens together with its 4x-downsampled spatial tokens, and writes the result back to the spatial stream. - **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every later frame from the *first* frame's features plus that frame's motion tokens, so the motion tokens carry what the first frame cannot explain. | | | |---|---| | Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) | | Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining | | Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` | | Weights | fp32 `safetensors`, encoder only | ## Quick start **With `transformers` only** (the model code ships in this repository): ```python import torch from transformers import AutoModel model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval() video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1] with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16): motion = model(video).motion_output # [B, T, 1, 768] ``` **With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training, evaluation, feature extraction): ```python from ttvidt.hub import load_model model = load_model("KBlueLeaf/TTVidT", device="cuda") motion = model.encoder(video).motion_output ``` Loading needs no access to the (gated) DINOv3 base weights: every weight is in this repository. ## Results Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds, 8 frames (evaluation protocol of the paper): | HMDB51 | ARID | IARD | Jester | SSv2 | |:---:|:---:|:---:|:---:|:---:| | 25.2 | 36.1 | 86.0 | 72.9 | 25.4 | For the full study (24 architecture–objective pairs at matched scale, fine-tuning and diagnostics), see the [paper](https://huggingface.co/papers/2609.33419). ## Possible downstream uses The encoder gives a compact sequence of per-frame motion tokens alongside the DINOv3 appearance features. Some directions we think are worth trying: - **From image models to video models**: pair an existing image model with the motion token sequence to get a video model for understanding in the broad sense: classification, retrieval, captioning, question answering, or any other task. - **Generation and motion transfer**: use the motion tokens as a conditioning signal for video generation, or take them from one clip and apply them to another subject or scene. > [!TIP] > Further exploration and feedback are very welcome, and so are attempts at larger > scale (bigger backbones, more data, longer training). Please open an issue or a > discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the > Community tab. ## Related resources | Resource | Link | |---|---| | Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) | | Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) | | Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) | | Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) | ## Files | File | Content | |---|---| | `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) | | `model.safetensors` | fp32 encoder weights | | `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) | | `assets/` | figures of this card | ## License Apache-2.0. The spatial path is initialised from [DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is released under its own license.