Video Classification
Transformers
Safetensors
ttvidt
feature-extraction
video
video-representation-learning
self-supervised-learning
motion
temporal-modeling
dinov3
vision-transformer
custom_code
Eval Results (legacy)
Instructions to use KBlueLeaf/TTVidT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use KBlueLeaf/TTVidT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("video-classification", model="KBlueLeaf/TTVidT", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
|
Download README.md from KBlueLeaf/TTVidT: direct link, hf CLI and curl.
- Browser
- Download file 7.26 kB
-
https://huggingface.co/KBlueLeaf/TTVidT/resolve/main/README.md
- Command line
-
hf download hf://KBlueLeaf/TTVidT/README.md
-
curl -L -o README.md https://huggingface.co/KBlueLeaf/TTVidT/resolve/main/README.md
7.26 kB
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: video-classification | |
| tags: | |
| - video | |
| - video-representation-learning | |
| - self-supervised-learning | |
| - motion | |
| - temporal-modeling | |
| - dinov3 | |
| - vision-transformer | |
| - custom_code | |
| base_model: facebook/dinov3-vitb16-pretrain-lvd1689m | |
| base_model_relation: finetune | |
| datasets: | |
| - nkp37/OpenVid-1M | |
| metrics: | |
| - accuracy | |
| model-index: | |
| - name: TT-VidT (TT3D) | |
| results: | |
| - task: | |
| type: video-classification | |
| name: Frozen attentive probe | |
| dataset: | |
| type: hmdb51 | |
| name: HMDB51 | |
| metrics: | |
| - type: accuracy | |
| value: 25.2 | |
| name: Top-1 accuracy (mean of 3 seeds) | |
| - task: | |
| type: video-classification | |
| name: Frozen attentive probe | |
| dataset: | |
| type: arid | |
| name: ARID | |
| metrics: | |
| - type: accuracy | |
| value: 36.1 | |
| name: Top-1 accuracy (mean of 3 seeds) | |
| - task: | |
| type: video-classification | |
| name: Frozen attentive probe | |
| dataset: | |
| type: iard | |
| name: IARD | |
| metrics: | |
| - type: accuracy | |
| value: 86.0 | |
| name: Top-1 accuracy (mean of 3 seeds) | |
| - task: | |
| type: video-classification | |
| name: Frozen attentive probe | |
| dataset: | |
| type: jester | |
| name: Jester | |
| metrics: | |
| - type: accuracy | |
| value: 72.9 | |
| name: Top-1 accuracy (mean of 3 seeds) | |
| - task: | |
| type: video-classification | |
| name: Frozen attentive probe | |
| dataset: | |
| type: something-something-v2 | |
| name: Something-Something v2 | |
| metrics: | |
| - type: accuracy | |
| value: 25.4 | |
| name: Top-1 accuracy (mean of 3 seeds) | |
| <div align="center"> | |
| # TT-VidT | |
| ### Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining | |
| **NeurIPS 2026 (main track)** | |
| [](https://kohakublueleaf.github.io/TTVidT/) | |
| [](https://huggingface.co/papers/2609.33419) | |
| [](https://arxiv.org/abs/2609.33419) | |
| [](https://github.com/KohakuBlueleaf/TTVidT) | |
| [](https://huggingface.co/KBlueLeaf/TTVidT-decoders) | |
| [](#license) | |
| </div> | |
| <!--  --> | |
|  | |
| **TT-VidT** is a self-supervised video encoder built for *motion*. A DINOv3 ViT-B/16 | |
| processes every frame independently (the appearance path), while a compact | |
| **Temporal Transfer** pathway turns each frame into a few motion tokens that | |
| exchange information across time. This repository holds the pretrained **TT3D** | |
| encoder (195.5M parameters). | |
| ## Model | |
|  | |
| - **Encoder**: 12 DINOv3 ViT-B/16 layers interleaved with 12 Temporal Transfer (TT3D) | |
| layers. Each TT3D layer runs block-causal attention over the frame's K = 8 motion | |
| tokens together with its 4x-downsampled spatial tokens, and writes the result back | |
| to the spatial stream. | |
| - **Pretraining objective**: Diff Compression. A DiT decoder reconstructs every | |
| later frame from the *first* frame's features plus that frame's motion tokens, so | |
| the motion tokens carry what the first frame cannot explain. | |
| | | | | |
| |---|---| | |
| | Parameters | 195.5M (incl. the 85.1M DINOv3 spatial path) | | |
| | Input | 8 RGB frames, 256 x 256, pixels in [-1, 1] (mean = std = 0.5), sampled at 6 fps in pretraining | | |
| | Output | `motion_output`: one 768-d motion embedding per frame, `[B, T, 1, 768]` | | |
| | Weights | fp32 `safetensors`, encoder only | | |
| ## Quick start | |
| **With `transformers` only** (the model code ships in this repository): | |
| ```python | |
| import torch | |
| from transformers import AutoModel | |
| model = AutoModel.from_pretrained("KBlueLeaf/TTVidT", trust_remote_code=True).cuda().eval() | |
| video = torch.rand(1, 8, 3, 256, 256, device="cuda") * 2 - 1 # [B, T, C, H, W] in [-1, 1] | |
| with torch.no_grad(), torch.autocast("cuda", dtype=torch.float16): | |
| motion = model(video).motion_output # [B, T, 1, 768] | |
| ``` | |
| **With the [TT-VidT codebase](https://github.com/KohakuBlueleaf/TTVidT)** (training, | |
| evaluation, feature extraction): | |
| ```python | |
| from ttvidt.hub import load_model | |
| model = load_model("KBlueLeaf/TTVidT", device="cuda") | |
| motion = model.encoder(video).motion_output | |
| ``` | |
| Loading needs no access to the (gated) DINOv3 base weights: every weight is in this | |
| repository. | |
| ## Results | |
| Frozen attentive probe on the motion embeddings, top-1 accuracy (%), mean of 3 seeds, | |
| 8 frames (evaluation protocol of the paper): | |
| | HMDB51 | ARID | IARD | Jester | SSv2 | | |
| |:---:|:---:|:---:|:---:|:---:| | |
| | 25.2 | 36.1 | 86.0 | 72.9 | 25.4 | | |
| For the full study (24 architecture–objective pairs at matched scale, fine-tuning and | |
| diagnostics), see the [paper](https://huggingface.co/papers/2609.33419). | |
| ## Possible downstream uses | |
| The encoder gives a compact sequence of per-frame motion tokens alongside the | |
| DINOv3 appearance features. Some directions we think are worth trying: | |
| - **From image models to video models**: pair an existing image model with the | |
| motion token sequence to get a video model for understanding in the broad sense: | |
| classification, retrieval, captioning, question answering, or any other task. | |
| - **Generation and motion transfer**: use the motion tokens as a conditioning | |
| signal for video generation, or take them from one clip and apply them to another | |
| subject or scene. | |
| > [!TIP] | |
| > Further exploration and feedback are very welcome, and so are attempts at larger | |
| > scale (bigger backbones, more data, longer training). Please open an issue or a | |
| > discussion on [GitHub](https://github.com/KohakuBlueleaf/TTVidT) or in the | |
| > Community tab. | |
| ## Related resources | |
| | Resource | Link | | |
| |---|---| | |
| | Project page | [kohakublueleaf.github.io/TTVidT](https://kohakublueleaf.github.io/TTVidT/) | | |
| | Paper | [huggingface.co/papers/2609.33419](https://huggingface.co/papers/2609.33419) · [arXiv:2609.33419](https://arxiv.org/abs/2609.33419) | | |
| | Source code (training, evaluation, all paper configs) | [github.com/KohakuBlueleaf/TTVidT](https://github.com/KohakuBlueleaf/TTVidT) | | |
| | Pretrained DiT decoders (Diff Compression and the other objectives) | [KBlueLeaf/TTVidT-decoders](https://huggingface.co/KBlueLeaf/TTVidT-decoders) | | |
| ## Files | |
| | File | Content | | |
| |---|---| | |
| | `config.json` | architecture and loader configuration (`transformers` + TT-VidT codebase) | | |
| | `model.safetensors` | fp32 encoder weights | | |
| | `*.py` | encoder code for `trust_remote_code` (needs only `torch` and `transformers`) | | |
| | `assets/` | figures of this card | | |
| ## License | |
| Apache-2.0. The spatial path is initialised from | |
| [DINOv3](https://huggingface.co/facebook/dinov3-vitb16-pretrain-lvd1689m), which is | |
| released under its own license. | |