Image Feature Extraction
Transformers
Safetensors
motif_vision
feature-extraction
motif
vision-transformer
self-supervised
video
custom_code
Instructions to use Motif-Technologies/Motif-Vision-Encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Motif-Technologies/Motif-Vision-Encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: mit | |
| library_name: transformers | |
| pipeline_tag: image-feature-extraction | |
| tags: | |
| - motif | |
| - vision-transformer | |
| - self-supervised | |
| - image-feature-extraction | |
| - video | |
| - custom_code | |
| # Motif Vision Encoder | |
| Motif Vision Encoder is a unified image + video self-supervised vision encoder on a ViT | |
| backbone. A single 3D-convolutional tokenizer ingests both modalities β an image is a | |
| 1-frame clip (`T=1`), a video is `T>1` β so the same weights produce dense patch-level | |
| features and a language-aligned global (CLS) representation. | |
| Trained on **~1/3 the data of DINOv3** (0.5B vs 1.7B samples), it still reaches competitive | |
| performance across image and video benchmarks β and leads on DAVIS video tracking. | |
| <p align="center"> | |
| <img src="assets/haaland_attn_blk20.gif" width="480" alt="Point tracking on a video clip: Motif vs V-JEPA 2.1"/> | |
| </p> | |
| <p align="center"><em>Point tracking on a video clip (top: Motif, bottom: V-JEPA 2.1) β a query point propagated across frames by patch-feature cosine similarity. Motif tracks the subject more reliably than V-JEPA 2.1.</em></p> | |
| - **Architecture**: ViT-7B (embed 4096 / depth 40 / heads 32), patch 16, 3D axial RoPE | |
| (`base=100`), SwiGLU FFN, LayerScale, per-head QK-norm, gated attention, 4 register tokens. | |
| - **Tokenizer**: `Conv3d(kernel=stride=(tubelet, patch, patch))` β image `(B,3,H,W)` β `T=1`, | |
| video `(B,T,3,H,W)`. Token layout `[CLS] + [register Γ 4] + [patch Γ N]`. | |
| ## Usage | |
| The model ships a self-contained `modeling_motif_vision_encoder.py`, so it loads with `trust_remote_code=True`. | |
| ### Image | |
| ```python | |
| import torch | |
| from transformers import AutoImageProcessor, AutoModel | |
| from transformers.image_utils import load_image | |
| url = "http://images.cocodataset.org/val2017/000000039769.jpg" | |
| image = load_image(url) | |
| repo = "Motif-Technologies/Motif-Vision-Encoder" | |
| processor = AutoImageProcessor.from_pretrained(repo) | |
| model = AutoModel.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).to("cuda").eval() | |
| inputs = processor(images=image, return_tensors="pt").to(model.device, torch.bfloat16) | |
| with torch.inference_mode(): | |
| outputs = model(**inputs) | |
| outputs.last_hidden_state # (1, 1 + 4 + N, 4096) CLS + registers + patch tokens | |
| outputs.pooler_output # (1, 4096) global (CLS) representation | |
| patch_tokens = outputs.last_hidden_state[:, 5:, :] # (1, N, 4096), N = (H/16)*(W/16) | |
| ``` | |
| The processor resizes the shorter side to 512, center-crops to 512Γ512, and normalizes with | |
| ImageNet mean/std (BICUBIC). `H`/`W` must be multiples of 16. | |
| ### Video | |
| An image is a 1-frame clip; a video is the same call with a `(B, T, 3, H, W)` tensor. Apply the | |
| same per-frame transform (resize β center-crop β ImageNet norm) and stack over time: | |
| ```python | |
| import torch | |
| video = torch.randn(1, 8, 3, 256, 256, device="cuda", dtype=torch.bfloat16) # (B, T, 3, H, W) | |
| with torch.inference_mode(): | |
| outputs = model(pixel_values=video) | |
| ``` | |
| ## Model details | |
| <p align="center"> | |
| <img src="assets/architecture.png" width="820" alt="Motif Vision Encoder architecture: image and video inputs, patch embedding, 40-block transformer stack, and transformer block internals"/> | |
| </p> | |
| | | | | |
| |---|---| | |
| | Backbone | ViT-7B, patch 16, embed 4096, depth 40, heads 32, SwiGLU | | |
| | Register tokens | 4 | | |
| | Position encoding | 3D axial RoPE (T,H,W), `base=100.0` | | |
| | Video tokenizer | 3D Conv, tubelet size 2 | | |
| | Precision | bf16 weights | | |
| | Training | DINO + iBOT + KoLeo self-distillation, Gram anchoring, contrastive caption alignment | | |
| | Training data | ~0.47B samples β 448.6M images (96%) + 18.5M video clips (4%) | | |
| Outputs (`BaseModelOutputWithPooling`): `last_hidden_state` `(B, 1+4+N, 4096)`, | |
| `pooler_output` `(B, 4096)`. | |
| ## Evaluation | |
| Compared against the strongest publicly reported self-supervised / vision backbones. Higher is | |
| better for every column. Best comparable value per column in bold, second best <u>underlined</u>. | |
| DAVIS S/M/L follow the DINOv3 protocol (J&F-mean at video short side 420/480, 840/960, | |
| 1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its | |
| single-resolution (S) figure is available. | |
| | Model | Training<br>data | DAVIS S<br>J&F β | DAVIS M<br>J&F β | DAVIS L<br>J&F β | ImageNet-1K<br>lin. probe β | ADE20K<br>mIoU β | K400 β | | |
| |---|---|---|---|---|---|---|---| | |
| | **Motif Vision Encoder** | 0.5B | **73.8** | **80.4** | **83.4** | 87.4 | <u>52.0</u> | 87.4 | | |
| | DINOv3 | 1.7B | <u>71.1</u> | <u>79.7</u> | <u>83.3</u> | 88.4 | **55.9** | <u>87.8</u> | | |
| | Web-DINO | 2B | 57.2 | 65.8 | 69.5 | 85.9 | 42.7 | 86.8 | | |
| | PEcore | 5.4B | 48.2 | 53.1 | 49.8 | **89.3** | 38.9 | **87.9** | | |
| | SigLIP2 | 10B | 56.1 | 62.3 | 62.9 | <u>89.1</u> | 45.4 | 86.9 | | |
| | OpenCLIP | 2B | β | β | β | β | β | β | | |
| | V-JEPA 2.1 | 0.022B | 69.0 | β | β | 85.5 | 47.9 | 87.7 | | |
| Protocol: DINOv3-style linear/attentive probes for image tasks; V-JEPA 2-style protocol for | |
| video. Baseline DAVIS / ADE20K / K400 figures are taken from the DINOv3 technical report's | |
| unified evaluation (Tab. 3, 5, 6) and ImageNet from Tab. 7; OpenCLIP is not in that report and | |
| its benchmarks are not reported under a comparable protocol. | |
| Motif is state of the art on DAVIS video tracking at every resolution (74.0 / 80.5 / 83.5 J&F) | |
| and stays competitive on the other image and video benchmarks, using roughly 1/3 of DINOv3's | |
| training data (~0.5B samples). | |
| <p align="center"> | |
| <img src="assets/davis_mask_propagation.gif" width="820" alt="Mask propagation: ground truth vs DINOv3 vs Motif"/> | |
| </p> | |
| <p align="center"> | |
| <img src="assets/dense_attention_comparison.png" width="820" alt="Dense attention and feature-similarity comparison across Motif, DINOv3, V-JEPA 2.1, and SigLIP2"/> | |
| </p> | |
| <p align="center"><em>Dense features on a single image (768px). Columns: query point, CLS attention, query-point attention, patch-feature cosine similarity. Motif and DINOv3 keep attention and similarity tightly localized on the queried object, while V-JEPA 2.1 and SigLIP2 are noticeably noisier.</em></p> | |
| ## License | |
| Released under the **MIT License** (see `LICENSE`). The model was trained on data governed by the | |
| respective dataset licenses; downstream users are responsible for compliance with those terms. | |