Image Feature Extraction
Transformers
Safetensors
motif_vision
feature-extraction
motif
vision-transformer
self-supervised
video
custom_code
Instructions to use Motif-Technologies/Motif-Vision-Encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Motif-Technologies/Motif-Vision-Encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
docs(README): Training-data column + corrected ImageNet numbers
Browse filesReplace Publisher column with Training data (Motif 0.5B, DINOv3 1.7B, DINOv2 142M, V-JEPA 22M, SigLIP2 10B); correct ImageNet-1K lin. probe (DINOv3 88.4, DINOv2 87.3, SigLIP2 89.1); move ImageNet bold to SigLIP2; update narrative bullets accordingly.
README.md
CHANGED
|
@@ -94,19 +94,22 @@ DAVIS S/M/L follow the DINOv3 protocol (J&F-mean at video short side 420/480, 84
|
|
| 94 |
1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
|
| 95 |
single-resolution (S) figure is available.
|
| 96 |
|
| 97 |
-
| Model |
|
| 98 |
|---|---|---|---|---|---|---|---|---|
|
| 99 |
-
| **Motif Vision Encoder** |
|
| 100 |
-
| DINOv3 |
|
| 101 |
-
| DINOv2 |
|
| 102 |
-
| V-JEPA 2.1 |
|
| 103 |
-
| SigLIP2 |
|
| 104 |
|
| 105 |
- **DAVIS video segmentation (J&F)** β leads at every resolution: **74.0 (S) / 80.5 (M) /
|
| 106 |
83.5 (L)**, surpassing DINOv3 7B (71.1 / 79.7 / 83.3) across the board, reflecting the
|
| 107 |
encoder's dense, temporally-coherent patch features on video.
|
| 108 |
-
- **
|
| 109 |
-
|
|
|
|
|
|
|
|
|
|
| 110 |
- **Data efficiency** β these results come from **~0.47B training samples** (448.6M images +
|
| 111 |
18.5M video clips), roughly **3.6Γ less data than DINOv3**, which is trained on LVD-1689M
|
| 112 |
(1,689M images). Despite the smaller corpus β and with only ~4% of samples being video β the
|
|
|
|
| 94 |
1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
|
| 95 |
single-resolution (S) figure is available.
|
| 96 |
|
| 97 |
+
| Model | Training<br>data | DAVIS S<br>J&F β | DAVIS M<br>J&F β | DAVIS L<br>J&F β | ImageNet-1K<br>lin. probe β | ADE20K<br>mIoU β | K400 β | KITTI<br>depth MSE β |
|
| 98 |
|---|---|---|---|---|---|---|---|---|
|
| 99 |
+
| **Motif Vision Encoder** | 0.5B | **74.0** | **80.5** | **83.5** | 87.2 | 52.0 | *in progress* | *in progress* |
|
| 100 |
+
| DINOv3 | 1.7B | 71.1 | 79.7 | 83.3 | 88.4 | **55.9** | **87.8** | **2.3** |
|
| 101 |
+
| DINOv2 | 142M | 63.9 | 73.6 | 76.6 | 87.3 | 49.0 | 84.4 | β |
|
| 102 |
+
| V-JEPA 2.1 | 22M | 69.0 | β | β | 85.5 | 47.9 | 87.7 | 3.x |
|
| 103 |
+
| SigLIP2 | 10B | 56.1 | 62.3 | 62.9 | **89.1** | 45.4 | 86.9 | β |
|
| 104 |
|
| 105 |
- **DAVIS video segmentation (J&F)** β leads at every resolution: **74.0 (S) / 80.5 (M) /
|
| 106 |
83.5 (L)**, surpassing DINOv3 7B (71.1 / 79.7 / 83.3) across the board, reflecting the
|
| 107 |
encoder's dense, temporally-coherent patch features on video.
|
| 108 |
+
- **ADE20K semantic segmentation (52.0 mIoU)** is second only to DINOv3 7B (55.9), well ahead
|
| 109 |
+
of DINOv2, V-JEPA 2.1, and SigLIP2.
|
| 110 |
+
- **ImageNet-1K linear probe (87.2)** is on par with DINOv2 (87.3) and DINOv3 (88.4), trailing
|
| 111 |
+
the contrastively-trained SigLIP2 (89.1) β notable given Motif sees **0.5B** samples versus
|
| 112 |
+
SigLIP2's **10B** imageβtext pairs and DINOv3's **1.7B** images.
|
| 113 |
- **Data efficiency** β these results come from **~0.47B training samples** (448.6M images +
|
| 114 |
18.5M video clips), roughly **3.6Γ less data than DINOv3**, which is trained on LVD-1689M
|
| 115 |
(1,689M images). Despite the smaller corpus β and with only ~4% of samples being video β the
|