Image Feature Extraction
Transformers
Safetensors
motif_vision
feature-extraction
motif
vision-transformer
self-supervised
video
custom_code
Instructions to use Motif-Technologies/Motif-Vision-Encoder with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Motif-Technologies/Motif-Vision-Encoder with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Motif-Technologies/Motif-Vision-Encoder", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
docs(README): underline second-best value per column
Browse filesAdd <u>underline</u> to the runner-up in each column (comparable values only); document convention in caption.
README.md
CHANGED
|
@@ -96,8 +96,8 @@ Outputs (`BaseModelOutputWithPooling`): `last_hidden_state` `(B, 1+4+N, 4096)`,
|
|
| 96 |
|
| 97 |
Compared against the strongest publicly reported self-supervised / vision backbones. Higher is
|
| 98 |
better for every column **except KITTI depth** (lower is better). Best **comparable** value per
|
| 99 |
-
column in **bold**; the footnoted figure (β‘) is measured under a
|
| 100 |
-
excluded from the
|
| 101 |
|
| 102 |
DAVIS S/M/L follow the DINOv3 protocol (J&F-mean at video short side 420/480, 840/960,
|
| 103 |
1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
|
|
@@ -105,12 +105,12 @@ single-resolution (S) figure is available.
|
|
| 105 |
|
| 106 |
| Model | Training<br>data | DAVIS S<br>J&F β | DAVIS M<br>J&F β | DAVIS L<br>J&F β | ImageNet-1K<br>lin. probe β | ADE20K<br>mIoU β | K400 β | KITTI<br>depth MSE β |
|
| 107 |
|---|---|---|---|---|---|---|---|---|
|
| 108 |
-
| **Motif Vision Encoder** | 0.5B | **74.0** | **80.5** | **83.5** | 87.2 | 52.0 | 87.4 | *in progress* |
|
| 109 |
-
| DINOv3 | 1.7B | 71.1 | 79.7 | 83.3 | 88.4 | **55.9** | 87.8 | **2.3** |
|
| 110 |
| PEcore | 5.4B | 48.2 | 53.1 | 49.8 | **89.3** | 38.9 | **87.9** | 4.1 |
|
| 111 |
-
| SigLIP2 | 10B | 56.1 | 62.3 | 62.9 | 89.1 | 45.4 | 86.9 | β |
|
| 112 |
| OpenCLIP | 2B | β | β | β | β | β | β | β |
|
| 113 |
-
| V-JEPA 2.1 | 0.022B | 69.0 | β | β | 85.5 | 47.9 | 87.7 | 2.5 |
|
| 114 |
| VideoMAEv2 | 1.35M | β | β | β | β | β | 90.0<sup>β‘</sup> | β |
|
| 115 |
|
| 116 |
<sub>β‘ VideoMAEv2 ViT-g: K400 top-1 after **full fine-tuning**, not the frozen attentive-probe protocol used for the other rows.</sub>
|
|
|
|
| 96 |
|
| 97 |
Compared against the strongest publicly reported self-supervised / vision backbones. Higher is
|
| 98 |
better for every column **except KITTI depth** (lower is better). Best **comparable** value per
|
| 99 |
+
column in **bold**, second best <u>underlined</u>; the footnoted figure (β‘) is measured under a
|
| 100 |
+
different protocol and is excluded from the ranking.
|
| 101 |
|
| 102 |
DAVIS S/M/L follow the DINOv3 protocol (J&F-mean at video short side 420/480, 840/960,
|
| 103 |
1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
|
|
|
|
| 105 |
|
| 106 |
| Model | Training<br>data | DAVIS S<br>J&F β | DAVIS M<br>J&F β | DAVIS L<br>J&F β | ImageNet-1K<br>lin. probe β | ADE20K<br>mIoU β | K400 β | KITTI<br>depth MSE β |
|
| 107 |
|---|---|---|---|---|---|---|---|---|
|
| 108 |
+
| **Motif Vision Encoder** | 0.5B | **74.0** | **80.5** | **83.5** | 87.2 | <u>52.0</u> | 87.4 | *in progress* |
|
| 109 |
+
| DINOv3 | 1.7B | <u>71.1</u> | <u>79.7</u> | <u>83.3</u> | 88.4 | **55.9** | <u>87.8</u> | **2.3** |
|
| 110 |
| PEcore | 5.4B | 48.2 | 53.1 | 49.8 | **89.3** | 38.9 | **87.9** | 4.1 |
|
| 111 |
+
| SigLIP2 | 10B | 56.1 | 62.3 | 62.9 | <u>89.1</u> | 45.4 | 86.9 | β |
|
| 112 |
| OpenCLIP | 2B | β | β | β | β | β | β | β |
|
| 113 |
+
| V-JEPA 2.1 | 0.022B | 69.0 | β | β | 85.5 | 47.9 | 87.7 | <u>2.5</u> |
|
| 114 |
| VideoMAEv2 | 1.35M | β | β | β | β | β | 90.0<sup>β‘</sup> | β |
|
| 115 |
|
| 116 |
<sub>β‘ VideoMAEv2 ViT-g: K400 top-1 after **full fine-tuning**, not the frozen attentive-probe protocol used for the other rows.</sub>
|