taewhan commited on
Commit
2517d39
Β·
verified Β·
1 Parent(s): a8fa992

docs(README): Training-data column + corrected ImageNet numbers

Browse files

Replace Publisher column with Training data (Motif 0.5B, DINOv3 1.7B, DINOv2 142M, V-JEPA 22M, SigLIP2 10B); correct ImageNet-1K lin. probe (DINOv3 88.4, DINOv2 87.3, SigLIP2 89.1); move ImageNet bold to SigLIP2; update narrative bullets accordingly.

Files changed (1) hide show
  1. README.md +11 -8
README.md CHANGED
@@ -94,19 +94,22 @@ DAVIS S/M/L follow the DINOv3 protocol (J&F-mean at video short side 420/480, 84
94
  1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
95
  single-resolution (S) figure is available.
96
 
97
- | Model | Publisher | DAVIS S<br>J&F ↑ | DAVIS M<br>J&F ↑ | DAVIS L<br>J&F ↑ | ImageNet-1K<br>lin. probe ↑ | ADE20K<br>mIoU ↑ | K400 ↑ | KITTI<br>depth MSE ↓ |
98
  |---|---|---|---|---|---|---|---|---|
99
- | **Motif Vision Encoder** | Motif | **74.0** | **80.5** | **83.5** | 87.2 | 52.0 | *in progress* | *in progress* |
100
- | DINOv3 | Meta | 71.1 | 79.7 | 83.3 | **88.2** | **55.9** | **87.8** | **2.3** |
101
- | DINOv2 | Meta | 63.9 | 73.6 | 76.6 | 86.5 | 49.0 | 84.4 | – |
102
- | V-JEPA 2.1 | Meta | 69.0 | – | – | 85.5 | 47.9 | 87.7 | 3.x |
103
- | SigLIP2 | Google | 56.1 | 62.3 | 62.9 | 84.5 | 45.4 | 86.9 | – |
104
 
105
  - **DAVIS video segmentation (J&F)** β€” leads at every resolution: **74.0 (S) / 80.5 (M) /
106
  83.5 (L)**, surpassing DINOv3 7B (71.1 / 79.7 / 83.3) across the board, reflecting the
107
  encoder's dense, temporally-coherent patch features on video.
108
- - **ImageNet-1K linear probe (87.2)** and **ADE20K semantic segmentation (52.0 mIoU)** are
109
- second only to DINOv3 7B while ahead of DINOv2, V-JEPA 2.1, and SigLIP2.
 
 
 
110
  - **Data efficiency** β€” these results come from **~0.47B training samples** (448.6M images +
111
  18.5M video clips), roughly **3.6Γ— less data than DINOv3**, which is trained on LVD-1689M
112
  (1,689M images). Despite the smaller corpus β€” and with only ~4% of samples being video β€” the
 
94
  1260/1440 px). V-JEPA 2.1 is not part of the DINOv3 Table 5 tracking benchmark, so only its
95
  single-resolution (S) figure is available.
96
 
97
+ | Model | Training<br>data | DAVIS S<br>J&F ↑ | DAVIS M<br>J&F ↑ | DAVIS L<br>J&F ↑ | ImageNet-1K<br>lin. probe ↑ | ADE20K<br>mIoU ↑ | K400 ↑ | KITTI<br>depth MSE ↓ |
98
  |---|---|---|---|---|---|---|---|---|
99
+ | **Motif Vision Encoder** | 0.5B | **74.0** | **80.5** | **83.5** | 87.2 | 52.0 | *in progress* | *in progress* |
100
+ | DINOv3 | 1.7B | 71.1 | 79.7 | 83.3 | 88.4 | **55.9** | **87.8** | **2.3** |
101
+ | DINOv2 | 142M | 63.9 | 73.6 | 76.6 | 87.3 | 49.0 | 84.4 | – |
102
+ | V-JEPA 2.1 | 22M | 69.0 | – | – | 85.5 | 47.9 | 87.7 | 3.x |
103
+ | SigLIP2 | 10B | 56.1 | 62.3 | 62.9 | **89.1** | 45.4 | 86.9 | – |
104
 
105
  - **DAVIS video segmentation (J&F)** β€” leads at every resolution: **74.0 (S) / 80.5 (M) /
106
  83.5 (L)**, surpassing DINOv3 7B (71.1 / 79.7 / 83.3) across the board, reflecting the
107
  encoder's dense, temporally-coherent patch features on video.
108
+ - **ADE20K semantic segmentation (52.0 mIoU)** is second only to DINOv3 7B (55.9), well ahead
109
+ of DINOv2, V-JEPA 2.1, and SigLIP2.
110
+ - **ImageNet-1K linear probe (87.2)** is on par with DINOv2 (87.3) and DINOv3 (88.4), trailing
111
+ the contrastively-trained SigLIP2 (89.1) β€” notable given Motif sees **0.5B** samples versus
112
+ SigLIP2's **10B** image–text pairs and DINOv3's **1.7B** images.
113
  - **Data efficiency** β€” these results come from **~0.47B training samples** (448.6M images +
114
  18.5M video clips), roughly **3.6Γ— less data than DINOv3**, which is trained on LVD-1689M
115
  (1,689M images). Despite the smaller corpus β€” and with only ~4% of samples being video β€” the