File size: 4,785 Bytes
41db4e7 297878d 21959e5 297878d 41db4e7 21959e5 297878d 21959e5 41db4e7 297878d 41db4e7 21959e5 41db4e7 297878d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 | ---
license: apache-2.0
tags:
- image-classification
- probing
- frozen-features
- imagenet-1k
library_name: pytorch
---
# Efficient Probing heads
**Paper:** [Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency](https://huggingface.co/papers/2506.10178)
(ICLR 2026) · [arXiv:2506.10178](https://arxiv.org/abs/2506.10178) ·
[code & leaderboard](https://github.com/billpsomas/efficient-probing)
Trained EP (efficient probing) heads for all 37 frozen encoders of the standing
ImageNet-1k benchmark at [billpsomas/efficient-probing](https://github.com/billpsomas/efficient-probing).
Each head is a multi-query cross-attention pooling + BatchNorm + linear classifier,
trained for 90 epochs with LARS on frozen features. **No backbone weights are
included** -- each `config.json` records the exact `--model` / loader flags to
rebuild the encoder from its original source.
Why two kinds of head: the training code initially kept only a rolling per-epoch
checkpoint, so a finished run left its **final** epoch behind rather than its best.
13 early-peaking models were later re-run with best-epoch checkpointing, so their
heads are the **peak** and reproduce the leaderboard number exactly. Every file's
metadata records both its own accuracy at the saved epoch and the table's
best-epoch figure, so nothing has to be taken on trust.
## Loading
```python
from huggingface_hub import hf_hub_download
import torch
path = hf_hub_download("billpsomas/efficient-probing-heads",
"dinov3_vit7b/ep_head.pth")
ck = torch.load(path, map_location="cpu", weights_only=False)
head_state, meta = ck["state_dict"], ck["meta"]
# with the benchmark repo on PYTHONPATH:
# model = backbones.build_backbone(args, device) # args from meta
# probe_heads.build_probe_head(model, args)
# model.head.load_state_dict(head_state, strict=True)
```
Or evaluate directly with the benchmark's tool:
```
python tools/eval_reimagenet.py predict <backbone flags from config.json> \
--head_ckpt ep_head.pth --pred_out preds.json
```
## Heads
| encoder | EP variant | top-1 @ saved epoch | epoch | checkpoint is |
|---|---|---:|---:|---|
| DINOv3 ViT-7B/16 | ep_all | 88.36 | 6 | peak |
| MetaCLIP2 ViT-bigG/14-378 | ep | 88.12 | 6 | peak |
| EVA02-CLIP E-14-plus | ep | 87.98 | 6 | peak |
| EVA02-CLIP E-14 | ep | 87.70 | 6 | peak |
| SigLIP2 SO400M/14 | ep | 87.68 | 6 | peak |
| PE-Core L-14/336 | ep | 87.25 | 12 | peak |
| MetaCLIP2 ViT-bigG/14 | ep | 87.11 | 6 | peak |
| SigLIP2 ViT-L/16 | ep | 87.06 | 6 | peak |
| DINOv3 ViT-L/16 | ep_all | 86.73 | 19 | final epoch |
| AIMv2 ViT-L/14 | ep | 85.62 | 19 | final epoch |
| SigLIP ViT-L/16 | ep | 85.93 | 6 | peak |
| DINOv2 ViT-L/14 | ep_all | 85.56 | 15 | peak |
| Franca ViT-L/14 | ep_all | 84.28 | 14 | peak |
| DINOv3 ViT-B/16 | ep_all | 83.77 | 20 | final epoch |
| DINOv2 ViT-B/14 | ep | 83.61 | 25 | final epoch |
| RADIO ViT-L/16 | ep | 83.40 | 89 | final epoch |
| EVA02 ViT-L/14 | ep | 83.22 | 89 | final epoch |
| CLIP ViT-L/14 | ep | 83.22 | 11 | peak |
| CAPI ViT-L/14 | ep | 82.43 | 89 | final epoch |
| BEiTv2 ViT-B/16 | ep | 81.32 | 89 | final epoch |
| RADIO ViT-B/16 | ep | 80.26 | 89 | final epoch |
| iBOT ViT-L/16 | ep_all | 79.43 | 89 | final epoch |
| Hiera ViT-H/16 | ep | 79.82 | 89 | final epoch |
| MAE ViT-L/16 | ep | 79.43 | 89 | final epoch |
| I-JEPA ViT-H/14 | ep | 78.80 | 89 | final epoch |
| iBOT ViT-B/16 | ep_all | 78.62 | 89 | final epoch |
| Hiera ViT-L/16 | ep | 78.51 | 83 | final epoch |
| CLIP ViT-B/16 | ep_all | 77.85 | 11 | peak |
| DINO ViT-B/16 | ep_all | 77.08 | 89 | final epoch |
| MoCov3 ViT-B/16 | ep_all | 76.21 | 89 | final epoch |
| Hiera ViT-B/16 | ep | 75.63 | 88 | final epoch |
| MAE ViT-B/16 | ep | 75.35 | 86 | final epoch |
| MaskFeat ViT-B/16 | ep | 71.68 | 89 | final epoch |
| MaskFeat ViT-L/16 | ep | 69.56 | 89 | final epoch |
| SimMIM ViT-B/16 | ep | 64.81 | 89 | final epoch |
| MAE ViT-S/16 | ep | 64.56 | 89 | final epoch |
| DiT DiT-XL/2 | ep | 56.94 | 86 | final epoch |
Full provenance (training logs, exact commands, the leaderboard itself) lives in the
[GitHub repo](https://github.com/billpsomas/efficient-probing). Heads were trained on
ImageNet-1k; use accordingly.
## Citation
```bibtex
@inproceedings{psomas2026attention,
title = {Attention, Please! Revisiting Attentive Probing Through the Lens of Efficiency},
author = {Bill Psomas and Dionysis Christopoulos and Eirini Baltzi and Ioannis Kakogeorgiou and Tilemachos Aravanis and Nikos Komodakis and Konstantinos Karantzalos and Yannis Avrithis and Giorgos Tolias},
booktitle = {The Fourteenth International Conference on Learning Representations},
year = {2026},
url = {https://openreview.net/forum?id=PXo0gtT7Al}
}
```
|