File size: 5,032 Bytes
ac5bcb0
 
a93c459
ac5bcb0
a93c459
ac5bcb0
 
a93c459
ac5bcb0
 
 
 
 
a93c459
ac5bcb0
 
a93c459
 
ac5bcb0
a93c459
ac5bcb0
 
a93c459
ac5bcb0
a93c459
 
 
ac5bcb0
a93c459
ac5bcb0
a93c459
ac5bcb0
a93c459
ac5bcb0
 
 
 
 
 
a93c459
ac5bcb0
a93c459
ac5bcb0
 
 
 
 
 
 
 
a93c459
ac5bcb0
a93c459
ac5bcb0
a93c459
ac5bcb0
 
 
 
 
a93c459
ac5bcb0
 
 
 
 
 
 
 
 
a93c459
ac5bcb0
a93c459
ac5bcb0
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a93c459
ac5bcb0
 
 
a93c459
ac5bcb0
 
 
 
 
 
a93c459
 
ac5bcb0
 
 
a93c459
ac5bcb0
a93c459
 
 
 
ac5bcb0
 
a93c459
ac5bcb0
a93c459
 
ac5bcb0
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
# MegaLoc
An image retrieval model for any localization task, which achieves SOTA on most VPR datasets, including indoor and outdoor ones.

[Gradio Demo](https://gmberton.github.io/demos-url/megaloc/) - [ArXiv](https://arxiv.org/abs/2502.17237) - [Paper on ArXiv](https://arxiv.org/pdf/2502.17237) - [Paper on HF](https://huggingface.co/papers/2502.17237) - [Model on HF](https://huggingface.co/gberton/MegaLoc).

### Demo
Try the demo on your own images to see how good MegaLoc is! The demo uses a database of ~5M street-view images from San Francisco, and when you upload one it will find the most similar one from the same place.

<img width="746" height="576" alt="image" src="https://github.com/user-attachments/assets/4e7a3eec-dfee-4aae-83cc-f5146a1b421d" />


### Using the model
You can use the model with torch.hub, as simple as this
```
import torch
model = torch.hub.load("gmberton/MegaLoc", "get_trained_model")
```

For more complex uses, like computing results on VPR datasets, visualizing predictions and so on, you can use our [VPR-methods-evaluation](https://github.com/gmberton/VPR-methods-evaluation), which lets you do all this for MegaLoc and multiple other VPR methods on labelled or unlabelled datasets.

### Qualitative examples
Here are some examples of top-1 retrieved images from the SF-XL test set, which has 2.8M images as database.

![teaser](https://github.com/user-attachments/assets/a90b8d4c-ab53-4151-aacc-93493d583713)



---

## Available Models

Pre-quantized models on HuggingFace: [Realcat/image_retrieval_checkpoints](https://huggingface.co/Realcat/image_retrieval_checkpoints)

```bash
# ONNX
huggingface-cli download Realcat/image_retrieval_checkpoints --include "megaloc/onnx/*" --local-dir .
# CoreML
huggingface-cli download Realcat/image_retrieval_checkpoints --include "megaloc/coreml/*" --local-dir .
```

### Benchmark (MacBook Pro M-series, 518×518, B=1 vs PyTorch FP32)

| Model | Size | CosSim | SNR | Latency | FPS | Multi-Batch |
|-------|------|--------|-----|---------|-----|-------------|
| **PyTorch FP32** | (mem) | 1.000000 | ∞ | 367 ms | 2.7 | — |
| ONNX FP32 | 873 MB | 1.000000 | 80.5 dB | 367 ms | 2.7 | **Y** |
| ONNX FP16 | 437 MB | 0.999989 | 46.4 dB | 345 ms | 2.9 | **Y** |
| ONNX INT8 | 223 MB | 0.996625 | 21.7 dB | 395 ms | 2.5 | **Y** |
| CoreML FP32 | 873 MB | 1.000000 | 80.5 dB | **74 ms** | **13.6** | N |
| CoreML INT8 | 219 MB | 0.998699 | 25.8 dB | **74 ms** | **13.5** | N |

> CoreML FP16 excluded due to poor accuracy (cos=0.944, SNR=9.5dB). CoreML's FP16 path uses half-precision for all intermediate activations, causing error accumulation in this deep ViT.

### ONNX Multi-Batch Throughput

| Model | B=1 | B=2 | B=4 | B=8 |
|-------|-----|-----|-----|-----|
| ONNX FP32 | 2.7 img/s | 2.3 img/s | 2.7 img/s | 2.9 img/s |
| ONNX FP16 | 2.9 img/s | 2.6 img/s | 2.7 img/s | 2.9 img/s |
| ONNX INT8 | 2.5 img/s | 2.1 img/s | 2.3 img/s | 2.5 img/s |

### Recommendations

- **Best accuracy/size**: CoreML INT8 — 219 MB, 74ms, cos=0.9987, 13.5 FPS (macOS only)
- **Cross-platform**: ONNX FP16 — 437 MB, 345ms, cos=0.99999, multi-batch
- **Maximum compression**: ONNX INT8 — 223 MB, all platforms, multi-batch

### On sub-8-bit quantization

4-bit/1-bit quantization fails catastrophically (cos_sim ≈ 0). DINOv2 ViT has 12 blocks + aggregator (~200 weight tensors). Error from only 16 levels accumulates across 48+ MatMul layers. ViT architectures below 8-bit require **quantization-aware training**.

### Usage

```bash
uv sync

# ONNX (cross-platform, multi-batch)
python3 -c "
import onnxruntime as ort, numpy as np
from PIL import Image
sess = ort.InferenceSession('megaloc_fp16.onnx', providers=['CPUExecutionProvider'])
img = Image.open('img.jpg').convert('RGB').resize((518,518), Image.LANCZOS)
data = np.expand_dims(np.array(img).astype(np.float32).transpose(2,0,1)/255.0, 0)
desc = sess.run(None, {'images': data})[0]  # [B, 8448]
"

# CoreML (macOS, 5x faster)
python3 -c "
import coremltools as ct, numpy as np
from PIL import Image
model = ct.models.MLModel('megaloc_coreml_int8.mlpackage')
img = Image.open('img.jpg').convert('RGB').resize((518,518), Image.LANCZOS)
data = np.expand_dims(np.array(img).astype(np.float32).transpose(2,0,1)/255.0, 0)
desc = list(model.predict({'images': data}).values())[0]
"
```

### Reproduce

```bash
uv run python export_onnx_batch.py                             # ONNX multi-batch
uv run python quantize_megaloc.py --method 8bit --input megaloc_fp32.onnx  # ONNX INT8
uv run python convert_coreml.py                                # CoreML FP32/FP16/INT8
```

---

## Acknowledgements / Cite / BibTex

If you use this repository please cite the following
```bibtex
@inproceedings{Berton_2025_MegaLoc,
    author    = {Berton, Gabriele and Masone, Carlo},
    title     = {MegaLoc: One Retrieval to Place Them All},
    booktitle = {IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops},
    month     = {June},
    year      = {2025},
    pages     = {2886--2892}
}
```