# MegaLoc An image retrieval model for any localization task, which achieves SOTA on most VPR datasets, including indoor and outdoor ones. [Gradio Demo](https://gmberton.github.io/demos-url/megaloc/) - [ArXiv](https://arxiv.org/abs/2502.17237) - [Paper on ArXiv](https://arxiv.org/pdf/2502.17237) - [Paper on HF](https://huggingface.co/papers/2502.17237) - [Model on HF](https://huggingface.co/gberton/MegaLoc). ### Demo Try the demo on your own images to see how good MegaLoc is! The demo uses a database of ~5M street-view images from San Francisco, and when you upload one it will find the most similar one from the same place. image ### Using the model You can use the model with torch.hub, as simple as this ``` import torch model = torch.hub.load("gmberton/MegaLoc", "get_trained_model") ``` For more complex uses, like computing results on VPR datasets, visualizing predictions and so on, you can use our [VPR-methods-evaluation](https://github.com/gmberton/VPR-methods-evaluation), which lets you do all this for MegaLoc and multiple other VPR methods on labelled or unlabelled datasets. ### Qualitative examples Here are some examples of top-1 retrieved images from the SF-XL test set, which has 2.8M images as database. ![teaser](https://github.com/user-attachments/assets/a90b8d4c-ab53-4151-aacc-93493d583713) --- ## Available Models Pre-quantized models on HuggingFace: [Realcat/image_retrieval_checkpoints](https://huggingface.co/Realcat/image_retrieval_checkpoints) ```bash # ONNX huggingface-cli download Realcat/image_retrieval_checkpoints --include "megaloc/onnx/*" --local-dir . # CoreML huggingface-cli download Realcat/image_retrieval_checkpoints --include "megaloc/coreml/*" --local-dir . ``` ### Benchmark (MacBook Pro M-series, 518×518, B=1 vs PyTorch FP32) | Model | Size | CosSim | SNR | Latency | FPS | Multi-Batch | |-------|------|--------|-----|---------|-----|-------------| | **PyTorch FP32** | (mem) | 1.000000 | ∞ | 367 ms | 2.7 | — | | ONNX FP32 | 873 MB | 1.000000 | 80.5 dB | 367 ms | 2.7 | **Y** | | ONNX FP16 | 437 MB | 0.999989 | 46.4 dB | 345 ms | 2.9 | **Y** | | ONNX INT8 | 223 MB | 0.996625 | 21.7 dB | 395 ms | 2.5 | **Y** | | CoreML FP32 | 873 MB | 1.000000 | 80.5 dB | **74 ms** | **13.6** | N | | CoreML INT8 | 219 MB | 0.998699 | 25.8 dB | **74 ms** | **13.5** | N | > CoreML FP16 excluded due to poor accuracy (cos=0.944, SNR=9.5dB). CoreML's FP16 path uses half-precision for all intermediate activations, causing error accumulation in this deep ViT. ### ONNX Multi-Batch Throughput | Model | B=1 | B=2 | B=4 | B=8 | |-------|-----|-----|-----|-----| | ONNX FP32 | 2.7 img/s | 2.3 img/s | 2.7 img/s | 2.9 img/s | | ONNX FP16 | 2.9 img/s | 2.6 img/s | 2.7 img/s | 2.9 img/s | | ONNX INT8 | 2.5 img/s | 2.1 img/s | 2.3 img/s | 2.5 img/s | ### Recommendations - **Best accuracy/size**: CoreML INT8 — 219 MB, 74ms, cos=0.9987, 13.5 FPS (macOS only) - **Cross-platform**: ONNX FP16 — 437 MB, 345ms, cos=0.99999, multi-batch - **Maximum compression**: ONNX INT8 — 223 MB, all platforms, multi-batch ### On sub-8-bit quantization 4-bit/1-bit quantization fails catastrophically (cos_sim ≈ 0). DINOv2 ViT has 12 blocks + aggregator (~200 weight tensors). Error from only 16 levels accumulates across 48+ MatMul layers. ViT architectures below 8-bit require **quantization-aware training**. ### Usage ```bash uv sync # ONNX (cross-platform, multi-batch) python3 -c " import onnxruntime as ort, numpy as np from PIL import Image sess = ort.InferenceSession('megaloc_fp16.onnx', providers=['CPUExecutionProvider']) img = Image.open('img.jpg').convert('RGB').resize((518,518), Image.LANCZOS) data = np.expand_dims(np.array(img).astype(np.float32).transpose(2,0,1)/255.0, 0) desc = sess.run(None, {'images': data})[0] # [B, 8448] " # CoreML (macOS, 5x faster) python3 -c " import coremltools as ct, numpy as np from PIL import Image model = ct.models.MLModel('megaloc_coreml_int8.mlpackage') img = Image.open('img.jpg').convert('RGB').resize((518,518), Image.LANCZOS) data = np.expand_dims(np.array(img).astype(np.float32).transpose(2,0,1)/255.0, 0) desc = list(model.predict({'images': data}).values())[0] " ``` ### Reproduce ```bash uv run python export_onnx_batch.py # ONNX multi-batch uv run python quantize_megaloc.py --method 8bit --input megaloc_fp32.onnx # ONNX INT8 uv run python convert_coreml.py # CoreML FP32/FP16/INT8 ``` --- ## Acknowledgements / Cite / BibTex If you use this repository please cite the following ```bibtex @inproceedings{Berton_2025_MegaLoc, author = {Berton, Gabriele and Masone, Carlo}, title = {MegaLoc: One Retrieval to Place Them All}, booktitle = {IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops}, month = {June}, year = {2025}, pages = {2886--2892} } ```