| # Reproducing the numbers |
|
|
| Everything in `results.json` comes from one pipeline. This file records what to run, what the |
| evaluation harness is, where each input comes from, and what the full run costs. |
|
|
| ## Evaluation harness |
|
|
| `compute_map` from [filipradenovic/revisitop](https://github.com/filipradenovic/revisitop) |
| (`python/evaluate.py`), vendored verbatim. A stage-0 audit scored the vendored copy against the |
| file fetched byte-for-byte from that repository and recomputed mAP for three descriptor sets on |
| both datasets: maximum absolute difference 0.0. The audit also verified that ROxford and RParis |
| queries are absent from their own galleries (70 queries, 4,993 and 6,322 database images), that |
| ground-truth bounding boxes are present, and that queries are cropped to those boxes before |
| embedding while gallery images are not. The crop check is a recompute: the cached query |
| descriptor matches a crop-then-embed at cosine 1.0 and diverges from an uncropped embed at |
| cosine 0.49 to 0.97 depending on how much of the frame the box covers. |
|
|
| Medium uses `ok = easy + hard`, `junk = junk`. Hard uses `ok = hard`, `junk = junk + easy`. |
|
|
| ## Stages |
|
|
| All stages run as Modal functions against the `fusion-data` volume. Scripts live in the |
| `fusion-embeddings` repository under `scripts/`. |
|
|
| | Stage | Script | What it produces | |
| | --- | --- | --- | |
| | 0. Protocol audit | `fp_phase3a_audit.py` | `fp_phase3a/audit_verdict.json` | |
| | 1a. GLDv2 ingest plan and image streaming | `fp_phase3a_data.py --action {plan-full,stream,delta}` | `fp_phase3a/plan/plan_full.json`, `plan/leak_class_ids.json`, packed image shards | |
| | 1b. Feature extraction | `fp_phase2_extract.py` | `fp_phase2/gld_feats/feats_*.pt` (float16 CLS at three scales) | |
| | 2. Head training, both protocols | `fp_phase3a_retrain.py --action train --protocol {standard,decon}` | `fp_phase3a/full_head_{protocol}.pt`, `.json`, eval descriptors, first-stage `nn_*.pkl` | |
| | 2b. Seed spread for stage 2 | `fp_phase3a_seeds.py` (3 seeds per protocol) | `fp_phase3a/seeds/seed_results.json`, `seed_analysis.json`, per-run heads | |
| | 3. AMES reranking, no distractors | `fp_phase3a_ames.py --action ours` | `fp_phase3a/ames_nn_full_{protocol}_dinov2_ames.json` | |
| | 4. Distractor extraction | `fp_phase3a_r1m.py --action pipeline` | `fp_phase3a/r1m/locals_*.hdf5`, `cls_*.pt` | |
| | 5. Distractor finalize | `fp_phase3a_r1m.py --action finalize` | `fp_phase3a/r1m/r1m_order.json`, `r1m_desc_{protocol}.pt` | |
| | 6. +1M evaluation | `fp_phase3a_r1m.py --action eval1m --protocol {standard,decon}` | `fp_phase3a/r1m_eval_{protocol}_dinov2_ames.json` | |
| | 7. Semantic-embedding comparison | `fp_fe2_placerec.py --action {extract,score}` | `fp_fe2_placerec/fe2_{roxford,rparis}.pt`, `fp_fe2_placerec/results.json` | |
|
|
| ### Feature extraction (stages 1 and 4) |
|
|
| `facebook/dinov2-large`, frozen, float16. For each of three scales (1.0, 1.414, 2.0) the short |
| side is set to `round(224 * scale / 14) * 14`, the long side follows the aspect ratio, both are |
| snapped to a multiple of 14, and the long side is clamped to 1022. Bicubic resize, ImageNet mean |
| and standard deviation. The CLS token is L2-normalized per scale, the three are averaged, and |
| the result is re-normalized. `inference.py` in this repository implements exactly this path and |
| reproduces the cached evaluation descriptors to cosine 0.9996 or better. |
|
|
| ### Head training (stage 2) |
|
|
| `Linear(1024, 2048) -> GELU -> Linear(2048, 512, bias=False) -> BatchNorm1d(512)`, ArcFace with |
| margin 0.3 and scale 32, label smoothing 0.1, class-balanced sampling with weight |
| `1/sqrt(class_count)`, AdamW at learning rate 1e-3 and weight decay 5e-4, cosine schedule, batch |
| 4096, 40 epochs, early stop after 8 epochs without improvement. Classes with fewer than three |
| images are dropped. Checkpoint selection is on ROxford Medium per epoch; RParis is not consulted |
| until the final table. The `decon` protocol additionally masks out the 87 classes listed in |
| `fp_phase3a/plan/leak_class_ids.json` before the class filter runs. |
|
|
| Best epoch was 12 for `standard` and 0 for `decon`. |
|
|
| ### Reranking (stages 3 and 6) |
|
|
| AMES is used as published. The image is built from a shallow clone of |
| [pavelsuma/ames](https://github.com/pavelsuma/ames); the local descriptors |
| (`dinov2_gallery_local.hdf5`, `dinov2_query_local.hdf5`) and the `dinov2_ames.pt` checkpoint are |
| downloaded from the authors' host. The only substitution is the first-stage ranking file |
| `nn_*.pkl`, which is generated from our global descriptors. |
|
|
| Before trusting the transplant, `fp_phase3a_ames.py --action repro` runs the authors' own |
| `nn_superglobal.pkl` through the same harness and lands at ROxford 92.70 / 83.75 and RParis |
| 95.26 / 90.66 without distractors, against their published 92.4 +/- 0.9 / 83.1 +/- 1.1 and |
| 95.2 +/- 0.1 / 90.2 +/- 0.4 (supplementary Table 9, three seeds). That is the harness check. |
|
|
| The shortlist is junk-aware, matching the AMES and DELG evaluation code: ground-truth junk ids |
| are moved behind the shortlist before reranking. Every published two-stage row in the comparison |
| does the same. |
|
|
| The reported cell is `k1600_l0.55_t0.3`, the AMES paper default: top-1600 shortlist, |
| `score = 0.55 * global_cosine + 0.45 * sigmoid(0.3 * ames_logit)`. A 15-cell lambda and |
| temperature grid was also measured. Those cells are diagnostics, they are not reported as |
| results, and they are not in `results.json`. For reference, the best grid cell exceeds the |
| paper-default cell by between 0.07 and 1.90 mAP depending on the dataset, difficulty and |
| protocol, which is the size of the effect avoided by fixing the cell in advance. |
|
|
| ## Two-stage reproduction from this repository |
|
|
| `rerank.py` in this repository is the second stage as a standalone script, so the two-stage |
| rows can be reproduced without the Modal pipeline. It builds the junk-aware shortlist from our |
| global descriptors, reranks it with the authors' AMES model over their local descriptors, |
| fuses the two scores and reports Medium and Hard mAP with the same vendored `compute_map`. |
| It is the logic of stage 3 and stage 6 in one file; the Modal scripts remain the record of how |
| the published artifacts were produced. |
|
|
| ### Third-party assets |
|
|
| None of the following is redistributed with this model. |
|
|
| | Asset | Where it comes from | Terms | |
| | --- | --- | --- | |
| | AMES code | `github.com/pavelsuma/ames`, fetched by `torch.hub` on first use, or `--ames-repo <clone>` | Apache-2.0 | |
| | `dinov2_ames.pt` | downloaded by the authors' model class from `ptak.felk.cvut.cz/personal/sumapave/public/ames/networks/`, cached under `TORCH_HOME` | authors' host, no separate license stated | |
| | `dinov2_query_local.hdf5`, `dinov2_gallery_local.hdf5`, per dataset | `ptak.felk.cvut.cz/personal/sumapave/public/ames/data/<dataset>/`, downloaded by `--fetch` | same | |
| | `gnd_<dataset>.pkl` | revisited Oxford / Paris release, mirrored on the same host | `github.com/filipradenovic/revisitop` | |
|
|
| Cite AMES (Suma et al., ECCV 2024, arXiv 2408.03282) for any use of the two-stage numbers. |
| The local descriptors are DINOv2-B features, which is why the pretraining caveat in |
| `README_hf.md` applies to the two-stage rows as well as the global-only ones. |
|
|
| ### Without distractors |
|
|
| One-time asset download. The gallery local descriptors are 5.44 GB for ROxford and 6.89 GB for |
| RParis, plus 0.08 GB of query descriptors each: |
|
|
| ```bash |
| python rerank.py --dataset roxford5k --ames-dir ./ames_assets --fetch-only |
| python rerank.py --dataset rparis6k --ames-dir ./ames_assets --fetch-only |
| ``` |
|
|
| `--fetch-only` writes to a `.tmp` name and renames, so an interrupted download is never |
| mistaken for a complete one, but it does not resume. For a resumable download use `wget -c` |
| against the same host, which is also what the AMES README suggests. |
|
|
| Then, per dataset, computing our descriptors from the benchmark images with `inference.py`: |
|
|
| ```bash |
| python rerank.py --dataset roxford5k --ames-dir ./ames_assets \ |
| --images-root /data/roxford5k/jpg --head standard \ |
| --save-descriptors roxford5k_standard_desc.pt --out roxford5k_standard.json |
| |
| python rerank.py --dataset rparis6k --ames-dir ./ames_assets \ |
| --images-root /data/rparis6k/jpg --head standard \ |
| --save-descriptors rparis6k_standard_desc.pt --out rparis6k_standard.json |
| ``` |
|
|
| `--images-root` is the flat `jpg` directory of the benchmark; file names come from the |
| ground-truth pickle. Queries are cropped to the ground-truth box, gallery images are not. |
| `--descriptors <file>.pt` skips the embedding pass on a rerun, and `--head decon` selects the |
| decontaminated head. Defaults are `--topk 1600 --lambdas 0.55 --temps 0.3`, the reported cell; |
| the arguments take comma-separated lists if you want the diagnostic grid. |
|
|
| Each dataset is one pass of 70 x 1600 query-candidate pairs per difficulty setting, and Medium |
| and Hard are separate passes because their junk sets differ. Measured on an A10G: 12.4 minutes |
| per pass, so about 25 minutes per dataset and just under an hour for both, a little over a |
| dollar, plus roughly 10 minutes per dataset if the descriptors are being computed from images |
| rather than supplied. |
|
|
| ### What the verification run produced |
|
|
| `rerank.py` was run against both benchmarks with the standard head and the default cell. Two |
| descriptor sources were tried, because they answer different questions. |
|
|
| Feeding in the descriptor set the reported cell was computed from, which is the shipped head |
| weights applied to the cached multi-scale CLS features in float32: |
|
|
| | Source | ROxf M | ROxf H | RPar M | RPar H | |
| | --- | --- | --- | --- | --- | |
| | `results.json` | 91.00 | 80.49 | 95.53 | 91.43 | |
| | `rerank.py`, supplied descriptors | 91.00 | 80.49 | 95.53 | 91.43 | |
|
|
| Every cell matches. The global-only rows match too, at 76.51 / 58.09 and 92.85 / 84.88. The |
| stored artifact holds 0.80488 for ROxford Hard and 0.91429 for RParis Hard where `rerank.py` |
| reports 0.8049 and 0.9143, because the two paths round at different points: the Modal stage |
| used the AMES metric wrapper, which rounds to three decimals of a percent, while `rerank.py` |
| rounds the fraction to four places. Nothing else differs. |
|
|
| Recomputing the descriptors from the JPEGs with `--images-root`, which is the path a user |
| without our cached features takes, on ROxford: |
|
|
| | Source | ROxf M global | ROxf H global | ROxf M two stage | ROxf H two stage | |
| | --- | --- | --- | --- | --- | |
| | `results.json` | 76.51 | 58.09 | 91.00 | 80.49 | |
| | `rerank.py`, `--images-root` | 76.64 | 58.26 | 91.01 | 80.50 | |
|
|
| The first stage moves by 0.13 and 0.17 mAP and the two-stage cells by 0.01. This is the |
| descriptor-level difference already recorded above: `inference.py` reproduces the cached |
| evaluation descriptors at cosine 0.9996 or better, with the residual concentrated in the |
| bounding-box-cropped queries, and 0.9996 on 70 queries is worth about a tenth of a mAP point. |
| The reranker absorbs almost all of it, because a shortlist of 1600 out of 4993 is insensitive |
| to a reordering that small. Expect a tenth of a point either way from this path; the exact |
| figures come from supplying the descriptors. This second run covers ROxford only, since it is |
| the tighter of the two cells and one dataset is enough to size the effect. |
|
|
| ### With the +1M distractors |
|
|
| `--distractor-locals` and `--distractor-desc` extend the gallery. Both are required together: |
|
|
| ```bash |
| python rerank.py --dataset roxford5k --ames-dir ./ames_assets \ |
| --descriptors roxford5k_standard_desc.pt --head standard \ |
| --distractor-locals /data/r1m --distractor-desc /data/r1m/r1m_desc_standard.pt \ |
| --out roxford5k_standard_1m.json |
| ``` |
|
|
| `--distractor-locals` is a directory holding `r1m_order.json` and the `locals_XXXX.hdf5` |
| shards it names; `--distractor-desc` is `{"desc": [1001001, 512]}` in exactly that order. The |
| gallery becomes `[benchmark database ; distractor shards]` and the shortlist is built over all |
| of it. Both files come from the pipeline above: |
|
|
| ``` |
| fp_phase3a_r1m.py --action pipeline # -> r1m/locals_XXXX.hdf5 and cls_XXXX.pt |
| fp_phase3a_r1m.py --action finalize # -> r1m/r1m_order.json and r1m_desc_{protocol}.pt |
| ``` |
|
|
| The extraction, not the reranking, is what makes this expensive: about 1.08 TB of float16 |
| local descriptors, roughly $100 of GPU time, and the whole set has to be resident because the |
| shards are opened lazily across the full gallery. The cost table below is the measured |
| projection. Nothing about the +1M setting changes the reranker or the reported cell; it only |
| changes what is in the gallery. |
|
|
| What was verified through `rerank.py` is the no-distractor setting, above. The +1M cells in |
| `results.json` stand on the `fp_phase3a_r1m.py --action eval1m` run recorded in |
| `r1m_eval_{protocol}_dinov2_ames.json`; `rerank.py` is the same shortlist, model, fusion and |
| scorer with a longer gallery, and it has not been re-run at +1M, because that run is the |
| expensive one. A user who has the shards can run it. The gallery alignment is checked before |
| the model loads, so a mismatch between the number of distractor descriptors and the rows in |
| the shards fails immediately rather than producing a wrong number. |
|
|
| ### Distractor set (stages 4 to 6) |
|
|
| `revisitop1m` from `http://ptak.felk.cvut.cz/revisitop/revisitop1m`, 100 gzipped archives. |
| Each archive is downloaded to container-local scratch, unpacked, run through (a) the AMES |
| authors' `extract_descriptors.py` verbatim to produce 700 local descriptors per image, stored as |
| float16 HDF5, and (b) our DINOv2-L three-scale CLS extraction. The raw JPEGs are deleted |
| immediately after. The step is resumable, skips archives whose two outputs both exist, and is |
| capped at 10 concurrent containers. |
|
|
| `revisitop1m` contains a small number of truncated JPEGs. `ImageFile.LOAD_TRUNCATED_IMAGES` is |
| enabled so those decode with padding instead of raising, which keeps our CLS rows aligned with |
| the AMES local-descriptor rows. |
|
|
| `finalize` concatenates the per-archive CLS shards in canonical archive order, writes |
| `r1m_order.json` (100 shards, 1,001,001 images total), and applies each head to produce |
| `r1m_desc_{protocol}.pt`. `eval1m` builds the gallery as `[benchmark database ; R1M shards]` in |
| that order, scores global-only mAP, then reranks the top-1600. |
|
|
| Before the full run, `fp_phase3a_r1m.py --action valcheck` reproduces the authors' local |
| extraction on the first 50 ROxford gallery images and cosine-compares against their published |
| HDF5, and `--action measure` times one archive and writes a cost projection to |
| `fp_phase3a/r1m_projection.json`. The projection had a hard pause threshold of $200. |
|
|
| ### Semantic-embedding comparison (stage 7) |
|
|
| `EximiusLabs/fusion-embedding-2-2b-preview` at revision |
| `1720d8b16af578d794d7b21ee7b829281941899d`, bfloat16 on an A10G. Every vector comes from that |
| model's own released image path: the base chat template with the system instruction, the |
| `<|vision_start|><|image_pad|><|vision_end|>` user turn, last-token pooling, Matryoshka prefix |
| truncation and L2 normalization. The extractor is checked against the released `embed_image` |
| API on sampled images from each dataset before extraction starts and aborts if the maximum |
| absolute difference exceeds 1e-6; measured 1.5e-8. Image geometry is whatever the base |
| processor does natively, aspect ratio preserved, no fixed square crop and no multi-scale |
| averaging, which is the documented path for that model. Gallery images carry the document |
| instruction and are never cropped; queries are cropped to the ground-truth box, exactly as in |
| every other row of the table. |
|
|
| Twelve read-outs are scored per dataset from the stored 2048-d pooled vectors, at no additional |
| GPU cost: document instruction on both sides against query instruction on the query side, |
| 1024-d against 2048-d, mean-centering on against off. The reported row is the strongest |
| configuration with protocol-cropped queries, selected on the sum of Medium mAP over both |
| datasets, and the full grid stays in the artifact. Extraction is 5,206 and 6,535 forward passes |
| and took 37 and 31 minutes on one A10G, under $2. |
|
|
| ## Cost and storage |
|
|
| From `fp_phase3a/r1m_projection.json`, measured on one archive of 510 images and extrapolated: |
|
|
| | Item | Measured rate | Projected for 1M | |
| | --- | --- | --- | |
| | Download and unpack | 11.7 MB/s | 11.5 h serial, ~$12 CPU across parallel containers | |
| | AMES local extraction | 4.4 img/s on A10G | 63.1 A10G-hours, ~$69 | |
| | Our global extraction | 18.0 img/s on A10G | 15.4 A10G-hours, ~$17 | |
| | Reranking | | ~$2 | |
| | **Total** | | **~$100** | |
|
|
| Storage is the part that surprises people: the float16 local descriptors for 1M images occupy |
| **1.083 TB** across the 100 shards, measured on the volume. All of it has to be resident at |
| once, because the evaluation opens shards lazily across the whole gallery. |
|
|
| Everything before the distractor run (audit, ingest, feature extraction, head training, |
| no-distractor reranking, preflight) is estimated at $80 for this round and $110 cumulative |
| including earlier phases, in `fp_phase3a/final_report.json`. Adding the +1M run puts the project |
| near $200 all in. Treat all of these as estimates; the authoritative number is the Modal |
| dashboard. |
|
|
| ## Artifact index |
|
|
| Every value in `results.json` traces to one of these. |
|
|
| | Artifact | Contains | |
| | --- | --- | |
| | `fusion-data:/fp_phase3a/r1m_eval_standard_dinov2_ames.json` | +1M global-only and reranked, standard head | |
| | `fusion-data:/fp_phase3a/r1m_eval_decon_dinov2_ames.json` | +1M global-only and reranked, decon head | |
| | `fusion-data:/fp_phase3a/ames_nn_full_standard_dinov2_ames.json` | no-distractor reranked grid, standard head | |
| | `fusion-data:/fp_phase3a/ames_nn_full_decon_dinov2_ames.json` | no-distractor reranked grid, decon head | |
| | `fusion-data:/fp_phase3a/full_head_standard.json` | no-distractor global-only, training rows and classes | |
| | `fusion-data:/fp_phase3a/full_head_decon.json` | same, decon head | |
| | `fusion-data:/fp_phase3a/full_head_{protocol}.pt` | the shipped head weights | |
| | `fusion-data:/fp_phase3a/r1m/r1m_order.json` | distractor gallery composition, 1,001,001 images | |
| | `fusion-data:/fp_phase3a/plan/plan_full.json` | GLDv2 audit counts: 87 leak classes, 2,529 leak images | |
| | `fusion-data:/fp_phase3a/plan/leak_class_ids.json` | the 87 class ids | |
| | `fusion-data:/fp_phase2/plan/leakage_report.json` | per-class match records: landmark id, Wikimedia category, matching rule | |
| | `leak_class_ids.json` (in this package) | the 87 classes as published, merged from the three artifacts above | |
| | `fusion-data:/fp_phase3a/audit_verdict.json` | stage-0 protocol audit | |
| | `fusion-data:/fp_phase3a/r1m_projection.json` | measured rates and cost projection | |
| | `fusion-data:/fp_fe2_placerec/results.json` | the fusion-embedding-2 image-tower comparison, every read-out variant measured | |
| | `fusion-data:/fp_fe2_placerec/fe2_{roxford,rparis}.pt` | the raw fusion-embedding-2 gallery and query vectors those numbers are scored from | |
|
|