# Reproducing the numbers Everything in `results.json` comes from one pipeline. This file records what to run, what the evaluation harness is, where each input comes from, and what the full run costs. ## Evaluation harness `compute_map` from [filipradenovic/revisitop](https://github.com/filipradenovic/revisitop) (`python/evaluate.py`), vendored verbatim. A stage-0 audit scored the vendored copy against the file fetched byte-for-byte from that repository and recomputed mAP for three descriptor sets on both datasets: maximum absolute difference 0.0. The audit also verified that ROxford and RParis queries are absent from their own galleries (70 queries, 4,993 and 6,322 database images), that ground-truth bounding boxes are present, and that queries are cropped to those boxes before embedding while gallery images are not. The crop check is a recompute: the cached query descriptor matches a crop-then-embed at cosine 1.0 and diverges from an uncropped embed at cosine 0.49 to 0.97 depending on how much of the frame the box covers. Medium uses `ok = easy + hard`, `junk = junk`. Hard uses `ok = hard`, `junk = junk + easy`. ## Stages All stages run as Modal functions against the `fusion-data` volume. Scripts live in the `fusion-embeddings` repository under `scripts/`. | Stage | Script | What it produces | | --- | --- | --- | | 0. Protocol audit | `fp_phase3a_audit.py` | `fp_phase3a/audit_verdict.json` | | 1a. GLDv2 ingest plan and image streaming | `fp_phase3a_data.py --action {plan-full,stream,delta}` | `fp_phase3a/plan/plan_full.json`, `plan/leak_class_ids.json`, packed image shards | | 1b. Feature extraction | `fp_phase2_extract.py` | `fp_phase2/gld_feats/feats_*.pt` (float16 CLS at three scales) | | 2. Head training, both protocols | `fp_phase3a_retrain.py --action train --protocol {standard,decon}` | `fp_phase3a/full_head_{protocol}.pt`, `.json`, eval descriptors, first-stage `nn_*.pkl` | | 2b. Seed spread for stage 2 | `fp_phase3a_seeds.py` (3 seeds per protocol) | `fp_phase3a/seeds/seed_results.json`, `seed_analysis.json`, per-run heads | | 3. AMES reranking, no distractors | `fp_phase3a_ames.py --action ours` | `fp_phase3a/ames_nn_full_{protocol}_dinov2_ames.json` | | 4. Distractor extraction | `fp_phase3a_r1m.py --action pipeline` | `fp_phase3a/r1m/locals_*.hdf5`, `cls_*.pt` | | 5. Distractor finalize | `fp_phase3a_r1m.py --action finalize` | `fp_phase3a/r1m/r1m_order.json`, `r1m_desc_{protocol}.pt` | | 6. +1M evaluation | `fp_phase3a_r1m.py --action eval1m --protocol {standard,decon}` | `fp_phase3a/r1m_eval_{protocol}_dinov2_ames.json` | | 7. Semantic-embedding comparison | `fp_fe2_placerec.py --action {extract,score}` | `fp_fe2_placerec/fe2_{roxford,rparis}.pt`, `fp_fe2_placerec/results.json` | ### Feature extraction (stages 1 and 4) `facebook/dinov2-large`, frozen, float16. For each of three scales (1.0, 1.414, 2.0) the short side is set to `round(224 * scale / 14) * 14`, the long side follows the aspect ratio, both are snapped to a multiple of 14, and the long side is clamped to 1022. Bicubic resize, ImageNet mean and standard deviation. The CLS token is L2-normalized per scale, the three are averaged, and the result is re-normalized. `inference.py` in this repository implements exactly this path and reproduces the cached evaluation descriptors to cosine 0.9996 or better. ### Head training (stage 2) `Linear(1024, 2048) -> GELU -> Linear(2048, 512, bias=False) -> BatchNorm1d(512)`, ArcFace with margin 0.3 and scale 32, label smoothing 0.1, class-balanced sampling with weight `1/sqrt(class_count)`, AdamW at learning rate 1e-3 and weight decay 5e-4, cosine schedule, batch 4096, 40 epochs, early stop after 8 epochs without improvement. Classes with fewer than three images are dropped. Checkpoint selection is on ROxford Medium per epoch; RParis is not consulted until the final table. The `decon` protocol additionally masks out the 87 classes listed in `fp_phase3a/plan/leak_class_ids.json` before the class filter runs. Best epoch was 12 for `standard` and 0 for `decon`. ### Reranking (stages 3 and 6) AMES is used as published. The image is built from a shallow clone of [pavelsuma/ames](https://github.com/pavelsuma/ames); the local descriptors (`dinov2_gallery_local.hdf5`, `dinov2_query_local.hdf5`) and the `dinov2_ames.pt` checkpoint are downloaded from the authors' host. The only substitution is the first-stage ranking file `nn_*.pkl`, which is generated from our global descriptors. Before trusting the transplant, `fp_phase3a_ames.py --action repro` runs the authors' own `nn_superglobal.pkl` through the same harness and lands at ROxford 92.70 / 83.75 and RParis 95.26 / 90.66 without distractors, against their published 92.4 +/- 0.9 / 83.1 +/- 1.1 and 95.2 +/- 0.1 / 90.2 +/- 0.4 (supplementary Table 9, three seeds). That is the harness check. The shortlist is junk-aware, matching the AMES and DELG evaluation code: ground-truth junk ids are moved behind the shortlist before reranking. Every published two-stage row in the comparison does the same. The reported cell is `k1600_l0.55_t0.3`, the AMES paper default: top-1600 shortlist, `score = 0.55 * global_cosine + 0.45 * sigmoid(0.3 * ames_logit)`. A 15-cell lambda and temperature grid was also measured. Those cells are diagnostics, they are not reported as results, and they are not in `results.json`. For reference, the best grid cell exceeds the paper-default cell by between 0.07 and 1.90 mAP depending on the dataset, difficulty and protocol, which is the size of the effect avoided by fixing the cell in advance. ## Two-stage reproduction from this repository `rerank.py` in this repository is the second stage as a standalone script, so the two-stage rows can be reproduced without the Modal pipeline. It builds the junk-aware shortlist from our global descriptors, reranks it with the authors' AMES model over their local descriptors, fuses the two scores and reports Medium and Hard mAP with the same vendored `compute_map`. It is the logic of stage 3 and stage 6 in one file; the Modal scripts remain the record of how the published artifacts were produced. ### Third-party assets None of the following is redistributed with this model. | Asset | Where it comes from | Terms | | --- | --- | --- | | AMES code | `github.com/pavelsuma/ames`, fetched by `torch.hub` on first use, or `--ames-repo ` | Apache-2.0 | | `dinov2_ames.pt` | downloaded by the authors' model class from `ptak.felk.cvut.cz/personal/sumapave/public/ames/networks/`, cached under `TORCH_HOME` | authors' host, no separate license stated | | `dinov2_query_local.hdf5`, `dinov2_gallery_local.hdf5`, per dataset | `ptak.felk.cvut.cz/personal/sumapave/public/ames/data//`, downloaded by `--fetch` | same | | `gnd_.pkl` | revisited Oxford / Paris release, mirrored on the same host | `github.com/filipradenovic/revisitop` | Cite AMES (Suma et al., ECCV 2024, arXiv 2408.03282) for any use of the two-stage numbers. The local descriptors are DINOv2-B features, which is why the pretraining caveat in `README_hf.md` applies to the two-stage rows as well as the global-only ones. ### Without distractors One-time asset download. The gallery local descriptors are 5.44 GB for ROxford and 6.89 GB for RParis, plus 0.08 GB of query descriptors each: ```bash python rerank.py --dataset roxford5k --ames-dir ./ames_assets --fetch-only python rerank.py --dataset rparis6k --ames-dir ./ames_assets --fetch-only ``` `--fetch-only` writes to a `.tmp` name and renames, so an interrupted download is never mistaken for a complete one, but it does not resume. For a resumable download use `wget -c` against the same host, which is also what the AMES README suggests. Then, per dataset, computing our descriptors from the benchmark images with `inference.py`: ```bash python rerank.py --dataset roxford5k --ames-dir ./ames_assets \ --images-root /data/roxford5k/jpg --head standard \ --save-descriptors roxford5k_standard_desc.pt --out roxford5k_standard.json python rerank.py --dataset rparis6k --ames-dir ./ames_assets \ --images-root /data/rparis6k/jpg --head standard \ --save-descriptors rparis6k_standard_desc.pt --out rparis6k_standard.json ``` `--images-root` is the flat `jpg` directory of the benchmark; file names come from the ground-truth pickle. Queries are cropped to the ground-truth box, gallery images are not. `--descriptors .pt` skips the embedding pass on a rerun, and `--head decon` selects the decontaminated head. Defaults are `--topk 1600 --lambdas 0.55 --temps 0.3`, the reported cell; the arguments take comma-separated lists if you want the diagnostic grid. Each dataset is one pass of 70 x 1600 query-candidate pairs per difficulty setting, and Medium and Hard are separate passes because their junk sets differ. Measured on an A10G: 12.4 minutes per pass, so about 25 minutes per dataset and just under an hour for both, a little over a dollar, plus roughly 10 minutes per dataset if the descriptors are being computed from images rather than supplied. ### What the verification run produced `rerank.py` was run against both benchmarks with the standard head and the default cell. Two descriptor sources were tried, because they answer different questions. Feeding in the descriptor set the reported cell was computed from, which is the shipped head weights applied to the cached multi-scale CLS features in float32: | Source | ROxf M | ROxf H | RPar M | RPar H | | --- | --- | --- | --- | --- | | `results.json` | 91.00 | 80.49 | 95.53 | 91.43 | | `rerank.py`, supplied descriptors | 91.00 | 80.49 | 95.53 | 91.43 | Every cell matches. The global-only rows match too, at 76.51 / 58.09 and 92.85 / 84.88. The stored artifact holds 0.80488 for ROxford Hard and 0.91429 for RParis Hard where `rerank.py` reports 0.8049 and 0.9143, because the two paths round at different points: the Modal stage used the AMES metric wrapper, which rounds to three decimals of a percent, while `rerank.py` rounds the fraction to four places. Nothing else differs. Recomputing the descriptors from the JPEGs with `--images-root`, which is the path a user without our cached features takes, on ROxford: | Source | ROxf M global | ROxf H global | ROxf M two stage | ROxf H two stage | | --- | --- | --- | --- | --- | | `results.json` | 76.51 | 58.09 | 91.00 | 80.49 | | `rerank.py`, `--images-root` | 76.64 | 58.26 | 91.01 | 80.50 | The first stage moves by 0.13 and 0.17 mAP and the two-stage cells by 0.01. This is the descriptor-level difference already recorded above: `inference.py` reproduces the cached evaluation descriptors at cosine 0.9996 or better, with the residual concentrated in the bounding-box-cropped queries, and 0.9996 on 70 queries is worth about a tenth of a mAP point. The reranker absorbs almost all of it, because a shortlist of 1600 out of 4993 is insensitive to a reordering that small. Expect a tenth of a point either way from this path; the exact figures come from supplying the descriptors. This second run covers ROxford only, since it is the tighter of the two cells and one dataset is enough to size the effect. ### With the +1M distractors `--distractor-locals` and `--distractor-desc` extend the gallery. Both are required together: ```bash python rerank.py --dataset roxford5k --ames-dir ./ames_assets \ --descriptors roxford5k_standard_desc.pt --head standard \ --distractor-locals /data/r1m --distractor-desc /data/r1m/r1m_desc_standard.pt \ --out roxford5k_standard_1m.json ``` `--distractor-locals` is a directory holding `r1m_order.json` and the `locals_XXXX.hdf5` shards it names; `--distractor-desc` is `{"desc": [1001001, 512]}` in exactly that order. The gallery becomes `[benchmark database ; distractor shards]` and the shortlist is built over all of it. Both files come from the pipeline above: ``` fp_phase3a_r1m.py --action pipeline # -> r1m/locals_XXXX.hdf5 and cls_XXXX.pt fp_phase3a_r1m.py --action finalize # -> r1m/r1m_order.json and r1m_desc_{protocol}.pt ``` The extraction, not the reranking, is what makes this expensive: about 1.08 TB of float16 local descriptors, roughly $100 of GPU time, and the whole set has to be resident because the shards are opened lazily across the full gallery. The cost table below is the measured projection. Nothing about the +1M setting changes the reranker or the reported cell; it only changes what is in the gallery. What was verified through `rerank.py` is the no-distractor setting, above. The +1M cells in `results.json` stand on the `fp_phase3a_r1m.py --action eval1m` run recorded in `r1m_eval_{protocol}_dinov2_ames.json`; `rerank.py` is the same shortlist, model, fusion and scorer with a longer gallery, and it has not been re-run at +1M, because that run is the expensive one. A user who has the shards can run it. The gallery alignment is checked before the model loads, so a mismatch between the number of distractor descriptors and the rows in the shards fails immediately rather than producing a wrong number. ### Distractor set (stages 4 to 6) `revisitop1m` from `http://ptak.felk.cvut.cz/revisitop/revisitop1m`, 100 gzipped archives. Each archive is downloaded to container-local scratch, unpacked, run through (a) the AMES authors' `extract_descriptors.py` verbatim to produce 700 local descriptors per image, stored as float16 HDF5, and (b) our DINOv2-L three-scale CLS extraction. The raw JPEGs are deleted immediately after. The step is resumable, skips archives whose two outputs both exist, and is capped at 10 concurrent containers. `revisitop1m` contains a small number of truncated JPEGs. `ImageFile.LOAD_TRUNCATED_IMAGES` is enabled so those decode with padding instead of raising, which keeps our CLS rows aligned with the AMES local-descriptor rows. `finalize` concatenates the per-archive CLS shards in canonical archive order, writes `r1m_order.json` (100 shards, 1,001,001 images total), and applies each head to produce `r1m_desc_{protocol}.pt`. `eval1m` builds the gallery as `[benchmark database ; R1M shards]` in that order, scores global-only mAP, then reranks the top-1600. Before the full run, `fp_phase3a_r1m.py --action valcheck` reproduces the authors' local extraction on the first 50 ROxford gallery images and cosine-compares against their published HDF5, and `--action measure` times one archive and writes a cost projection to `fp_phase3a/r1m_projection.json`. The projection had a hard pause threshold of $200. ### Semantic-embedding comparison (stage 7) `EximiusLabs/fusion-embedding-2-2b-preview` at revision `1720d8b16af578d794d7b21ee7b829281941899d`, bfloat16 on an A10G. Every vector comes from that model's own released image path: the base chat template with the system instruction, the `<|vision_start|><|image_pad|><|vision_end|>` user turn, last-token pooling, Matryoshka prefix truncation and L2 normalization. The extractor is checked against the released `embed_image` API on sampled images from each dataset before extraction starts and aborts if the maximum absolute difference exceeds 1e-6; measured 1.5e-8. Image geometry is whatever the base processor does natively, aspect ratio preserved, no fixed square crop and no multi-scale averaging, which is the documented path for that model. Gallery images carry the document instruction and are never cropped; queries are cropped to the ground-truth box, exactly as in every other row of the table. Twelve read-outs are scored per dataset from the stored 2048-d pooled vectors, at no additional GPU cost: document instruction on both sides against query instruction on the query side, 1024-d against 2048-d, mean-centering on against off. The reported row is the strongest configuration with protocol-cropped queries, selected on the sum of Medium mAP over both datasets, and the full grid stays in the artifact. Extraction is 5,206 and 6,535 forward passes and took 37 and 31 minutes on one A10G, under $2. ## Cost and storage From `fp_phase3a/r1m_projection.json`, measured on one archive of 510 images and extrapolated: | Item | Measured rate | Projected for 1M | | --- | --- | --- | | Download and unpack | 11.7 MB/s | 11.5 h serial, ~$12 CPU across parallel containers | | AMES local extraction | 4.4 img/s on A10G | 63.1 A10G-hours, ~$69 | | Our global extraction | 18.0 img/s on A10G | 15.4 A10G-hours, ~$17 | | Reranking | | ~$2 | | **Total** | | **~$100** | Storage is the part that surprises people: the float16 local descriptors for 1M images occupy **1.083 TB** across the 100 shards, measured on the volume. All of it has to be resident at once, because the evaluation opens shards lazily across the whole gallery. Everything before the distractor run (audit, ingest, feature extraction, head training, no-distractor reranking, preflight) is estimated at $80 for this round and $110 cumulative including earlier phases, in `fp_phase3a/final_report.json`. Adding the +1M run puts the project near $200 all in. Treat all of these as estimates; the authoritative number is the Modal dashboard. ## Artifact index Every value in `results.json` traces to one of these. | Artifact | Contains | | --- | --- | | `fusion-data:/fp_phase3a/r1m_eval_standard_dinov2_ames.json` | +1M global-only and reranked, standard head | | `fusion-data:/fp_phase3a/r1m_eval_decon_dinov2_ames.json` | +1M global-only and reranked, decon head | | `fusion-data:/fp_phase3a/ames_nn_full_standard_dinov2_ames.json` | no-distractor reranked grid, standard head | | `fusion-data:/fp_phase3a/ames_nn_full_decon_dinov2_ames.json` | no-distractor reranked grid, decon head | | `fusion-data:/fp_phase3a/full_head_standard.json` | no-distractor global-only, training rows and classes | | `fusion-data:/fp_phase3a/full_head_decon.json` | same, decon head | | `fusion-data:/fp_phase3a/full_head_{protocol}.pt` | the shipped head weights | | `fusion-data:/fp_phase3a/r1m/r1m_order.json` | distractor gallery composition, 1,001,001 images | | `fusion-data:/fp_phase3a/plan/plan_full.json` | GLDv2 audit counts: 87 leak classes, 2,529 leak images | | `fusion-data:/fp_phase3a/plan/leak_class_ids.json` | the 87 class ids | | `fusion-data:/fp_phase2/plan/leakage_report.json` | per-class match records: landmark id, Wikimedia category, matching rule | | `leak_class_ids.json` (in this package) | the 87 classes as published, merged from the three artifacts above | | `fusion-data:/fp_phase3a/audit_verdict.json` | stage-0 protocol audit | | `fusion-data:/fp_phase3a/r1m_projection.json` | measured rates and cost projection | | `fusion-data:/fp_fe2_placerec/results.json` | the fusion-embedding-2 image-tower comparison, every read-out variant measured | | `fusion-data:/fp_fe2_placerec/fe2_{roxford,rparis}.pt` | the raw fusion-embedding-2 gallery and query vectors those numbers are scored from |