File size: 5,889 Bytes
3a87185 1bb5a3e 3a87185 1bb5a3e | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 | ---
license: apache-2.0
tags:
- geolocation
- image-regression
- dinov3
- street-view
- jakarta
- computer-vision
library_name: pytorch
language:
- id
datasets:
- nadh0708/JKTSV-primary
---
# Img2LocJakarta — DINOv3 street-view geolocation for Jakarta
Predict **(latitude, longitude)** of a Google Street View perspective crop taken
on a Jakarta road. A frozen **DINOv3 ViT-L/16** backbone extracts features; a
U-shaped MLP head regresses a local flat-earth (x, y) offset which is converted
to WGS-84 degrees.
| | |
|---|---|
| **Input** | RGB street-level image (any size, resized to 224×224) |
| **Output** | `{"lat": float, "lon": float}` — WGS-84 degrees |
| **Scope** | Trained on Jakarta (5 administrative cities, motorway + primary roads) |
| **Backbone** | DINOv3 ViT-L/16 — frozen; only the regression head was trained |
## Quick Start
```bash
pip install torch torchvision pillow huggingface_hub safetensors
```
```python
from huggingface_hub import hf_hub_download
import sys, os
# Download the inference code
for fname in ["inference.py", "modeling_geotag.py"]:
hf_hub_download("nadh0708/Img2LocJakarta", fname, local_dir=".")
from inference import GeoTagPredictor
predictor = GeoTagPredictor("nadh0708/Img2LocJakarta")
print(predictor.predict("street.jpg"))
# {'lat': -6.2261, 'lon': 106.8123}
# Batched
print(predictor.predict(["a.jpg", "b.jpg"]))
```
### From a local checkpoint
```python
predictor = GeoTagPredictor("modelD_40e_2.pth")
```
### CLI
```bash
python inference.py nadh0708/Img2LocJakarta street.jpg
```
### Low-level API
```python
import torch
from torchvision import transforms
from PIL import Image
from modeling_geotag import DinoGeoRegressor
model = DinoGeoRegressor.from_pretrained("nadh0708/Img2LocJakarta")
model.eval()
tf = transforms.Compose([
transforms.Resize((224, 224)),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
x = tf(Image.open("street.jpg").convert("RGB")).unsqueeze(0)
lonlat = model.predict_lonlat(x) # tensor([[lon, lat]])
```
## Architecture
```
DinoGeoRegressor
├─ encoder : DINOv3 ViT-L/16 (frozen, patch_size=16)
│ last block tokens (B, 201, 1024) ──flatten──► (B, 205824)
│ tokens: 196 patch + 1 CLS + 4 storage
└─ head : UNet-MLP
205824 ──enc1──► 512 ──enc2──► 512
↓ bottleneck 512
512 ◄──dec2── 1024 (cat)
512 ◄──dec1── 1024 (cat) skip connections from enc1, enc2
↓ out linear
2 (x_metres, y_metres)
──coord_convert──► (lon, lat) degrees
```
The output is a flat-earth offset in metres relative to Jakarta origin
`(lon 106.828320, lat -6.227468)`, inverted to degrees at inference time.
The published checkpoint bundles the **frozen backbone + trained head** — no
separate LVD-1689M weight download is required. Only the DINOv3 **code** is
pulled from `torch.hub` on first load.
## Performance
Evaluated on **137,173 test samples** (Google Street View perspective crops of
Jakarta roads, 8 headings: 0°–315° in 45° steps).
| Epoch | Mean error | Median error | % < 1 km | % < 5 km | % < 25 km |
|-------|-----------|-------------|---------|---------|---------|
| e15 | 3.313 km | 1.966 km | 25.6% | 79.4% | 100.0% |
| e35 | 3.158 km | 1.588 km | 33.8% | 80.2% | 100.0% |
| **e40** | **2.743 km** | **1.272 km** | **42.2%** | **83.1%** | **100.0%** |
Error is geodesic (haversine) distance between the predicted and true GPS
coordinate. **This checkpoint is epoch 40** (`modelD_40e_2.pth`).
## Training
| Hyperparameter | Value |
|---|---|
| Backbone | DINOv3 ViT-L/16 (`facebookresearch/dinov3`) — frozen |
| Head | UNet-MLP, hidden 512, skip connections |
| Optimiser | AdamW |
| Learning rate | 1e-3 |
| Weight decay | 1e-4 |
| Batch size | 64 |
| Epochs | 40 |
| Input size | 224 × 224 |
| Loss | MSE on flat-earth (x, y) metres |
**Dataset:** [`nadh0708/JKTSV-primary`](https://huggingface.co/datasets/nadh0708/JKTSV-primary) —
685,848 perspective crops (548,675 train / 137,173 test) sampled along
Jakarta motorway and primary road segments at 50 m intervals, 8 headings.
## Preprocessing
Images are resized to **224 × 224** and normalised with ImageNet statistics:
```python
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
```
`GeoTagPredictor` applies this automatically. If you call `DinoGeoRegressor`
directly, apply the transform before passing `pixel_values`.
## Limitations
- **Geographically restricted to Jakarta.** Out-of-domain images return
coordinates near the projection origin.
- Trained on 2023+ Street View panoramas projected to 8 fixed headings
(0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°).
- Backbone is frozen; only the 3 M-parameter regression head was trained.
- Hard examples (48% of the test set, defined as error > 1 km across all
measured epochs) tend to cluster in peripheral and coastal areas with
low visual distinctiveness.
## Files
| File | Purpose |
|---|---|
| `modeling_geotag.py` | Model classes + `from_pretrained` loader + coordinate conversion |
| `inference.py` | `GeoTagPredictor` high-level API + CLI |
| `model.safetensors` | Full state dict (frozen backbone + trained head, ~1.2 GB) |
| `requirements.txt` | Runtime dependencies |
## Citation
If you use this model or the dataset, please cite:
```bibtex
@misc{img2locjakarta2025,
author = {Nadhif},
title = {Img2LocJakarta: DINOv3 Street-View Geolocation for Jakarta},
year = {2025},
url = {https://huggingface.co/nadh0708/Img2LocJakarta}
}
```
## License
This model is released under the **Apache License 2.0**, consistent with the
DINOv3 backbone license. See `LICENSE` for details.
|