Img2LocJakarta / README.md
nadh0708's picture
Upload folder using huggingface_hub
1bb5a3e verified
|
Raw
History Blame Contribute Delete
5.89 kB
---
license: apache-2.0
tags:
- geolocation
- image-regression
- dinov3
- street-view
- jakarta
- computer-vision
library_name: pytorch
language:
- id
datasets:
- nadh0708/JKTSV-primary
---
# Img2LocJakarta — DINOv3 street-view geolocation for Jakarta
Predict **(latitude, longitude)** of a Google Street View perspective crop taken
on a Jakarta road. A frozen **DINOv3 ViT-L/16** backbone extracts features; a
U-shaped MLP head regresses a local flat-earth (x, y) offset which is converted
to WGS-84 degrees.
| | |
|---|---|
| **Input** | RGB street-level image (any size, resized to 224×224) |
| **Output** | `{"lat": float, "lon": float}` — WGS-84 degrees |
| **Scope** | Trained on Jakarta (5 administrative cities, motorway + primary roads) |
| **Backbone** | DINOv3 ViT-L/16 — frozen; only the regression head was trained |
## Quick Start
```bash
pip install torch torchvision pillow huggingface_hub safetensors
```
```python
from huggingface_hub import hf_hub_download
import sys, os
# Download the inference code
for fname in ["inference.py", "modeling_geotag.py"]:
hf_hub_download("nadh0708/Img2LocJakarta", fname, local_dir=".")
from inference import GeoTagPredictor
predictor = GeoTagPredictor("nadh0708/Img2LocJakarta")
print(predictor.predict("street.jpg"))
# {'lat': -6.2261, 'lon': 106.8123}
# Batched
print(predictor.predict(["a.jpg", "b.jpg"]))
```
### From a local checkpoint
```python
predictor = GeoTagPredictor("modelD_40e_2.pth")
```
### CLI
```bash
python inference.py nadh0708/Img2LocJakarta street.jpg
```
### Low-level API
```python
import torch
from torchvision import transforms
from PIL import Image
from modeling_geotag import DinoGeoRegressor
model = DinoGeoRegressor.from_pretrained("nadh0708/Img2LocJakarta")
model.eval()
tf = transforms.Compose([
transforms.Resize((224, 224)),
transforms.ToTensor(),
transforms.Normalize([0.485, 0.456, 0.406], [0.229, 0.224, 0.225]),
])
x = tf(Image.open("street.jpg").convert("RGB")).unsqueeze(0)
lonlat = model.predict_lonlat(x) # tensor([[lon, lat]])
```
## Architecture
```
DinoGeoRegressor
├─ encoder : DINOv3 ViT-L/16 (frozen, patch_size=16)
│ last block tokens (B, 201, 1024) ──flatten──► (B, 205824)
│ tokens: 196 patch + 1 CLS + 4 storage
└─ head : UNet-MLP
205824 ──enc1──► 512 ──enc2──► 512
↓ bottleneck 512
512 ◄──dec2── 1024 (cat)
512 ◄──dec1── 1024 (cat) skip connections from enc1, enc2
↓ out linear
2 (x_metres, y_metres)
──coord_convert──► (lon, lat) degrees
```
The output is a flat-earth offset in metres relative to Jakarta origin
`(lon 106.828320, lat -6.227468)`, inverted to degrees at inference time.
The published checkpoint bundles the **frozen backbone + trained head** — no
separate LVD-1689M weight download is required. Only the DINOv3 **code** is
pulled from `torch.hub` on first load.
## Performance
Evaluated on **137,173 test samples** (Google Street View perspective crops of
Jakarta roads, 8 headings: 0°–315° in 45° steps).
| Epoch | Mean error | Median error | % < 1 km | % < 5 km | % < 25 km |
|-------|-----------|-------------|---------|---------|---------|
| e15 | 3.313 km | 1.966 km | 25.6% | 79.4% | 100.0% |
| e35 | 3.158 km | 1.588 km | 33.8% | 80.2% | 100.0% |
| **e40** | **2.743 km** | **1.272 km** | **42.2%** | **83.1%** | **100.0%** |
Error is geodesic (haversine) distance between the predicted and true GPS
coordinate. **This checkpoint is epoch 40** (`modelD_40e_2.pth`).
## Training
| Hyperparameter | Value |
|---|---|
| Backbone | DINOv3 ViT-L/16 (`facebookresearch/dinov3`) — frozen |
| Head | UNet-MLP, hidden 512, skip connections |
| Optimiser | AdamW |
| Learning rate | 1e-3 |
| Weight decay | 1e-4 |
| Batch size | 64 |
| Epochs | 40 |
| Input size | 224 × 224 |
| Loss | MSE on flat-earth (x, y) metres |
**Dataset:** [`nadh0708/JKTSV-primary`](https://huggingface.co/datasets/nadh0708/JKTSV-primary) —
685,848 perspective crops (548,675 train / 137,173 test) sampled along
Jakarta motorway and primary road segments at 50 m intervals, 8 headings.
## Preprocessing
Images are resized to **224 × 224** and normalised with ImageNet statistics:
```python
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
```
`GeoTagPredictor` applies this automatically. If you call `DinoGeoRegressor`
directly, apply the transform before passing `pixel_values`.
## Limitations
- **Geographically restricted to Jakarta.** Out-of-domain images return
coordinates near the projection origin.
- Trained on 2023+ Street View panoramas projected to 8 fixed headings
(0°, 45°, 90°, 135°, 180°, 225°, 270°, 315°).
- Backbone is frozen; only the 3 M-parameter regression head was trained.
- Hard examples (48% of the test set, defined as error > 1 km across all
measured epochs) tend to cluster in peripheral and coastal areas with
low visual distinctiveness.
## Files
| File | Purpose |
|---|---|
| `modeling_geotag.py` | Model classes + `from_pretrained` loader + coordinate conversion |
| `inference.py` | `GeoTagPredictor` high-level API + CLI |
| `model.safetensors` | Full state dict (frozen backbone + trained head, ~1.2 GB) |
| `requirements.txt` | Runtime dependencies |
## Citation
If you use this model or the dataset, please cite:
```bibtex
@misc{img2locjakarta2025,
author = {Nadhif},
title = {Img2LocJakarta: DINOv3 Street-View Geolocation for Jakarta},
year = {2025},
url = {https://huggingface.co/nadh0708/Img2LocJakarta}
}
```
## License
This model is released under the **Apache License 2.0**, consistent with the
DINOv3 backbone license. See `LICENSE` for details.