miniHui
/

Geo-R1

Image-Text-to-Text

text-generation-inference

Model card Files Files and versions

Geo-R1 / README.md

nielsr's picture

nielsr HF Staff

Add model card for Geo-R1

aa1ea12 verified 4 months ago

|

888 Bytes

	---
	library_name: transformers
	pipeline_tag: image-text-to-text
	---

	# Geo-R1: Unlocking VLM Geospatial Reasoning with Cross-View Reinforcement Learning

	This repository contains the Geo-R1 model, a reasoning-centric post-training framework that unlocks geospatial reasoning in vision-language models, as introduced in the paper:

	[Geo-R1: Unlocking VLM Geospatial Reasoning with Cross-View Reinforcement Learning](https://huggingface.co/papers/2510.00072)

	Geo-R1 combines "thinking scaffolding" (supervised fine-tuning on synthetic chain-of-thought exemplars) and an "elevating" stage using GRPO-based reinforcement learning on a weakly-supervised cross-view pairing proxy. This approach enables models to connect visual cues with geographic priors and harness reasoning for accurate prediction, achieving state-of-the-art performance across various geospatial reasoning benchmarks.