Papers
arxiv:2604.11415

Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding

Published on Apr 13
Authors:
,
,
,
,
,
,

Abstract

A unified cost-aware framework couples fine-grained high-resolution sampling with cross-patch representation prediction for efficient remote sensing understanding, supported by a large-scale cross-resolution pretraining dataset.

Remote sensing understanding inherently requires multi-resolution observation, since different targets and application tasks demand different levels of spatial detail. While low-resolution (LR) imagery enables efficient global observation, high-resolution (HR) imagery provides critical local details at a much higher acquisition cost and with limited coverage. This motivates a cross-scale sensing strategy that selectively acquires HR imagery guided by LR-based global perception to improve task performance under constrained cost. Existing HR sampling methods typically make selection decisions from isolated LR patches, thereby ignoring fine-grained intra-patch importance and cross-patch contextual interactions, leading to fragmented feature representation and suboptimal scene reasoning under sparse HR observations. To address this issue, we formulate cross-scale remote sensing understanding as a unified cost-aware problem that couples fine-grained HR sampling with cross-patch representation prediction, enabling more effective task reasoning with fewer HR observations. Furthermore, we present GL-10M, a high- and low-resolution dataset with nearly 100,000 scene pairs and 10 million images for large-scale cross-resolution pretraining. Extensive experiments on recognition and retrieval tasks show that our method consistently achieves a superior performance-cost trade-off. The code is publicly available at https://github.com/xzhacc/CrossSO.

Community

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2604.11415
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2604.11415 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2604.11415 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2604.11415 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.