Skip to main navigation Skip to search Skip to main content

Advances in open vocabulary perception for remote sensing images

  • Xi'an Jiaotong University
  • Xi'an Jiaotong University
  • School of Mathematics and Statistics

Research output: Contribution to journalArticlepeer-review

Abstract

Remote sensing technology serves as the core mechanism for the observation of the Earth and the understanding of surface environments. It plays an irreplaceable role in critical fields, such as natural disaster monitoring, urban planning, resource exploration, and ecological protection. Driven by the rapid advancement of deep learning over the past decade, the intelligent interpretation of remote sensing images has achieved breakthrough progress in fundamental vision tasks. However, the traditional deep learning paradigm is intrinsically built upon a close-set assumption, i.e., that models can only recognize a predefined and human-annotated set of fixed categories during the inference stage. When confronted with highly complex surface environments in real-world Earth observation scenarios, dynamic object morphology, and rare ground objects with long-tail distributions, this traditional paradigm not only incurs predialitive costs for the construction of massive pixel-level annotated datasets but also easily falls into the trap of domain-specific overfitting. Consequently, the generalization and response capabilities of this paradigm are severely challenged by unseen categories or sudden events, making this paradigm inadequate for meeting the highly dynamic interpretation demands of the open world. In recent years, the rapid development of vision-language models has catalyzed a paradigm shift in artificial intelligence from task-specific models into general-purpose perception models. By mapping visual representations and natural language into a unified feature space through contrastive learning on massive image-text pairs, these models have broken the constraints of discrete labels, enabling a direct response to arbitrary natural language prompts. This capability is known as open vocabulary perception. Although this technology has demonstrated remarkable zero-shot generalization and cross-modal reasoning capabilities in the natural image domain, the direct application of these general vision-language models to the remote sensing domain encounters a severe domain gap. The uniqueness of remote sensing data poses multiple challenges to the adaptability of existing models. First, the distinct overread imaging perspective causes drastic variations in object scale and complex background textures. Second, Earth observation tasks rely on multisource heterogeneous data from synthetic-aperture radar, multispectral or hyperspectral imaging, and thermal infrared sensors. The underlying physical mechanisms of these sensors exceed the inherent inductive biases of models that are pretrained solely on natural RGB images. Third, remote sensing objects often exhibit strong geospatial attributes and complex topological associations. To address these critical challenges, this study provides a comprehensive and systematic review of recent advancements in open vocabulary perception for remote sensing images. We first delve into the foundational aspect of this field: vision-language pretraining for remote sensing. We extensively review the evolution of construction strategies for large-scale datasets. We highlight the transition from limited, human-annotated image-text pairs into massive datasets generated via heuristic rules, the integration of geographic metadata, and advanced multimodal large language models, including innovative approaches that leverage OpenStreetMap and geographical coordinates to produce fine-grained, physics-aware descriptions across multiple modalities. Concurrently, we systematically summarize the progression of pretraining methodologies. Although early approaches have primarily focused on simple domain adaptation through continuous pretraining, recent state-of-the-art frameworks emphasize physics-aware encoding, fine-grained multilevel consistency learning, and geography-enhanced architectures. These frameworks better capture the intricate spatial relationships and modality diversities that are inherent in Earth observation data. Subsequently, this review conducts an in-depth analysis of the adaptation and optimization of open vocabulary perception techniques across a wide spectrum of crucial downstream tasks. For zero-shot scene classification and cross-modal retrieval, we discuss advanced strategies that are designed to mitigate the high intra-class similarity and complex interclass variances typical in remote sensing. We emphasize the shift toward fine-grained local-global alignment, hard negative mining, dynamic soft labeling, and prompt engineering. In the realm of open vocabulary image segmentation, we categorize the existing literature into training-based methods and training-free or annotation-free paradigms. Training-based methods leverage base categories to adapt models while preventing catastrophic forgetting through pseudo-label distillation and knowledge retention mechanisms. Training-free paradigms synergize foundational models, such as CLIP and the segment anything model, to extract structural masks and align semantics without the updating of network weights. For open vocabulary object detection and remote sensing visual grounding, we explore the approaches of researchers to deal with extreme scale variations, arbitrary orientations, and dense object distributions. These approaches include innovative frameworks for pseudo-label generation, multi-scale feature alignment, cross-modality context modeling, and interactive grounding mechanisms. Furthermore, we examine open vocabulary change detection, wherein recent studies employ either combinations of pretrained vision-language models or generative models to generate large-scale data. These approaches aim to identify arbitrary, text-specified surface transitions and simulate complex spatiotemporal changes without reliance on massive and costly litemporal pixel-level annotations. We also briefly discuss emerging open vocabulary applications in 3D urban point clouds and cross-domain archeological remote sensing, illustrating the expanding horizon of this technology. Despite remarkable progress, the field of open vocabulary perception for remote sensing remains in a crucial developmental stage and faces several critical bottlenecks. This study critically identifies the limitations of current research, including the severe scarcity of high-quality and geographically balanced training data. This scarcity leads to geographic biases and performance degradation in data-prior regions. In addition, genuinely fine-grained and long-tailed open vocabulary evaluation benchmarks that can accurately reflect the performance of a model are prominently absent in extreme or unknown real-world scenarios. The inadequate physical understanding of heterogeneous modalities and the inherent black box unreliability of current large models in high-stake decision-making scenarios further constrain practical deployments. To chart the course for future research, we outline several promising and essential trajectories. First, we anticipate a paradigm shift toward generative perception that is driven by multimodal large language models. This shift unifies various spatial localization tasks into the direct generation of coordinate sequences or geometric property tokens to utilize fully the logical reasoning capabilities of foundational models. Second, we strongly advocate for the construction of ripeness, real-world, and fine-grained evaluation systems that incorporate complex spatiotemporal logic, diverse geographic conditions, and comprehensive evaluation metrics. Third, the development of omni-modal foundation models that explicitly integrate physical priors and deep learning is considered crucial for the achievement of all-weather and all-spectrum Earth observations, moving beyond pure data-driven approaches. Furthermore, we highlight the necessity to extend perception from static spatial analysis to dynamic spatiotemporal causal reasoning to decode the evolutionary processes of the Earth. Finally, addressing the severe conflict between the massive parameter scale of foundation models and the limited computing power of aerospace edge devices requires focused research into efficient, trustworthy, and safe edge-closed collaborative computing architectures. By systematically synthesizing these advancements and challenges, this comprehensive review aims to serve as a foundational road map for researchers and practitioners. It accelerates the transition of the intelligent interpretation of remote sensing from isolated, close-set recognition toward artificial general intelligence that is capable of highly reliable, dynamic, and open world perception.

Translated title of the contribution遥感图像开放词汇感知进展
Original languageEnglish
Pages (from-to)2381-2407
Number of pages27
JournalJournal of Image and Graphics
Volume31
Issue number7
DOIs
StatePublished - 2026
Externally publishedYes

UN SDGs

This output contributes to the following UN Sustainable Development Goals (SDGs)

  1. SDG 11 - Sustainable Cities and Communities
    SDG 11 Sustainable Cities and Communities

Keywords

  • intelligent interpretation
  • open vocabulary perception
  • remote sensing image
  • vision-language model (VLM)
  • zero-shot learning

Fingerprint

Dive into the research topics of 'Advances in open vocabulary perception for remote sensing images'. Together they form a unique fingerprint.

Cite this