Skip to main navigation Skip to search Skip to main content

Window-to-window BEV representation learning for limited FoV cross-view geo-localization

  • Lei Cheng
  • , Daikun Liu
  • , Lingquan Meng
  • , Teng Wang
  • , Changyin Sun
  • Southeast University, Nanjing

Research output: Contribution to journalArticlepeer-review

Abstract

Cross-view geo-localization confronts significant challenges due to large perspective changes, especially when the ground-view query image has a limited field of view (FoV) with unknown orientation. To bridge the cross-view domain gap, we explore learning the bird’s eye view (BEV) representation directly from the ground feature. However, the unknown orientation of ground images combined with the absence of camera parameters leads to ambiguity between BEV queries and ground references. To tackle this challenge, we propose a novel Window-to-Window BEV representation learning framework, termed W2W-BEV, which adaptively matches BEV queries to ground reference at window-scale and generates the BEV feature for each grid by attending to the semantically relevant ground features within the matched window. Additionally, we employ depth-guide BEV initialization to project the ground feature into BEV space, thus facilitating the high-quality window matching. By performing window-scale matching in a learnable manner, our W2W-BEV allows the BEV space to learn features by borrowing information from limited FoV, and aligning it with satellite imagery. Extensive experimental results on benchmark datasets demonstrate significant superiority of our W2W-BEV over previous state-of-the-art methods under challenging conditions of unknown orientation and limited FoV.

Original languageEnglish
Article number109099
JournalNeural Networks
Volume203
DOIs
StatePublished - Nov 2026
Externally publishedYes

Keywords

  • Bird’s eye view
  • Cross-view geo-localization
  • Limited field of view

Fingerprint

Dive into the research topics of 'Window-to-window BEV representation learning for limited FoV cross-view geo-localization'. Together they form a unique fingerprint.

Cite this