Abstract
Cross-view geo-localization confronts significant challenges due to large perspective changes, especially when the ground-view query image has a limited field of view (FoV) with unknown orientation. To bridge the cross-view domain gap, we explore learning the bird’s eye view (BEV) representation directly from the ground feature. However, the unknown orientation of ground images combined with the absence of camera parameters leads to ambiguity between BEV queries and ground references. To tackle this challenge, we propose a novel Window-to-Window BEV representation learning framework, termed W2W-BEV, which adaptively matches BEV queries to ground reference at window-scale and generates the BEV feature for each grid by attending to the semantically relevant ground features within the matched window. Additionally, we employ depth-guide BEV initialization to project the ground feature into BEV space, thus facilitating the high-quality window matching. By performing window-scale matching in a learnable manner, our W2W-BEV allows the BEV space to learn features by borrowing information from limited FoV, and aligning it with satellite imagery. Extensive experimental results on benchmark datasets demonstrate significant superiority of our W2W-BEV over previous state-of-the-art methods under challenging conditions of unknown orientation and limited FoV.
| Original language | English |
|---|---|
| Article number | 109099 |
| Journal | Neural Networks |
| Volume | 203 |
| DOIs | |
| State | Published - Nov 2026 |
| Externally published | Yes |
Keywords
- Bird’s eye view
- Cross-view geo-localization
- Limited field of view
Fingerprint
Dive into the research topics of 'Window-to-window BEV representation learning for limited FoV cross-view geo-localization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver