摘要
Remote sensing scene categorisation is a task to distinguish the basic level scene images in accordance with the contents of the subordinate level feature representations. This gives rise to a significant semantic gap between subordinate level features and the basic level scene contents. In this paper, we propose recurrent transformer networks (RTN) to mitigate the above problem. RTN incorporates learning transformation-invariant regions with transformer based attention mechanism, thus reducing the semantic gap efficiently. It also can learn the canonical appearance for the most relevant regions based on the subordinate level contents of the remote sensing scene images. The predictions of both transformation parameters and classification score are derived from the bilinear CNN pooling regression. The whole network is differentiable and can be learned end-by-end by only acquiring the basic level labels. Through extensive experiments, we demonstrate that our RTN is able to achieve state-of-the-art performance on several public remote sensing scene datasets.
| 源语言 | 英语 |
|---|---|
| 出版状态 | 已出版 - 1 1月 2018 |
| 活动 | 29th British Machine Vision Conference, BMVC 2018 - Newcastle, 英国 期限: 3 9月 2018 → 6 9月 2018 |
会议
| 会议 | 29th British Machine Vision Conference, BMVC 2018 |
|---|---|
| 国家/地区 | 英国 |
| 市 | Newcastle |
| 时期 | 3/09/18 → 6/09/18 |
学术指纹
探究 'Recurrent transformer networks for remote sensing scene categorisation' 的科研主题。它们共同构成独一无二的学术指纹。引用此
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver