Skip to main navigation Skip to search Skip to main content

Region Guide Grid Cross Transformer for Image Caption

  • Jiayu Bai
  • , Lihua Tian
  • , Qingyang Zhang
  • , Chen Li
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

In image captioning tasks, many studies have shown that using both grid feature and region feature from images helps models better understand visual content, leading to more accurate descriptions. However, to save training time and keep feature extraction efficient, most research uses pre-trained models to get these grid and region feature. This method could provide the model with diverse feature, but the pre-trained models used for feature extraction were trained for different purposes, leading to variations in their focus. As a result, many of the extracted visual feature may not be well-suited for the current task, introducing a significant amount of redundant information. To resolve the feature discrepancies and redundancy caused by the differing focuses of these models, we propose a model named Region Guide Grid Cross Transformer (RGGT) for image captioning. In our model, since region feature tend to lose more global visual-semantic information compared to grid feature, the model primarily uses grid feature as the main during encoding stage. We use multi-head cross-attention mechanism that allows region feature to guide the grid feature, generating new grid feature enriched with both global semantics and target-region semantics. Furthermore, a feature refinement module based on sparse scan attention is introduced to purify the visual feature and produce new region feature derived from the refined new grid feature. In the decoding stage, to better use the target region semantics from the new region feature while preserving global feature, we further integrate and control redundancy between the new grid feature and new region feature. To achieve this, we propose a feature deep fusion module based on a gate mechanism. This module combines text feature with both region and grid feature through their respective multi-head cross attention mechanisms. Using a gate mechanism, it automatically learns to control the proportion of each feature in the final fusion, enabling more accurate integration of the different feature information. We evaluate our RGGT model on the MSCOCO2014 dataset, with experimental results demonstrating its outstanding performance. The model significantly outperforms both comparable approaches and state-of-the-art methods. The code will be made available on https://github.com/Kickdog1022/RGGT_image_caption.

Original languageEnglish
Article numbere70622
JournalConcurrency and Computation: Practice and Experience
Volume38
Issue number6
DOIs
StatePublished - Mar 2026

Keywords

  • deep fusion
  • image captioning
  • sparse scan attention
  • transformer

Fingerprint

Dive into the research topics of 'Region Guide Grid Cross Transformer for Image Caption'. Together they form a unique fingerprint.

Cite this