Abstract
In image captioning tasks, many studies have shown that using both grid feature and region feature from images helps models better understand visual content, leading to more accurate descriptions. However, to save training time and keep feature extraction efficient, most research uses pre-trained models to get these grid and region feature. This method could provide the model with diverse feature, but the pre-trained models used for feature extraction were trained for different purposes, leading to variations in their focus. As a result, many of the extracted visual feature may not be well-suited for the current task, introducing a significant amount of redundant information. To resolve the feature discrepancies and redundancy caused by the differing focuses of these models, we propose a model named Region Guide Grid Cross Transformer (RGGT) for image captioning. In our model, since region feature tend to lose more global visual-semantic information compared to grid feature, the model primarily uses grid feature as the main during encoding stage. We use multi-head cross-attention mechanism that allows region feature to guide the grid feature, generating new grid feature enriched with both global semantics and target-region semantics. Furthermore, a feature refinement module based on sparse scan attention is introduced to purify the visual feature and produce new region feature derived from the refined new grid feature. In the decoding stage, to better use the target region semantics from the new region feature while preserving global feature, we further integrate and control redundancy between the new grid feature and new region feature. To achieve this, we propose a feature deep fusion module based on a gate mechanism. This module combines text feature with both region and grid feature through their respective multi-head cross attention mechanisms. Using a gate mechanism, it automatically learns to control the proportion of each feature in the final fusion, enabling more accurate integration of the different feature information. We evaluate our RGGT model on the MSCOCO2014 dataset, with experimental results demonstrating its outstanding performance. The model significantly outperforms both comparable approaches and state-of-the-art methods. The code will be made available on https://github.com/Kickdog1022/RGGT_image_caption.
| Original language | English |
|---|---|
| Article number | e70622 |
| Journal | Concurrency and Computation: Practice and Experience |
| Volume | 38 |
| Issue number | 6 |
| DOIs | |
| State | Published - Mar 2026 |
Keywords
- deep fusion
- image captioning
- sparse scan attention
- transformer
Fingerprint
Dive into the research topics of 'Region Guide Grid Cross Transformer for Image Caption'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver