Skip to main navigation Skip to search Skip to main content

Generative adversarial network for semi-supervised image captioning

  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

7 Scopus citations

Abstract

Traditional supervised image captioning methods usually rely on a large number of images and paired captions for training. However, the creation of such datasets necessitates considerable temporal and human resources. Therefore, we propose a new semi-supervised image captioning algorithm to solve this problem. The proposed method uses a generative adversarial network to generate images that match captions, and uses these generated images and captions as new training data. This avoids the error accumulation problem when generating pseudo captions with autoregressive method and the network can directly perform backpropagation. At the same time, in order to ensure the correlation between the generated images and captions, we introduced the CLIP model for constraints. The CLIP model has been pre-trained on a large amount of image–text data, so it shows excellent performance in semantic alignment of images and text. To verify the effectiveness of our method, we validate on MSCOCO offline “Karpathy” test split. Experiment results show that our method can significantly improve the performance of the model when using 1% paired data, with the CIDEr score increasing from 69.5% to 77.7%. This shows that our method can effectively utilize unlabeled data for image caption tasks.

Original languageEnglish
Article number104199
JournalComputer Vision and Image Understanding
Volume249
DOIs
StatePublished - Dec 2024

Keywords

  • CLIP
  • Generative adversarial network
  • Image captioning
  • Semi-supervised
  • Transformer

Fingerprint

Dive into the research topics of 'Generative adversarial network for semi-supervised image captioning'. Together they form a unique fingerprint.

Cite this