跳到主要导航 跳到搜索 跳到主要内容

FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training

  • Xi'an Jiaotong University
  • Shenzhen University of Advanced Technology
  • Guangdong Provincial Key Laboratory of Computility Microelectronics
  • Shenzhen Institute of Advanced Technology

科研成果: 期刊稿件会议文章同行评审

3 引用 (Scopus)

摘要

Language-image pre-training faces significant challenges due to limited data in specific formats and the constrained capacities of text encoders. While prevailing methods attempt to address these issues through data augmentation and architecture modifications, they continue to struggle with processing long-form text inputs, and the inherent limitations of traditional CLIP text encoders lead to suboptimal downstream generalization. In this paper, we propose FLAME (Frozen Large lAnguage Models Enable data-efficient language-image pre-training) that leverages frozen large language models as text encoders, naturally processing long text inputs and demonstrating impressive multilingual generalization. FLAME comprises two key components: 1) a multifaceted prompt distillation technique for extracting diverse semantic representations from long captions, which better aligns with the multifaceted nature of images, and 2) a facet-decoupled attention mechanism, complemented by an offline embedding strategy, to ensure efficient computation. Extensive empirical evaluations demonstrate FLAME's superior performance. When trained on CC3M, FLAME surpasses the previous state-of-the-art by 4.9% in ImageNet top-1 accuracy. On YFCC15M, FLAME surpasses the WIT-400M-trained CLIP by 44.4% in average image-to-text recall@1 across 36 languages, and by 34.6% in text-to-image recall@1 for long-context retrieval on Urban-1k. Code is available at https://github.com/MIV-XJTU/FLAME.

源语言英语
页(从-至)4080-4090
页数11
期刊Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
DOI
出版状态已出版 - 2025
活动2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025 - Nashville, 美国
期限: 11 6月 202515 6月 2025

学术指纹

探究 'FLAME: Frozen Large Language Models Enable Data-Efficient Language-Image Pre-training' 的科研主题。它们共同构成独一无二的学术指纹。

引用此