跳到主要导航 跳到搜索 跳到主要内容

SDA: Structure-aware Distribution Alignment for Vision-Language Models

  • Lin Peng
  • , Cong Wan
  • , Shaokun Wang
  • , Yuhang He
  • , Yihong Gong
  • Xi'an Jiaotong University
  • Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

摘要

This paper addresses a critical limitation in existing CLIP adaptation methods: the tendency to form tight text feature clusters that struggle to distinguish semantically similar categories and show limited generalization. We propose Structureaware Distribution Alignment (SDA), a novel framework that simultaneously optimizes inter-class separation and preserves intra-class diversity. SDA consists of three key components: (1) Equiangular Prototype Optimization to maximize class separability by arranging text prototypes in an optimal geometric configuration; (2) Text Distribution Modeling through Cross- Category Feature Synthesis and selective KNN modeling to create comprehensive class distributions with natural variations; and (3) Text-Image Structural Knowledge Alignment to transfer the optimized text structures to image features. Additionally, we introduce a parameter-efficient bias-adapter architecture that reduces parameters by approximately 50% without performance degradation. Extensive experiments across few-shot learning, domain generalization, and robust classification transfer demonstrate that our approach achieves strong and competitive performance, with clear improvements in multiple challenging settings.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'SDA: Structure-aware Distribution Alignment for Vision-Language Models' 的科研主题。它们共同构成独一无二的学术指纹。

引用此