Skip to main navigation Skip to search Skip to main content

SDA: Structure-aware Distribution Alignment for Vision-Language Models

  • Lin Peng
  • , Cong Wan
  • , Shaokun Wang
  • , Yuhang He
  • , Yihong Gong
  • Xi'an Jiaotong University
  • Xi'an Jiaotong University

Research output: Contribution to journalArticlepeer-review

Abstract

This paper addresses a critical limitation in existing CLIP adaptation methods: the tendency to form tight text feature clusters that struggle to distinguish semantically similar categories and show limited generalization. We propose Structureaware Distribution Alignment (SDA), a novel framework that simultaneously optimizes inter-class separation and preserves intra-class diversity. SDA consists of three key components: (1) Equiangular Prototype Optimization to maximize class separability by arranging text prototypes in an optimal geometric configuration; (2) Text Distribution Modeling through Cross- Category Feature Synthesis and selective KNN modeling to create comprehensive class distributions with natural variations; and (3) Text-Image Structural Knowledge Alignment to transfer the optimized text structures to image features. Additionally, we introduce a parameter-efficient bias-adapter architecture that reduces parameters by approximately 50% without performance degradation. Extensive experiments across few-shot learning, domain generalization, and robust classification transfer demonstrate that our approach achieves strong and competitive performance, with clear improvements in multiple challenging settings.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026

Keywords

  • Distribution Alignment
  • Domain Generation
  • Few-shot learning
  • Structure-aware
  • Vision- Language Models

Fingerprint

Dive into the research topics of 'SDA: Structure-aware Distribution Alignment for Vision-Language Models'. Together they form a unique fingerprint.

Cite this