跳到主要导航 跳到搜索 跳到主要内容

Enhancing Vision Transformer with Shift Expansion Linear Attention for Image Classification and Object Tracking

  • Xi'an Jiaotong University
  • Zhejiang University

科研成果: 期刊稿件文章同行评审

摘要

As an effective feature extractor, Vision Transformer (ViT) has been widely applied to both image classification and object tracking tasks. In this paper, we revisit and enhance the classic Data-efficient image Transformer (DeiT) for these two tasks. The DeiT is optimized step-by-step across different modules, including its patch stem, position embedding, and the development of efficient linear attention mechanisms. To address the performance degradation of linear attention, we propose Shift Expansion Linear Attention (SELA) which generates new heads with rich feature diversity through a simple but efficient cyclic shift operation. Additionally, SELA similarity minimization is added to cross-entropy loss to further enhance feature diversity. Based on these improvements, we develop SELA-ViT for image classification and further build SELA-Track for object tracking. With comparable model size and speed, SELA-ViT-T achieves a +4.8% improvement in Top-1 accuracy over DeiT-T on ImageNet-1K and establishes a new state-of-the-art performance among linear attention methods. Furthermore, we validate SELA-ViT on five small datasets. On four benchmark object tracking datasets, SELA-Track exhibits improved tracking performance.

学术指纹

探究 'Enhancing Vision Transformer with Shift Expansion Linear Attention for Image Classification and Object Tracking' 的科研主题。它们共同构成独一无二的学术指纹。

引用此