TY - JOUR
T1 - Enhancing Vision Transformer with Shift Expansion Linear Attention for Image Classification and Object Tracking
AU - Zhou, Sai
AU - Liu, Meiqin
AU - Zhou, Jing
AU - Zheng, Ronghao
N1 - Publisher Copyright:
© 1991-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - As an effective feature extractor, Vision Transformer (ViT) has been widely applied to both image classification and object tracking tasks. In this paper, we revisit and enhance the classic Data-efficient image Transformer (DeiT) for these two tasks. The DeiT is optimized step-by-step across different modules, including its patch stem, position embedding, and the development of efficient linear attention mechanisms. To address the performance degradation of linear attention, we propose Shift Expansion Linear Attention (SELA) which generates new heads with rich feature diversity through a simple but efficient cyclic shift operation. Additionally, SELA similarity minimization is added to cross-entropy loss to further enhance feature diversity. Based on these improvements, we develop SELA-ViT for image classification and further build SELA-Track for object tracking. With comparable model size and speed, SELA-ViT-T achieves a +4.8% improvement in Top-1 accuracy over DeiT-T on ImageNet-1K and establishes a new state-of-the-art performance among linear attention methods. Furthermore, we validate SELA-ViT on five small datasets. On four benchmark object tracking datasets, SELA-Track exhibits improved tracking performance.
AB - As an effective feature extractor, Vision Transformer (ViT) has been widely applied to both image classification and object tracking tasks. In this paper, we revisit and enhance the classic Data-efficient image Transformer (DeiT) for these two tasks. The DeiT is optimized step-by-step across different modules, including its patch stem, position embedding, and the development of efficient linear attention mechanisms. To address the performance degradation of linear attention, we propose Shift Expansion Linear Attention (SELA) which generates new heads with rich feature diversity through a simple but efficient cyclic shift operation. Additionally, SELA similarity minimization is added to cross-entropy loss to further enhance feature diversity. Based on these improvements, we develop SELA-ViT for image classification and further build SELA-Track for object tracking. With comparable model size and speed, SELA-ViT-T achieves a +4.8% improvement in Top-1 accuracy over DeiT-T on ImageNet-1K and establishes a new state-of-the-art performance among linear attention methods. Furthermore, we validate SELA-ViT on five small datasets. On four benchmark object tracking datasets, SELA-Track exhibits improved tracking performance.
KW - Data-efficient image Transformer (DeiT)
KW - image classification
KW - object tracking
KW - Shift Expansion Linear Attention (SELA)
KW - Vision Transformer (ViT)
UR - https://www.scopus.com/pages/publications/105029968953
U2 - 10.1109/TCSVT.2026.3662708
DO - 10.1109/TCSVT.2026.3662708
M3 - 文章
AN - SCOPUS:105029968953
SN - 1051-8215
JO - IEEE Transactions on Circuits and Systems for Video Technology
JF - IEEE Transactions on Circuits and Systems for Video Technology
ER -