Abstract
As an effective feature extractor, Vision Transformer (ViT) has been widely applied to both image classification and object tracking tasks. In this paper, we revisit and enhance the classic Data-efficient image Transformer (DeiT) for these two tasks. The DeiT is optimized step-by-step across different modules, including its patch stem, position embedding, and the development of efficient linear attention mechanisms. To address the performance degradation of linear attention, we propose Shift Expansion Linear Attention (SELA) which generates new heads with rich feature diversity through a simple but efficient cyclic shift operation. Additionally, SELA similarity minimization is added to cross-entropy loss to further enhance feature diversity. Based on these improvements, we develop SELA-ViT for image classification and further build SELA-Track for object tracking. With comparable model size and speed, SELA-ViT-T achieves a +4.8% improvement in Top-1 accuracy over DeiT-T on ImageNet-1K and establishes a new state-of-the-art performance among linear attention methods. Furthermore, we validate SELA-ViT on five small datasets. On four benchmark object tracking datasets, SELA-Track exhibits improved tracking performance.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Circuits and Systems for Video Technology |
| DOIs | |
| State | Accepted/In press - 2026 |
Keywords
- Data-efficient image Transformer (DeiT)
- image classification
- object tracking
- Shift Expansion Linear Attention (SELA)
- Vision Transformer (ViT)
Fingerprint
Dive into the research topics of 'Enhancing Vision Transformer with Shift Expansion Linear Attention for Image Classification and Object Tracking'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver