Skip to main navigation Skip to search Skip to main content

LEViT: Locally Enhanced Vision Transformer for Efficient Object Re-identification

  • Shenqi Lai
  • , Yuhui Wang
  • , Mingyuan Fan
  • , Junshi Huang
  • , Haifeng Liu
  • , Deng Cai
  • , Xueming Qian
  • , Yaxiong Wang
  • Zhejiang University
  • Xi'an Jiaotong University
  • Meituan
  • Zhibian Technology Co. Ltd.
  • Hefei University of Technology

Research output: Contribution to journalArticlepeer-review

Abstract

Vision Transformer (ViT) on object re-identification (ReID) has attracted significant attention recently. However, ViT-based ReID substantially increases computational complexity, imposing significant burdens during training and inference. This paper presents an efficient and effective ViT-based backbone for ReID tasks, called the Locally Enhanced Vision Transformer (LEViT). ViT models typically emphasize global relationship modeling, yet ReID tasks are more sensitive to local information. To address this gap, we propose a Locally Enhanced (LE) block to enhance local information by performing self-attention within local split windows. Since part-based models dominate ReID, calculating self-attention across all patches is computationally inefficient. We also replace the traditional Query-Key-Value projector with the Group Convolution (G-Conv) projector, enabling the model to capture local details more efficiently. Furthermore, G-Conv is integrated into the channel MLP to strengthen local feature sensitivity. Using these components, we develop two LEViT variants: LEViT-S and LEViT-L. To our knowledge, LEViT is the first highly adaptable ViT backbone for ReID tasks. Experimental evaluations demonstrate the effectiveness in five ReID datasets: Market1501, DukeMTMC, MSMT17, VeRi-776, and VehicleID. Notably, LEViT-S outperforms TransReID while requiring less than 10% computational complexity. Furthermore, LEViT obtains the state-of-the-art on three deep metric learning datasets: CUB-200-2011, Cars196, and University-1652.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2025

Keywords

  • Deep Metric Learning
  • Efficient
  • Re-identification
  • Vision Transformer

Fingerprint

Dive into the research topics of 'LEViT: Locally Enhanced Vision Transformer for Efficient Object Re-identification'. Together they form a unique fingerprint.

Cite this