跳到主要导航 跳到搜索 跳到主要内容

Dynamic Grained Encoder for Vision Transformers

  • Xi'an Jiaotong University
  • ShanghaiTech University
  • University of Chinese Academy of Sciences
  • CAS - Shanghai Institute of Microsystem and Information Technology
  • Megvii Inc. (Face++)

科研成果: 书/报告/会议事项章节会议稿件同行评审

27 引用 (Scopus)

摘要

Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained Encoder for vision transformers, which can adaptively assign a suitable number of queries to each spatial region. Thus it achieves a fine-grained representation in discriminative regions while keeping high efficiency. Besides, the dynamic grained encoder is compatible with most vision transformer frameworks. Without bells and whistles, our encoder allows the state-of-the-art vision transformers to reduce computational complexity by 40%-60% while maintaining comparable performance on image classification. Extensive experiments on object detection and segmentation further demonstrate the generalizability of our approach. Code is available at https://github.com/StevenGrove/vtpack.

源语言英语
主期刊名Advances in Neural Information Processing Systems 34 - 35th Conference on Neural Information Processing Systems, NeurIPS 2021
编辑Marc'Aurelio Ranzato, Alina Beygelzimer, Yann Dauphin, Percy S. Liang, Jenn Wortman Vaughan
出版商Neural information processing systems foundation
5770-5783
页数14
ISBN(电子版)9781713845393
出版状态已出版 - 2021
活动35th Conference on Neural Information Processing Systems, NeurIPS 2021 - Virtual, Online
期限: 6 12月 202114 12月 2021

出版系列

姓名Advances in Neural Information Processing Systems
7
ISSN(印刷版)1049-5258

会议

会议35th Conference on Neural Information Processing Systems, NeurIPS 2021
Virtual, Online
时期6/12/2114/12/21

学术指纹

探究 'Dynamic Grained Encoder for Vision Transformers' 的科研主题。它们共同构成独一无二的学术指纹。

引用此