TY - GEN
T1 - T-former
T2 - 30th ACM International Conference on Multimedia, MM 2022
AU - Deng, Ye
AU - Hui, Siqi
AU - Zhou, Sanping
AU - Meng, Deyu
AU - Wang, Jinjun
N1 - Publisher Copyright:
© 2022 ACM.
PY - 2022/10/10
Y1 - 2022/10/10
N2 - Benefiting from powerful convolutional neural networks (CNNs), learning-based image inpainting methods have made significant breakthroughs over the years. However, some nature of CNNs (e.g. local prior, spatially shared parameters) limit the performance in the face of broken images with diverse and complex forms. Recently, a class of attention-based network architectures, called transformer, has shown significant performance on natural language processing fields and high-level vision tasks. Compared with CNNs, attention operators are better at long-range modeling and have dynamic weights, but their computational complexity is quadratic in spatial resolution, and thus less suitable for applications involving higher resolution images, such as image inpainting. In this paper, we design a novel attention linearly related to the resolution according to Taylor expansion. And based on this attention, a network called T-former is designed for image inpainting. Experiments on several benchmark datasets demonstrate that our proposed method achieves state-of-The-Art accuracy while maintaining a relatively low number of parameters and computational complexity.
AB - Benefiting from powerful convolutional neural networks (CNNs), learning-based image inpainting methods have made significant breakthroughs over the years. However, some nature of CNNs (e.g. local prior, spatially shared parameters) limit the performance in the face of broken images with diverse and complex forms. Recently, a class of attention-based network architectures, called transformer, has shown significant performance on natural language processing fields and high-level vision tasks. Compared with CNNs, attention operators are better at long-range modeling and have dynamic weights, but their computational complexity is quadratic in spatial resolution, and thus less suitable for applications involving higher resolution images, such as image inpainting. In this paper, we design a novel attention linearly related to the resolution according to Taylor expansion. And based on this attention, a network called T-former is designed for image inpainting. Experiments on several benchmark datasets demonstrate that our proposed method achieves state-of-The-Art accuracy while maintaining a relatively low number of parameters and computational complexity.
KW - attention
KW - image inpainting
KW - neural networks
KW - transformer
UR - https://www.scopus.com/pages/publications/85148705719
U2 - 10.1145/3503161.3548446
DO - 10.1145/3503161.3548446
M3 - 会议稿件
AN - SCOPUS:85148705719
T3 - MM 2022 - Proceedings of the 30th ACM International Conference on Multimedia
SP - 6559
EP - 6568
BT - MM 2022 - Proceedings of the 30th ACM International Conference on Multimedia
PB - Association for Computing Machinery, Inc
Y2 - 10 October 2022 through 14 October 2022
ER -