TY - JOUR
T1 - Visual-Semantic Aligned Bidirectional Network for Zero-Shot Learning
AU - Gao, Rui
AU - Hou, Xingsong
AU - Qin, Jie
AU - Shen, Yuming
AU - Long, Yang
AU - Liu, Li
AU - Zhang, Zhao
AU - Shao, Ling
N1 - Publisher Copyright:
© 2021 IEEE.
PY - 2023
Y1 - 2023
N2 - Zero-shot learning (ZSL) aims to recognize unknown categories that are unavailable during training. Recently, generative models have shown the potential to address this challenging problem by synthesizing unseen features conditioned on semantic embeddings such as attributes. However, unidirectional generative models cannot guarantee the effective coupling between visual and semantic spaces. To this end, we propose a visual-semantic aligned bidirectional network with cycle consistency to alleviate the gap between these two spaces, generating unseen features of high quality. More importantly, we incorporate two carefully designed strategies into our bidirectional framework to improve the overall ZSL performance. Specifically, we enhance the intra-domain class divergence in both visual and semantic spaces, and in the meantime, mitigate the inter-domain shift to preserve seen-unseen domain discrimination. Experimental results on four standard benchmarks show the superiority of our framework over existing state-of-the-art methods under both conventional and generalized ZSL settings.
AB - Zero-shot learning (ZSL) aims to recognize unknown categories that are unavailable during training. Recently, generative models have shown the potential to address this challenging problem by synthesizing unseen features conditioned on semantic embeddings such as attributes. However, unidirectional generative models cannot guarantee the effective coupling between visual and semantic spaces. To this end, we propose a visual-semantic aligned bidirectional network with cycle consistency to alleviate the gap between these two spaces, generating unseen features of high quality. More importantly, we incorporate two carefully designed strategies into our bidirectional framework to improve the overall ZSL performance. Specifically, we enhance the intra-domain class divergence in both visual and semantic spaces, and in the meantime, mitigate the inter-domain shift to preserve seen-unseen domain discrimination. Experimental results on four standard benchmarks show the superiority of our framework over existing state-of-the-art methods under both conventional and generalized ZSL settings.
KW - Bidirectional network
KW - generative model
KW - zero-shot learning
UR - https://www.scopus.com/pages/publications/85123789326
U2 - 10.1109/TMM.2022.3145666
DO - 10.1109/TMM.2022.3145666
M3 - 文章
AN - SCOPUS:85123789326
SN - 1520-9210
VL - 25
SP - 1649
EP - 1664
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -