TY - JOUR
T1 - VLDUS
T2 - Vision-language distillated unseen synthesizer for zero-shot object detection
AU - Yan, Caixia
AU - Jiao, Muyan
AU - Xue, Nuohan
AU - Zhang, Weizhan
AU - Wang, Jiahao
AU - Chang, Xiaojun
AU - Tian, Feng
N1 - Publisher Copyright:
© 2026 Elsevier Ltd. All rights are reserved, including those for text and data mining, AI training, and similar technologies.
PY - 2026/9
Y1 - 2026/9
N2 - Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator’s embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.
AB - Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator’s embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.
KW - Generative adversarial networks
KW - Vision-language distillation
KW - Zero-shot object detection
UR - https://www.scopus.com/pages/publications/105034728275
U2 - 10.1016/j.neunet.2026.108899
DO - 10.1016/j.neunet.2026.108899
M3 - 文章
AN - SCOPUS:105034728275
SN - 0893-6080
VL - 201
JO - Neural Networks
JF - Neural Networks
M1 - 108899
ER -