Skip to main navigation Skip to search Skip to main content

VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection

  • Caixia Yan
  • , Muyan Jiao
  • , Nuohan Xue
  • , Weizhan Zhang
  • , Jiahao Wang
  • , Xiaojun Chang
  • , Feng Tian
  • Xi'an Jiaotong University
  • University of Science and Technology of China

Research output: Contribution to journalArticlepeer-review

Abstract

Generative methods have shown promising performance on zero-shot object detection (ZSD) by synthesizing visual features of unseen classes from semantic embeddings. Although largely compensating for the lack of training samples, they learn the feature synthesizer of unseen classes solely based on limited training data of seen classes, leading to poor diversity and generalization ability of synthesized unseen samples. To overcome this challenge, we develop a Vision-Language Distillated Unseen Synthesizer, namely VLDUS, to build up a novel knowledge distillation-based feature generation paradigm for ZSD. To regulate the synthesized feature space, VLDUS designs two complementary generative distillation strategies that can distill rich image-text knowledge from a pre-trained CLIP model to the synthesizer. To mitigate the over-fitting towards seen classes, VLDUS performs feature-aligned generative distillation on the discriminator’s embedding space to methodically learn from the CLIP embedding space, and thus endows the synthesizer with strong generalization ability. To guarantee the intra-class diversity of synthesized unseen features, relation-aligned generative distillation is further performed to distill the diversified image-text correlations from pre-trained CLIP model to the synthesizer. Extensive experiments on MS COCO 2014, PASCAL VOC 2007/2012 and DIOR demonstrate that the proposed VLDUS can generate unseen features of both high intra-class diversity and inter-class separability, and thus outperforms state-of-the-art methods by a large margin on both ZSD and GZSD tasks. Our code is publicly available at https://github.com/Xxxnh/VLDUS.

Original languageEnglish
Article number108899
JournalNeural Networks
Volume201
DOIs
StatePublished - Sep 2026

Keywords

  • Generative adversarial networks
  • Vision-language distillation
  • Zero-shot object detection

Fingerprint

Dive into the research topics of 'VLDUS: Vision-language distillated unseen synthesizer for zero-shot object detection'. Together they form a unique fingerprint.

Cite this