Aligned Pseudo Feature Generation for Zero-Shot Object Detection

摘要

The goal of zero-shot object detection (ZSD) is to localize and classify objects that haven’t been presented in the training set. Drawing inspiration from the success of generative models, many feature generation based approaches have been explored for the ZSD problem, which had some promising results. However, the absence of visual samples from unseen classes inevitably leads to subpar synthesized features, thereby constraining the generator’s efficacy in prior approaches. This limitation highlights a key challenge in current ZSD methods: generating discriminative and realistic visual features without real data support. To address this, our method is motivated by the insight that semantic embeddings alone are insufficient for generating high-quality features. Instead, we align them with pseudo-visual features extracted from a diffusion model, which provides more visually plausible guidance and significantly improves the quality of synthesized features for unseen categories. Additionally, unlike previous works, the real visual features of seen classes is also used to train the unseen classifier, which is considered as the background class. Such a trained unseen classifier has less chance to misclassify a seen object into an unseen category. While working with the pretrained seen classifier, it performs better in the real world application, i.e. the generalized ZSD situation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our method achieves state-of-the-art performance.

出版物
IEEE Transactions on Cognitive and Developmental Systems
戴鑫淼
戴鑫淼
硕士生

研究方向为零样本学习与目标检测,围绕伪特征生成、合成特征质量评估等问题提升零样本目标检测性能。

李哲浩
李哲浩
硕士生

研究方向为人物交互检测与目标检测,提出双重查询增强和上下文表征学习方法,提升人-物交互语义理解与检测性能。

王 翀
王 翀
副教授

研究兴趣:人机交互、人工智能、计算机视觉、多媒体计算.

陈翊
陈翊
硕士生

研究方向为开放词汇感知与目标检测,关注开放类别场景下的视觉识别、语义理解与检测泛化。