跳到主要导航 跳到搜索 跳到主要内容

Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global Attention

  • Xi'an Jiaotong University
  • Pengcheng Laboratory

科研成果: 书/报告/会议事项章节会议稿件同行评审

1 引用 (Scopus)

摘要

Text-guided zero-shot object counting leverages visionlanguage models (VLMs) to count objects of an arbitrary class given by a text prompt. Existing approaches for this challenging task only utilize local patch-level features to fuse with text feature, ignoring the important influence of the global image-level feature. In this paper, we propose a universal strategy that can exploit both local patchlevel features and global image-level feature simultaneously. Specifically, to improve the localization ability of VLMs, we propose Text-guided Local Ranking. Depending on the prior knowledge that foreground patches have higher similarity with the text prompt, a new local-text rank loss is designed to increase the differences between the similarity scores of foreground and background patches which push foreground and background patches apart. To enhance the counting ability of VLMs, Number-evoked Global Attention is introduced to first align global image-level feature with multiple number-conditioned text prompts. Then, the one with the highest similarity is selected to compute cross-attention with the global image-level feature. Through extensive experiments on widely used datasets and methods, the proposed approach has demonstrated superior advancements in performance, generalization, and scalability. Furthermore, to better evaluate text-guided zeroshot object counting methods, we propose a dataset named ZSC-8K, which is larger and more challenging, to establish a new benchmark. Codes and dataset are released at https://github.com/zaqai/LGCount.

源语言英语
主期刊名Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
出版商Institute of Electrical and Electronics Engineers Inc.
21097-21106
页数10
ISBN(电子版)9798331587758
DOI
出版状态已出版 - 2025
已对外发布
活动2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, 美国
期限: 19 10月 202523 10月 2025

丛书

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499
ISSN(电子版)2380-7504

会议

会议2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
国家/地区美国
Honolulu
时期19/10/2523/10/25

学术指纹

探究 'Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global Attention' 的科研主题。它们共同构成独一无二的学术指纹。

引用此