TY - GEN
T1 - Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global Attention
AU - Zhang, Shiwei
AU - Zhou, Qi
AU - Ke, Wei
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Text-guided zero-shot object counting leverages visionlanguage models (VLMs) to count objects of an arbitrary class given by a text prompt. Existing approaches for this challenging task only utilize local patch-level features to fuse with text feature, ignoring the important influence of the global image-level feature. In this paper, we propose a universal strategy that can exploit both local patchlevel features and global image-level feature simultaneously. Specifically, to improve the localization ability of VLMs, we propose Text-guided Local Ranking. Depending on the prior knowledge that foreground patches have higher similarity with the text prompt, a new local-text rank loss is designed to increase the differences between the similarity scores of foreground and background patches which push foreground and background patches apart. To enhance the counting ability of VLMs, Number-evoked Global Attention is introduced to first align global image-level feature with multiple number-conditioned text prompts. Then, the one with the highest similarity is selected to compute cross-attention with the global image-level feature. Through extensive experiments on widely used datasets and methods, the proposed approach has demonstrated superior advancements in performance, generalization, and scalability. Furthermore, to better evaluate text-guided zeroshot object counting methods, we propose a dataset named ZSC-8K, which is larger and more challenging, to establish a new benchmark. Codes and dataset are released at https://github.com/zaqai/LGCount.
AB - Text-guided zero-shot object counting leverages visionlanguage models (VLMs) to count objects of an arbitrary class given by a text prompt. Existing approaches for this challenging task only utilize local patch-level features to fuse with text feature, ignoring the important influence of the global image-level feature. In this paper, we propose a universal strategy that can exploit both local patchlevel features and global image-level feature simultaneously. Specifically, to improve the localization ability of VLMs, we propose Text-guided Local Ranking. Depending on the prior knowledge that foreground patches have higher similarity with the text prompt, a new local-text rank loss is designed to increase the differences between the similarity scores of foreground and background patches which push foreground and background patches apart. To enhance the counting ability of VLMs, Number-evoked Global Attention is introduced to first align global image-level feature with multiple number-conditioned text prompts. Then, the one with the highest similarity is selected to compute cross-attention with the global image-level feature. Through extensive experiments on widely used datasets and methods, the proposed approach has demonstrated superior advancements in performance, generalization, and scalability. Furthermore, to better evaluate text-guided zeroshot object counting methods, we propose a dataset named ZSC-8K, which is larger and more challenging, to establish a new benchmark. Codes and dataset are released at https://github.com/zaqai/LGCount.
UR - https://www.scopus.com/pages/publications/105044198190
U2 - 10.1109/ICCV51701.2025.01960
DO - 10.1109/ICCV51701.2025.01960
M3 - 会议稿件
AN - SCOPUS:105044198190
T3 - Proceedings of the IEEE International Conference on Computer Vision
SP - 21097
EP - 21106
BT - Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
Y2 - 19 October 2025 through 23 October 2025
ER -