跳到主要导航 跳到搜索 跳到主要内容

Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching

  • Yuhan Liu
  • , Jingwen Fu
  • , Yang Wu
  • , Kangyi Wu
  • , Pengna Li
  • , Jiayi Wu
  • , Sanping Zhou
  • , Jingmin Xin
  • Xi'an Jiaotong University

科研成果: 书/报告/会议事项章节会议稿件同行评审

1 引用 (Scopus)

摘要

Leveraging the vision foundation models has emerged as a mainstream paradigm that improves the performance of image feature matching. However, previous works have ignored the misalignment when introducing the foundation models into feature matching. The misalignment arises from the discrepancy between the foundation models focusing on single-image understanding and the cross-image understanding requirement of feature matching. Specifically, 1) the embeddings derived from commonly used foundation models exhibit discrepancies with the optimal embeddings required for feature matching; 2) lacking an effective mechanism to leverage the single-image understanding ability into cross-image understanding. A significant consequence of the misalignment is they struggle when addressing multiinstance feature matching problems. To address this, we introduce a simple but effective framework, called IMD (Image feature Matching with a pre-trained Diffusion model) with two parts: 1) Unlike the dominant solutions employing contrastive-learning based foundation models that emphasize global semantics, we integrate the generative-based diffusion models to effectively capture instance-level details. 2) We leverage the prompt mechanism in generative model as a natural tunnel, propose a novel cross-image interaction prompting module to facilitate bidirectional information interaction between image pairs. To more accurately measure the misalignment, we propose a new benchmark called IMIM, which focuses on multi-instance scenarios. Our proposed IMD establishes a new state-of-the-art in commonly evaluated benchmarks, and the superior improvement 12% in IMIM indicates our method efficiently mitigates the misalignment.

源语言英语
主期刊名Proceedings - 2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
出版商Institute of Electrical and Electronics Engineers Inc.
20313-20323
页数11
ISBN(电子版)9798331587758
DOI
出版状态已出版 - 2025
活动2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025 - Honolulu, 美国
期限: 19 10月 202523 10月 2025

出版系列

姓名Proceedings of the IEEE International Conference on Computer Vision
ISSN(印刷版)1550-5499
ISSN(电子版)2380-7504

会议

会议2025 IEEE/CVF International Conference on Computer Vision, ICCV 2025
国家/地区美国
Honolulu
时期19/10/2523/10/25

学术指纹

探究 'Mind the Gap: Aligning Vision Foundation Models to Image Feature Matching' 的科研主题。它们共同构成独一无二的学术指纹。

引用此