跳到主要导航 跳到搜索 跳到主要内容

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model

  • Xi'an Jiaotong University
  • Agency for Science, Technology and Research, Singapore

科研成果: 期刊稿件文章同行评审

摘要

Multi-modal Coreference Resolution (MCR) methodsrequires training with (partially) annotated data from the targetdataset before they can be applied, preventing their direct usability and raising concerns about generalization. While Vision Language Large Models (VLLMs) with billions of parameters offerpromising zero-shot capabilities, they remain largely inaccessible.Their massive size limits deployability, and many are only accessible through paid APIs. In this paper, we propose a plug-andadapt method that strategically adapts a carefully pre-trainedalignment model for immediate use in MCR tasks, designed toeliminate the need for training on scarce benchmark datasets orrelying on resource-intensive VLLMs. Specifically, we first pretrain a fine-grained alignment model between textual and visualcontextual information using vision-language alignment datasets.We then repurpose the alignment model to MCR throughsimilarity aggregation by fusing visual and categorical cueswith evidence theory, enhancing the effectiveness. Experimentson the Coreference Image Narratives (CIN) benchmark datasetdemonstrate the effectiveness of our method, achieving a 5.31%and 2.12% improvement in CoNLL F1 over SOTA dedicatedmethods and popular VLLMs, respectively. We further evaluateour method on a masked CIN dataset for robustness testing andon a specially constructed VCR-MCR dataset for generalizationassessment, with results confirming both capabilities.

源语言英语
期刊IEEE Transactions on Multimedia
DOI
出版状态已接受/待刊 - 2026

学术指纹

探究 'Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model' 的科研主题。它们共同构成独一无二的学术指纹。

引用此