Skip to main navigation Skip to search Skip to main content

Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model

  • Xi'an Jiaotong University
  • Agency for Science, Technology and Research, Singapore

Research output: Contribution to journalArticlepeer-review

Abstract

Multi-modal Coreference Resolution (MCR) methodsrequires training with (partially) annotated data from the targetdataset before they can be applied, preventing their direct usability and raising concerns about generalization. While Vision Language Large Models (VLLMs) with billions of parameters offerpromising zero-shot capabilities, they remain largely inaccessible.Their massive size limits deployability, and many are only accessible through paid APIs. In this paper, we propose a plug-andadapt method that strategically adapts a carefully pre-trainedalignment model for immediate use in MCR tasks, designed toeliminate the need for training on scarce benchmark datasets orrelying on resource-intensive VLLMs. Specifically, we first pretrain a fine-grained alignment model between textual and visualcontextual information using vision-language alignment datasets.We then repurpose the alignment model to MCR throughsimilarity aggregation by fusing visual and categorical cueswith evidence theory, enhancing the effectiveness. Experimentson the Coreference Image Narratives (CIN) benchmark datasetdemonstrate the effectiveness of our method, achieving a 5.31%and 2.12% improvement in CoNLL F1 over SOTA dedicatedmethods and popular VLLMs, respectively. We further evaluateour method on a masked CIN dataset for robustness testing andon a specially constructed VCR-MCR dataset for generalizationassessment, with results confirming both capabilities.

Original languageEnglish
JournalIEEE Transactions on Multimedia
DOIs
StateAccepted/In press - 2026

Keywords

  • coreference resolution
  • model adaptation
  • multi-modal learning

Fingerprint

Dive into the research topics of 'Plug-and-Adapt: Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model'. Together they form a unique fingerprint.

Cite this