TY - JOUR
T1 - Plug-and-Adapt
T2 - Multimodal Coreference Resolution at First Sight with a Pretrained Alignment Model
AU - Wu, Jinghan
AU - Li, Jing
AU - Tsang, Ivor W.
AU - Zhang, Xuetao
N1 - Publisher Copyright:
© 1999-2012 IEEE.
PY - 2026
Y1 - 2026
N2 - Multi-modal Coreference Resolution (MCR) methodsrequires training with (partially) annotated data from the targetdataset before they can be applied, preventing their direct usability and raising concerns about generalization. While Vision Language Large Models (VLLMs) with billions of parameters offerpromising zero-shot capabilities, they remain largely inaccessible.Their massive size limits deployability, and many are only accessible through paid APIs. In this paper, we propose a plug-andadapt method that strategically adapts a carefully pre-trainedalignment model for immediate use in MCR tasks, designed toeliminate the need for training on scarce benchmark datasets orrelying on resource-intensive VLLMs. Specifically, we first pretrain a fine-grained alignment model between textual and visualcontextual information using vision-language alignment datasets.We then repurpose the alignment model to MCR throughsimilarity aggregation by fusing visual and categorical cueswith evidence theory, enhancing the effectiveness. Experimentson the Coreference Image Narratives (CIN) benchmark datasetdemonstrate the effectiveness of our method, achieving a 5.31%and 2.12% improvement in CoNLL F1 over SOTA dedicatedmethods and popular VLLMs, respectively. We further evaluateour method on a masked CIN dataset for robustness testing andon a specially constructed VCR-MCR dataset for generalizationassessment, with results confirming both capabilities.
AB - Multi-modal Coreference Resolution (MCR) methodsrequires training with (partially) annotated data from the targetdataset before they can be applied, preventing their direct usability and raising concerns about generalization. While Vision Language Large Models (VLLMs) with billions of parameters offerpromising zero-shot capabilities, they remain largely inaccessible.Their massive size limits deployability, and many are only accessible through paid APIs. In this paper, we propose a plug-andadapt method that strategically adapts a carefully pre-trainedalignment model for immediate use in MCR tasks, designed toeliminate the need for training on scarce benchmark datasets orrelying on resource-intensive VLLMs. Specifically, we first pretrain a fine-grained alignment model between textual and visualcontextual information using vision-language alignment datasets.We then repurpose the alignment model to MCR throughsimilarity aggregation by fusing visual and categorical cueswith evidence theory, enhancing the effectiveness. Experimentson the Coreference Image Narratives (CIN) benchmark datasetdemonstrate the effectiveness of our method, achieving a 5.31%and 2.12% improvement in CoNLL F1 over SOTA dedicatedmethods and popular VLLMs, respectively. We further evaluateour method on a masked CIN dataset for robustness testing andon a specially constructed VCR-MCR dataset for generalizationassessment, with results confirming both capabilities.
KW - coreference resolution
KW - model adaptation
KW - multi-modal learning
UR - https://www.scopus.com/pages/publications/105046827584
U2 - 10.1109/TMM.2026.3719817
DO - 10.1109/TMM.2026.3719817
M3 - 文章
AN - SCOPUS:105046827584
SN - 1520-9210
JO - IEEE Transactions on Multimedia
JF - IEEE Transactions on Multimedia
ER -