TY - JOUR
T1 - Responsible AI in healthcare
T2 - Mitigating hallucinations and enhancing multimodal fusion - based reasoning in medical imaging
AU - Iqbal, Saeed
AU - Zhong, Xiaopin
AU - Khan, Muhammad Attique
AU - Wu, Zongze
AU - Almujally, Nouf Abdullah
AU - Liu, Weixiang
AU - Cambria, Erik
AU - Hussain, Amir
N1 - Publisher Copyright:
Copyright © 2026. Published by Elsevier B.V.
PY - 2026/12
Y1 - 2026/12
N2 - Multi-modal healthcare data challenges traditional machine learning models, causing misaligned features, overconfident predictions, and errors. Privacy concerns and dataset heterogeneity further limit AI scalability in healthcare. We propose a framework integrating multi-modal reasoning, hallucination mitigation, and uncertainty quantification. Our method uses Bayesian uncertainty quantification, cross-modal contrastive learning, and retrieval-augmented reasoning (RAR) to decrease hallucinations in large vision-language models (VLMs). We employ deep ensembles, Dirichlet-based calibration, and temperature scaling to improve the dependability of the model. Our system uses attention-driven fusion, dynamic modality weighting, and hierarchical interaction modeling to merge several modalities into a single representation through the Unified Multimodal Reasoning (UMR) Platform. It facilitates human-in-the-loop validation and temporal contextualization, which helps with adaptive learning and sound decision-making in dynamic clinical situations. We use rigorous assessment measures to show the efficacy of our approach. Innovative measuring criteria including the Clinical Consistency Score (CCS), Hallucination Rate Reduction (HRR), and Cross-Modal Alignment Index (CMAI) are employed. These metrics offer detailed information on how well the model performs in terms of clinical consistency, hallucination mitigation, and multi-modal alignment. Our framework exhibits notable enhancements when compared to state-of-the-art (SOTA) models. Our models (LLaVA-Med + RAR, LLaVA-Med + CMCL, and LLaVA-Med + UMR) outperform the best-performing SOTA model (LLaVA-Med) by an average of 0.05 to 0.1 points across datasets in terms of BertScore. In terms of BLEU, our approach generates medical reports more accurately than SOTA models, outperforming them by 3 to 5 in several tasks. We find gains of 0.03 to 0.05 in METEOR and 0.02 to 0.05 in ROUGE scores, respectively, indicating improved recall and similarity in text-based assessments. In comparison to SOTA models, our models exhibit a 10 - 15% improvement in cross-modal alignment, according to the CMAI, and a 15 - 20% decrease in hallucination rates, according to the HRR. With a 0.03 - 0.07 rise in CCS values, the CCS further confirms that our models are more in line with clinical knowledge. The framework addresses model overconfidence, multi-modal integration, and hallucination reduction, advancing reliable AI-driven medical solutions. Code and data are available at: MLVLM .
AB - Multi-modal healthcare data challenges traditional machine learning models, causing misaligned features, overconfident predictions, and errors. Privacy concerns and dataset heterogeneity further limit AI scalability in healthcare. We propose a framework integrating multi-modal reasoning, hallucination mitigation, and uncertainty quantification. Our method uses Bayesian uncertainty quantification, cross-modal contrastive learning, and retrieval-augmented reasoning (RAR) to decrease hallucinations in large vision-language models (VLMs). We employ deep ensembles, Dirichlet-based calibration, and temperature scaling to improve the dependability of the model. Our system uses attention-driven fusion, dynamic modality weighting, and hierarchical interaction modeling to merge several modalities into a single representation through the Unified Multimodal Reasoning (UMR) Platform. It facilitates human-in-the-loop validation and temporal contextualization, which helps with adaptive learning and sound decision-making in dynamic clinical situations. We use rigorous assessment measures to show the efficacy of our approach. Innovative measuring criteria including the Clinical Consistency Score (CCS), Hallucination Rate Reduction (HRR), and Cross-Modal Alignment Index (CMAI) are employed. These metrics offer detailed information on how well the model performs in terms of clinical consistency, hallucination mitigation, and multi-modal alignment. Our framework exhibits notable enhancements when compared to state-of-the-art (SOTA) models. Our models (LLaVA-Med + RAR, LLaVA-Med + CMCL, and LLaVA-Med + UMR) outperform the best-performing SOTA model (LLaVA-Med) by an average of 0.05 to 0.1 points across datasets in terms of BertScore. In terms of BLEU, our approach generates medical reports more accurately than SOTA models, outperforming them by 3 to 5 in several tasks. We find gains of 0.03 to 0.05 in METEOR and 0.02 to 0.05 in ROUGE scores, respectively, indicating improved recall and similarity in text-based assessments. In comparison to SOTA models, our models exhibit a 10 - 15% improvement in cross-modal alignment, according to the CMAI, and a 15 - 20% decrease in hallucination rates, according to the HRR. With a 0.03 - 0.07 rise in CCS values, the CCS further confirms that our models are more in line with clinical knowledge. The framework addresses model overconfidence, multi-modal integration, and hallucination reduction, advancing reliable AI-driven medical solutions. Code and data are available at: MLVLM .
KW - Clinical consistency
KW - Cross - modal alignment
KW - Hallucination mitigation
KW - Medical large vision - language models
KW - Multi - modal data fusion
KW - Uncertainty quantification
UR - https://www.scopus.com/pages/publications/105040724819
U2 - 10.1016/j.inffus.2026.104483
DO - 10.1016/j.inffus.2026.104483
M3 - 文章
AN - SCOPUS:105040724819
SN - 1566-2535
VL - 136
JO - Information Fusion
JF - Information Fusion
M1 - 104483
ER -