Abstract
Multi-modal healthcare data challenges traditional machine learning models, causing misaligned features, overconfident predictions, and errors. Privacy concerns and dataset heterogeneity further limit AI scalability in healthcare. We propose a framework integrating multi-modal reasoning, hallucination mitigation, and uncertainty quantification. Our method uses Bayesian uncertainty quantification, cross-modal contrastive learning, and retrieval-augmented reasoning (RAR) to decrease hallucinations in large vision-language models (VLMs). We employ deep ensembles, Dirichlet-based calibration, and temperature scaling to improve the dependability of the model. Our system uses attention-driven fusion, dynamic modality weighting, and hierarchical interaction modeling to merge several modalities into a single representation through the Unified Multimodal Reasoning (UMR) Platform. It facilitates human-in-the-loop validation and temporal contextualization, which helps with adaptive learning and sound decision-making in dynamic clinical situations. We use rigorous assessment measures to show the efficacy of our approach. Innovative measuring criteria including the Clinical Consistency Score (CCS), Hallucination Rate Reduction (HRR), and Cross-Modal Alignment Index (CMAI) are employed. These metrics offer detailed information on how well the model performs in terms of clinical consistency, hallucination mitigation, and multi-modal alignment. Our framework exhibits notable enhancements when compared to state-of-the-art (SOTA) models. Our models (LLaVA-Med + RAR, LLaVA-Med + CMCL, and LLaVA-Med + UMR) outperform the best-performing SOTA model (LLaVA-Med) by an average of 0.05 to 0.1 points across datasets in terms of BertScore. In terms of BLEU, our approach generates medical reports more accurately than SOTA models, outperforming them by 3 to 5 in several tasks. We find gains of 0.03 to 0.05 in METEOR and 0.02 to 0.05 in ROUGE scores, respectively, indicating improved recall and similarity in text-based assessments. In comparison to SOTA models, our models exhibit a 10 - 15% improvement in cross-modal alignment, according to the CMAI, and a 15 - 20% decrease in hallucination rates, according to the HRR. With a 0.03 - 0.07 rise in CCS values, the CCS further confirms that our models are more in line with clinical knowledge. The framework addresses model overconfidence, multi-modal integration, and hallucination reduction, advancing reliable AI-driven medical solutions. Code and data are available at: MLVLM .
| Original language | English |
|---|---|
| Article number | 104483 |
| Journal | Information Fusion |
| Volume | 136 |
| DOIs | |
| State | Published - Dec 2026 |
| Externally published | Yes |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- Clinical consistency
- Cross - modal alignment
- Hallucination mitigation
- Medical large vision - language models
- Multi - modal data fusion
- Uncertainty quantification
Fingerprint
Dive into the research topics of 'Responsible AI in healthcare: Mitigating hallucinations and enhancing multimodal fusion - based reasoning in medical imaging'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver