跳到主要导航 跳到搜索 跳到主要内容

Evaluating Generative AI Large Language Models for Urticaria Management: A Comparative Analysis of DeepSeek-R1 and ChatGPT-4o

  • Mengyao Yang
  • , Jingchen Liang
  • , Luyue Zhang
  • , Hongshan Liu
  • , Ying Chen
  • , Yawen Wang
  • , Chunzhi Qi
  • , Yuxin Ma
  • , Ziyun Gao
  • , Xinyue Zhang
  • , Xinwu Niu
  • , Xiaopeng Wang
  • , Jianwen Ren
  • , Jingyi Yuan
  • , Weihui Zeng
  • , Zhao Wang
  • The Second Affiliated Hospital of Xi'an Jiaotong University

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

Introduction: Urticaria is a prevalent condition affecting a significant portion of the global population. Both dermatologists and patients require access to up-to-date and accurate information. Traditional search engines often fall short in meeting these needs. Despite the growing reliance on AI for medical inquiries, the accuracy and quality of AI-generated remain understudied. This study aims to evaluate and compare the performance of two widely used AI models, ChatGPT-4o and DeepSeek-R1, in addressing urticaria-related queries. Methods: An e-Delphi procedure was employed to generate and refine a set of urticaria-related questions, as well as to develop an evaluation framework for AI-generated responses. ChatGPT-4o and DeepSeek-R1 were then prompted with the finalized questions, and their responses were recorded. A single-blind comparative assessment was conducted among 67 participants (29 dermatologists and 38 non-dermatologists). The responses from both AI models were assessed across simplicity, accuracy, professionalism, clinical feasibility, comprehensibility, and completeness. Results: DeepSeek-R1 outperformed ChatGPT-4o in most metrics. Dermatologists rated DeepSeek significantly higher in simplicity (p < 0.001), accuracy (p < 0.001), completeness (p = 0.001), professionalism (p < 0.001), and clinical feasibility (p < 0.001). Non-dermatologists found DeepSeek's responses more concise (p < 0.001) and comprehensible (p < 0.001). Both models showed comparable integration of cutting-edge knowledge (p = 0.06), though DeepSeek exhibited greater output stability, as evidenced by lower standard deviations. When compared with the guidelines, the answers provided by DeepSeek-R1 contained no errors, while ChatGPT-4o made errors in three clinical questions. Conclusion: AI-generated answers require rigorous evaluation to ensure their reliability and suitability for medical applications. Based on the current study, DeepSeek-R1 outperforms ChatGPT-4o in addressing urticaria-related queries, demonstrating higher potential for both clinical and patient use.

源语言英语
文章编号e70113
期刊Clinical and Translational Allergy
15
11
DOI
出版状态已出版 - 11月 2025
已对外发布

学术指纹

探究 'Evaluating Generative AI Large Language Models for Urticaria Management: A Comparative Analysis of DeepSeek-R1 and ChatGPT-4o' 的科研主题。它们共同构成独一无二的指纹。

引用此