TY - JOUR
T1 - Evaluating Generative AI Large Language Models for Urticaria Management
T2 - A Comparative Analysis of DeepSeek-R1 and ChatGPT-4o
AU - Yang, Mengyao
AU - Liang, Jingchen
AU - Zhang, Luyue
AU - Liu, Hongshan
AU - Chen, Ying
AU - Wang, Yawen
AU - Qi, Chunzhi
AU - Ma, Yuxin
AU - Gao, Ziyun
AU - Zhang, Xinyue
AU - Niu, Xinwu
AU - Wang, Xiaopeng
AU - Ren, Jianwen
AU - Yuan, Jingyi
AU - Zeng, Weihui
AU - Wang, Zhao
N1 - Publisher Copyright:
© 2025 The Author(s). Clinical and Translational Allergy published by John Wiley & Sons Ltd on behalf of European Academy of Allergy and Clinical Immunology.
PY - 2025/11
Y1 - 2025/11
N2 - Introduction: Urticaria is a prevalent condition affecting a significant portion of the global population. Both dermatologists and patients require access to up-to-date and accurate information. Traditional search engines often fall short in meeting these needs. Despite the growing reliance on AI for medical inquiries, the accuracy and quality of AI-generated remain understudied. This study aims to evaluate and compare the performance of two widely used AI models, ChatGPT-4o and DeepSeek-R1, in addressing urticaria-related queries. Methods: An e-Delphi procedure was employed to generate and refine a set of urticaria-related questions, as well as to develop an evaluation framework for AI-generated responses. ChatGPT-4o and DeepSeek-R1 were then prompted with the finalized questions, and their responses were recorded. A single-blind comparative assessment was conducted among 67 participants (29 dermatologists and 38 non-dermatologists). The responses from both AI models were assessed across simplicity, accuracy, professionalism, clinical feasibility, comprehensibility, and completeness. Results: DeepSeek-R1 outperformed ChatGPT-4o in most metrics. Dermatologists rated DeepSeek significantly higher in simplicity (p < 0.001), accuracy (p < 0.001), completeness (p = 0.001), professionalism (p < 0.001), and clinical feasibility (p < 0.001). Non-dermatologists found DeepSeek's responses more concise (p < 0.001) and comprehensible (p < 0.001). Both models showed comparable integration of cutting-edge knowledge (p = 0.06), though DeepSeek exhibited greater output stability, as evidenced by lower standard deviations. When compared with the guidelines, the answers provided by DeepSeek-R1 contained no errors, while ChatGPT-4o made errors in three clinical questions. Conclusion: AI-generated answers require rigorous evaluation to ensure their reliability and suitability for medical applications. Based on the current study, DeepSeek-R1 outperforms ChatGPT-4o in addressing urticaria-related queries, demonstrating higher potential for both clinical and patient use.
AB - Introduction: Urticaria is a prevalent condition affecting a significant portion of the global population. Both dermatologists and patients require access to up-to-date and accurate information. Traditional search engines often fall short in meeting these needs. Despite the growing reliance on AI for medical inquiries, the accuracy and quality of AI-generated remain understudied. This study aims to evaluate and compare the performance of two widely used AI models, ChatGPT-4o and DeepSeek-R1, in addressing urticaria-related queries. Methods: An e-Delphi procedure was employed to generate and refine a set of urticaria-related questions, as well as to develop an evaluation framework for AI-generated responses. ChatGPT-4o and DeepSeek-R1 were then prompted with the finalized questions, and their responses were recorded. A single-blind comparative assessment was conducted among 67 participants (29 dermatologists and 38 non-dermatologists). The responses from both AI models were assessed across simplicity, accuracy, professionalism, clinical feasibility, comprehensibility, and completeness. Results: DeepSeek-R1 outperformed ChatGPT-4o in most metrics. Dermatologists rated DeepSeek significantly higher in simplicity (p < 0.001), accuracy (p < 0.001), completeness (p = 0.001), professionalism (p < 0.001), and clinical feasibility (p < 0.001). Non-dermatologists found DeepSeek's responses more concise (p < 0.001) and comprehensible (p < 0.001). Both models showed comparable integration of cutting-edge knowledge (p = 0.06), though DeepSeek exhibited greater output stability, as evidenced by lower standard deviations. When compared with the guidelines, the answers provided by DeepSeek-R1 contained no errors, while ChatGPT-4o made errors in three clinical questions. Conclusion: AI-generated answers require rigorous evaluation to ensure their reliability and suitability for medical applications. Based on the current study, DeepSeek-R1 outperforms ChatGPT-4o in addressing urticaria-related queries, demonstrating higher potential for both clinical and patient use.
KW - AI
KW - ChatGPT-4o
KW - DeepSeek-R1
KW - large language model
KW - urticaria
UR - https://www.scopus.com/pages/publications/105023289110
U2 - 10.1002/clt2.70113
DO - 10.1002/clt2.70113
M3 - 文章
AN - SCOPUS:105023289110
SN - 2045-7022
VL - 15
JO - Clinical and Translational Allergy
JF - Clinical and Translational Allergy
IS - 11
M1 - e70113
ER -