Abstract
Objectives: To investigate the proficiency of artificial intelligence (ChatGPT-4 and DeepSeek-v3) in responding to frequently asked questions (FAQs) related to oral lichen planus (OLP).
Study Design: Twenty-three OLP-related FAQs were adapted from reputable sources. Responses generated by ChatGPT-4 and DeepSeek-v3 were assessed by expert panels across five domains: accuracy, comprehensibility, logic, fairness, and simplicity.
Results: DeepSeek-v3 outperformed ChatGPT-4 in comprehensiveness (4.65 ± 0.37 vs. 4.04 ± 0.54; P < 0.001), logic (4.41 ± 0.32 vs. 4.03 ± 0.35; P < 0.001), and fairness (4.34 ± 0.33 vs. 4.03 ± 0.27; P < 0.001). ChatGPT-4 scored higher in simplicity (4.25 ± 0.31 vs. 4.04 ± 0.28; P = 0.034), while DeepSeek-v3 had a modest advantage in accuracy (4.49 ± 0.38 vs. 4.22 ± 0.34; P = 0.025). DeepSeek-v3 performed better in comprehensiveness and logic for "Causes" and "Symptoms," and in accuracy for "Risk & Complications." ChatGPT-4’s simplicity was most evident in "Causes" and "Symptoms." In "Treatment & Management," DeepSeek-v3 excelled in logic and comprehensiveness, while ChatGPT-4 showed greater simplicity and occasional accuracy.
Conclusion: ChatGPT-4 provides concise, accessible information, whereas DeepSeek-v3 offers broader, more detailed responses. Combining both models may improve patient understanding of OLP.