Logo image
A real‐world analysis of AI chatbot performance for medicines information enquiries
Journal article   Open access   Peer reviewed

A real‐world analysis of AI chatbot performance for medicines information enquiries

Duncan Yorkston, Tracey Borrie and Paul Chin
Journal of pharmacy practice and research
03/08/2026
Handle:
https://hdl.handle.net/10523/52076

Abstract

artificial intelligence drug information medicines information evidence-based medicine clinical decision support clinical pharmacy
Background: The provision of medicines information (MI) services requires interpretation and clinical judgement of complex scenarios by pharmacists. To date, few studies have assessed the performance of artificial intelligence (AI) chatbots to assist pharmacists providing MI advice. Aim: To evaluate the performance and risk associated with two AI chatbots (Microsoft Copilot and Google Gemini) to answer medicines-related questions. Method: A sample of 20 questions answered by the local MI service in November 2023 was entered in the two chatbot applications in January 2024 (round 1) and May 2024 (round 2). All questions were preceded with the prompt 'I'm a pharmacist'. Chatbot responses were evaluated by comparing with a reference answer given by the MI service using a consensus process in the domains of content, patient management, risk of patient harm, and follow up review. Ethical approval was granted by the Canterbury District Health Board Research Office (Reference no: 20311) and the study conforms with the Declaration of Helsinki. Results: For the 20 questions answered by both chatbots, few of the round 1 responses (n = 4 for Copilot and n = 2 for Gemini) were considered complete and with adequate information to commence patient management with no risk of harm. Most were incomplete (n = 13 for Copilot and n = 15 for Gemini) regarding content, but none were high risk of causing harm. In round 1, four responses from Copilot and eight from Gemini were flagged for follow up review. There was no significant difference in performance between chatbots in round 1 (p = 0.68) or between rounds 1 and 2 (Copilot p = 0.25 and Gemini p > 0.99). Conclusion: Our study results demonstrated the chatbots' responses were typically suboptimal; albeit, a significant minority prompted a follow up to review the chatbot response.
pdf
Pharmacy Practice and Res - 2026 - Yorkston - A real‐world analysis of AI chatbot performance for medicines information499.14 kBDownloadView
Published (Version of record) Open Access CC BY-NC-ND V4.0
url
https://doi.org/10.1002/jppr.70090View
Published (Version of record) Open CC BY-NC-ND V4.0

Metrics

1 Record Views

Details

Logo image