Navigating Accuracy and Clarity in Patient Education: The Role of AI Chatbots in Communicating Testicular Prosthesis Information


COŞER Ş., ARAS B., Ozercan A. Y., Ozkaynak B. N., Alkis O., Basboga S., ...More

Archivos Espanoles de Urologia, vol.79, no.5, pp.830-836, 2026 (SCI-Expanded, Scopus)

  • Publication Type: Article / Article
  • Volume: 79 Issue: 5
  • Publication Date: 2026
  • Doi Number: 10.56434/j.arch.esp.urol.20267905.97
  • Journal Name: Archivos Espanoles de Urologia
  • Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, BIOSIS, DIALNET, Biomedical Reference Collection: Corporate Edition (EBSCO)
  • Page Numbers: pp.830-836
  • Keywords: artificial intelligence, chatbot, patient education, quality assessment, testicular prosthesis
  • Kütahya Health Sciences University Affiliated: Yes

Abstract

Background: This study investigated the quality and comprehensibility of responses generated by four different artificial intelligence (AI)-powered chatbots (ChatGPT, Gemini, DeepSeek, and Grok) when queried about testicular prostheses. Methods: A Google search using the keyword “testicular prosthesis” was conducted, and the 50 most frequently asked questions listed in the “People Also Ask” section were identified. These questions were categorized into preoperative, perioperative, and postoperative topics and were posed to four AI chatbots: ChatGPT, Google Gemini, DeepSeek, and Grok. The responses were independently evaluated by four urologists using the Global Quality Scale (GQS), Modified DISCERN, and Patient Education Materials Assessment Tool for Printed Materials (PEMAT-P) scales to assess quality, reliability, and readability. Results: According to the GQS evaluation, the median scores for ChatGPT, Gemini, DeepSeek, and Grok were 4.75, 4.5, 4.75 and 4.875, respectively (p < 0.001). According to the Modified DISCERN scale, the scores were 2.75, 3.0, 3.75 and 3.0, respectively. The PEMAT-P understandability scores were 79.1%, 75.0%, 84.2% and 78.6%, respectively (p < 0.001). Similarly, the PEMAT-P actionability scores were 73.5%, 68.5%, 78.9% and 73.2%, respectively (p < 0.001). Conclusions: While all chatbots provided high-quality responses, their reliability was moderate. DeepSeek demonstrated the best performance across the evaluated metrics. These findings suggest that AI chatbots may serve as useful supplementary tools for patient education regarding testicular prostheses.