Evaluation of large language models in a national orthopaedic proficiency examination: Implications for health informatics and medical education


ARI B.

Health Informatics Journal, cilt.32, sa.3, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 32 Sayı: 3
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1177/14604582261470637
  • Dergi Adı: Health Informatics Journal
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Applied Science & Technology Source, CINAHL, EBSCO Education Source, Educational research abstracts (ERA), EMBASE, INSPEC, Library, Information Science & Technology Abstracts (LISTA), MEDLINE, Directory of Open Access Journals, Information Science & Technology Abstracts (LISTA), Academic Search Ultimate (EBSCO), Education Source Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: artificial intelligence, continuing healthcare education, deep learning, digital health
  • Kütahya Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Objective: This study evaluates the performance of large language models (LLMs)—ChatGPT-4.0, Gemini 2.0 Pro, o3-mini, Doctor GPT and DeepSeek-V3—in a national orthopaedic proficiency examination and explores their implications for health informatics and medical education. The responses of these models were analysed to assess accuracy rates and differences between models. Method: A total of 100 multiple-choice questions from the 2024 TOTEK examination were administered to each AI model under identical conditions. Correct and incorrect responses were recorded, and differences in performance were evaluated using chi-square testing and frequency analysis. Question categories were also compared to identify domain-specific variations. Results: o3-mini achieved the highest accuracy rate (79%), while Gemini 2.0 showed the lowest (68%); all models exceeded the 60% pass threshold. A statistically significant difference between models was identified in the Surgical Procedures category, in which Gemini 2.0 answered fewer questions correctly (23/36) than the other models (30–32/36) (χ2 = 9.87, df = 4, p = 0.043). No significant differences were observed in the remaining categories (all p > 0.05), and the overall difference in accuracy between models did not reach statistical significance (χ2 = 4.01, df = 4, p = 0.405). Clinical decision-making and visual content-based questions were the most challenging for all models. Conclusion: AI models demonstrate generally high accuracy in medical examinations; however, they struggle with interpreting clinical context, recognising atypical medical scenarios and answering questions involving visual content.