Evaluating ChatGPT-5.0 in developmental dysplasia of the hip: A scenario-based validation study against AAOS and AAP guidelines


KOZLU S., Görgün B., ÖNER S. K.

Journal of Children's Orthopaedics, cilt.20, sa.2, ss.126-133, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 20 Sayı: 2
  • Basım Tarihi: 2026
  • Doi Numarası: 10.1177/18632521261419320
  • Dergi Adı: Journal of Children's Orthopaedics
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Directory of Open Access Journals, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
  • Sayfa Sayıları: ss.126-133
  • Anahtar Kelimeler: artificial intelligence (AI), ChatGPT-5.0, Developmental dysplasia of the hip (DDH)
  • Kütahya Sağlık Bilimleri Üniversitesi Adresli: Evet

Özet

Purpose: Developmental dysplasia of the hip (DDH) requires timely, guideline-concordant decisions to prevent long-term morbidity. ChatGPT-5.0 may support clinicians—especially where pediatric orthopedic expertise is limited, but their reliability across typical and discordant presentations is uncertain. This scenario-based validation study evaluated the accuracy of ChatGPT-5.0’s management recommendations for DDH using 30 structured clinical cases and compared these outputs against AAOS (2022) and AAP (2016) guidelines. Methods: Scenario-based validation using 30 unique cases: 20 concordant (aligned clinical and imaging findings) spanning Graf and acetabular index-based ages, and 10 mismatch scenarios with correct examinations but intentionally erroneous radiology. The primary outcome was guideline-concordant accuracy, categorized as correct, partially correct, undertreatment, overtreatment, or incorrect. Secondary outcomes included the effect of error-aware prompts and multilingual consistency. Results: In concordant scenarios, guided ChatGPT achieved 100% correct, while non-logged-in ChatGPT achieved 95% with one overtreatment. In mismatch scenarios, guided ChatGPT frequently tends toward overtreatment and failing to recommend repeat ultrasound or urgent pediatric orthopedic consultation. Non-logged-in ChatGPT performed better in mismatch cases but similarly under-emphasized remeasurement/consultation. Error-aware prompts did not materially alter recommendations in either environment. Swahili queries produced outputs clinically identical to English responses. Conclusions: ChatGPT-5.0 provides reliable, guideline-concordant guidance for DDH when clinical and radiologic data are concordant, supporting potential use as a decision aid in settings without immediate pediatric orthopedic access. Safe clinical implementation requires human oversight and integration of guideline-based safety checks to prevent mismanagement in ambiguous cases.