Subgroup Differences in Agreement Between an Algorithm Guided Large Language Model and Routine Emergency Department Triage


HALICI A., Cesur E., Çelik F.

Medical science monitor : international medical journal of experimental and clinical research, cilt.32, 2026 (SCI-Expanded, Scopus)

Özet

BACKGROUND Large language models (LLMs) are increasingly discussed as decision-support tools in emergency care, but their agreement with routine emergency department (ED) triage and subgroup behavior remain insufficiently characterized. We evaluated an algorithm-guided LLM against routine ED triage with emphasis on subgroup heterogeneity and safety-relevant discordance. MATERIAL AND METHODS This retrospective study included 1960 adult ED visits with complete triage data. A standardized prompt provided age, sex, chief complaint, comorbidities, systolic/diastolic blood pressure, heart rate, oxygen saturation, temperature, and Glasgow Coma Scale. The LLM assigned 1 triage category within a 5-level Emergency Severity Index-based system (green, yellow-1, yellow-2, red-1, red-2). Outputs were compared with routine ED triage. Performance for urgent vs non-urgent classification was assessed using AUC, sensitivity, specificity, positive predictive value, negative predictive value, F1, and accuracy. Five-level agreement was assessed using quadratic weighted Cohen's kappa and accuracy. Discordance (lower- and higher-acuity LLM vs routine triage) was analyzed across prespecified subgroups. RESULTS LLM achieved 71.3% five-level accuracy and substantial agreement with routine triage (weighted kappa=0.824). For urgent/non-urgent classification, AUC was 0.768, sensitivity 0.630, specificity 0.906. Lower- and higher-acuity discordance rates were 9.8% and 18.9%. Discordance varied across subgroups; lower-acuity assignments vs routine triage were more frequent in older adults, trauma, and diabetes, while infectious presentations showed the highest concordance. CONCLUSIONS The algorithm-guided LLM showed substantial concordance with routine ED triage but non-uniform subgroup discordance, particularly lower-acuity assignments in patients with older age, diabetes, and trauma. As routine triage served as an operational comparator rather than a gold standard, findings reflect agreement with local practice, not definitive accuracy or safety. Prospective outcome validation is required.