Diagnostic Accuracy of Two Large Language Models in Turkish Emergency Department Anamnesis Notes
NCT07632859 · Status: COMPLETED · Type: OBSERVATIONAL · Enrollment: 600
Last updated 2026-08-14
Summary
This retrospective diagnostic accuracy study evaluates two large language models - GPT-4.1 (gpt-4.1-2025-04-14; OpenAI) and Claude Sonnet 4.6 (claude-sonnet-4-6; Anthropic) - as retrospective coding-quality instruments applied to anonymized Turkish-language emergency department anamnesis notes.
The reference standard is the majority consensus of three board-certified emergency medicine specialists who independently coded each note in ICD-10, blinded to one another, to the code entered by the treating physician at case closure, and to the subsequent clinical course. Cases without chapter-level majority agreement are excluded without replacement.
Both models are queried once per note with a single locked prompt at temperature 0 in stateless application programming interface calls, with no retrieval augmentation, no external tools and no extended-reasoning mode. The primary outcome is the proportion of cases in which each model's rank-1 diagnosis matches the reference standard at ICD-10 chapter level, reported with a Wilson 95% confidence interval. Registered secondary outcome measures are chapter-level Cohen's kappa between each model's rank-1 diagnosis and the reference standard; top-3 chapter accuracy for each model; and chapter-level concordance between the closure ICD-10 code and the reference standard. Additional prespecified analyses set out in the statistical analysis plan (paired between-model difference, three-character accuracy, note-length association, confidence calibration and model-to-model agreement) are reported in the primary publication.
The ICD-10 code entered at case closure is characterised against the same reference standard as a description of current documentation practice; it is not a comparator, and no test of superiority or inferiority against model output is performed. The analysis plan was finalised and frozen before any accuracy computation. Reporting follows STARD-AI 2025.
Conditions
- Emergency Medicine
- Diagnostic Errors
- Artificial Intelligence (AI) in Diagnosis
Sponsors & Collaborators
-
Marmara University Pendik Training and Research Hospital
lead OTHER
Principal Investigators
-
Emir Ünal · Marmara University
Eligibility
- Min Age
- 18 Years
- Sex
- ALL
- Healthy Volunteers
- No
Timeline & Regulatory
- Start
- 2026-05-01
- Primary Completion
- 2026-08-03
- Completion
- 2026-08-07
Countries
- Turkey (Türkiye)
Study Locations
More Related Trials
-
Assessing the Accuracy of ChatGPT-4 in Interpreting Arterial Blood Gas Results
NCT06456866 ·Status: COMPLETED
-
The Predictability of the Necessity for Cardiology Consultation in Patients Scheduled for Non-Cardiac Surgery Using Artificial Intelligence Models in Preoperative Anesthesia Assessment
NCT07395713 ·Status: ACTIVE_NOT_RECRUITING
-
Comparison of AI-Generated Pain Scoring Visuals With Visual Analog Scale (VAS) for Pain Assessment
NCT06456853 ·Status: COMPLETED
-
The Success of ChatGPT in Providing American Society of Anesthesiologist (ASA) Scores
NCT06321445 ·Status: COMPLETED
-
Evaluation of the Success of Artificial Intelligence Models in Interpreting Arterial Waveform Analysis Data
NCT06828575 ·Status: RECRUITING
-
ChatGPT-5 vs. CDSS for Drug-Drug Interactions in ICU
NCT07314125 ·Status: COMPLETED
-
Capabilities ofArtificial Intelligence Models in Externation Decision of Patient Who Followed in Intensive Care Unit ()
NCT06584890 ·Status: COMPLETED
-
Evaluation of ChatGPT-5 and Anesthesiologist Predictions of Postoperative Intensive Care Unit Admission Using Preoperative Clinical Data
NCT07395791 ·Status: COMPLETED
-
AI vs. Anesthesiologists in Preoperative Triage
NCT07459491 ·Status: COMPLETED
-
Baseline Gastric Volume in Diabetic vs Non-Diabetic Patients
NCT07483593 ·Status: NOT_YET_RECRUITING
-
Assessment of the Patients of Emergency Consultation
NCT02216929 ·Status: COMPLETED
-
Artificial Intelligence in Lung Ultrasound for Preeclampsia
NCT05487014 ·Status: COMPLETED
-
Ultrasound-Based Airway Assessment for Predicting Difficult Intubation in Adult Female Patients
NCT07343557 ·Status: COMPLETED
-
Gastric Emptying Time After Turkish Breakfast
NCT04789785 ·Status: COMPLETED
-
Ultrasound vs Clinical Parameters for Difficult Airway Prediction in Obesity
NCT07528677 ·Status: NOT_YET_RECRUITING
-
Ultrasound Predictors of Difficult Airway in Adults
NCT07383610 ·Status: COMPLETED
-
Evaluation of Double Lumen Tube Intubation Difficulty With Photo-Based Artificial Intelligence Algorithms
NCT06839261 ·Status: COMPLETED
-
Determining the Consistency Between Nurses and Artificial Intelligence (ChatGPT-5) in Delivering Scenario-Based Discharge Education to Coronary Artery Bypass Graft Patients: A Methodological Study
NCT07263724 ·Status: NOT_YET_RECRUITING
-
Assessment Of Difficult Airway Predictors: A Prospective Comparison Of Upper Airway Ultrasound And Conventional Anthropometric Measures
NCT07309978 ·Status: RECRUITING
-
AI-Based ASA Classification in Preoperative Patients
NCT07407998 ·Status: COMPLETED
-
Evaluation of Difficult Airway With Ultrasonography
NCT04289597 ·Status: COMPLETED
-
Comparison of Artificial Intelligence and Anesthesiologist in Preoperative Risk Assessment
NCT07364942 ·Status: COMPLETED
-
Comparison of Gastric Volume in I-gel and ProSeal Laryngeal Mask Airways
NCT07138092 ·Status: NOT_YET_RECRUITING ·Phase: NA
-
The Impact of Ultrasound Measurements in Predicting Difficult Airway on Videolaryngoscopy Success
NCT07628868 ·Status: ACTIVE_NOT_RECRUITING
-
Anthropometric and US-Guided Difficult Intubation Prediction With ML Models
NCT06904586 ·Status: COMPLETED