Large Language Models Versus Human Examiners for Grading Physiotherapy Clinical Cases
NCT07677202 · Status: NOT_YET_RECRUITING · Type: OBSERVATIONAL · Enrollment: 65
Last updated 2026-06-30
Summary
This study evaluates whether large language models (LLMs) can reliably assess written clinical-reasoning case examinations completed by undergraduate physiotherapy students, compared with faculty assessment. In the course "Specific Methods in Physiotherapy" (third year of the Physiotherapy Degree), students solve complex clinical cases that require clinical reasoning, technical knowledge, and therapeutic decision-making. These cases are traditionally graded by faculty, a time-consuming process that may show inter-rater variability.
A set of de-identified student case examinations will be assessed using the rubric currently applied in the course, which covers clarity and structure of clinical reasoning, integration of the biopsychosocial model (ICF and APTA frameworks), accuracy in identifying pain mechanisms, coherence between diagnosis, hypotheses, and treatment, originality and depth of analysis, and professional writing. Each examination will be scored independently by three LLMs (for example, Claude, ChatGPT, and Gemini), each receiving an identical standardized prompt that embeds the same rubric, and by faculty serving as the reference standard.
To avoid overloading faculty, full double human grading may not be feasible; the human reference will therefore consist of expert faculty grading by one independent rater or, when resources allow, two independent raters. In contrast, paired assessment is fully implemented across the AI models: each examination is scored by several LLMs, and each model is queried in duplicate, allowing the study to estimate agreement between models and the test-retest stability of each model.
The primary aim is to quantify agreement between LLM-generated scores and the faculty reference score. Secondary aims include agreement among the LLMs, test-retest reliability of each model, criterion-level agreement, the quality and usefulness of the qualitative feedback generated, the time and cost associated with each approach, and students' perceptions of the usefulness of human versus AI feedback.
The findings will clarify the strengths and limitations of LLMs as supportive tools for formative assessment in health-professions education and will inform criteria for their responsible and effective use. No LLM output will affect students' official grades, which remain the sole responsibility of faculty.
Conditions
- Educational Assessment
- Artifical Intelligence
- Physical Therapy Education
Interventions
- DIAGNOSTIC_TEST
-
LLM-based assessment
Assessment of each anonymized examination by three large language models (for example, Claude, ChatGPT, and Gemini, in the versions available during data collection). Each model receives an identical standardized prompt embedding the study rubric and returns a score per criterion, a global score, and structured qualitative feedback. Each model is queried in duplicate in independent sessions under fixed generation parameters to estimate intra-model (test-retest) reliability, and outputs are compared across models to estimate inter-model agreement.
- DIAGNOSTIC_TEST
-
Faculty assessment (reference standard)
Assessment of the same anonymized examinations by faculty with expertise in the course, applying the identical rubric, serving as the reference standard. In the preferred scenario, two faculty members score each examination independently (paired human correction); if faculty workload precludes this, a single expert faculty rating, or the official course grade already assigned, is used as the reference. Faculty and LLM raters are blinded to one another's scores.
Sponsors & Collaborators
-
Centro Universitario La Salle
collaborator OTHER -
Neuron, Spain
lead OTHER
Eligibility
- Min Age
- 18 Years
- Sex
- ALL
- Healthy Volunteers
- Yes
Timeline & Regulatory
- Start
- 2026-08-01
- Primary Completion
- 2026-08-10
- Completion
- 2026-08-10
Countries
- Spain
Study Locations
More Related Trials
-
Video Analysis and Artificial Intelligence for the Analysis of Upper Limb Movement in Children
NCT07147478 ·Status: NOT_YET_RECRUITING
-
Functional Assessment Protocol for the Upper Limb for Pediatric Age
NCT06400667 ·Status: ACTIVE_NOT_RECRUITING ·Phase: NA
-
Influence of the Global Postural Reeducation and the Personality in the Posture
NCT02175667 ·Status: COMPLETED ·Phase: NA
-
The Influence of 3D Printed Prostheses on Neural Activation Patterns
NCT04110730 ·Status: RECRUITING ·Phase: NA
-
Reliability and Validity of the ACTIVE-mini for Quantifying Movement in Infants With Spinal Muscular Atrophy
NCT03808233 ·Status: COMPLETED
-
Robot Asissted Training on Neurodevelopmental Alterations
NCT06645795 ·Status: ACTIVE_NOT_RECRUITING ·Phase: NA
-
Evaluation of Functional, Neuroplastic and Biomechanical Changes Induced by an Intensive, Playful Early-morning Treatment Including Lower Limbs (EARLY-HABIT-ILE) in Preschool Children With Uni and Bilateral Cerebral Palsy
NCT04017871 ·Status: COMPLETED ·Phase: NA
-
AI-based Model for Rehabilitation Engagement and Motor Performance Evaluation in Pediatric Patients: A Pilot Study
NCT07664033 ·Status: NOT_YET_RECRUITING ·Phase: NA
-
Behavioral Assessment Method Index
NCT07291479 ·Status: RECRUITING
-
The Effects of Bilateral and Unilateral Perceptual-Motor Exercises on Manual Dexterity and Visuospatial Memory in Children With Nonverbal Learning Disorder
NCT07407270 ·Status: COMPLETED ·Phase: NA
-
Evaluation of Gait Performance in Children With Neuromuscular Diagnoses
NCT02237222 ·Status: COMPLETED
-
Assessment of Neural and Motor Performance
NCT04241588 ·Status: WITHDRAWN
-
Non-Immersive Virtual Reality for Gait and Balance Training in Children With Cerebral Palsy
NCT07612332 ·Status: NOT_YET_RECRUITING ·Phase: NA
-
Effects of an HABIT-ILE-Based Intervention in Children With Cerebral Palsy
NCT07506837 ·Status: NOT_YET_RECRUITING ·Phase: NA
-
Development and Effectiveness of the Participatory Adapted 3D Sedentation System (SSAP3D).
NCT07330921 ·Status: RECRUITING ·Phase: NA
-
Functional Effects and Impact on Motor Neuronal Activity of Early and Intensive Motrice (Hand and Arm Bimanual Intensive Therapy Including Lower Extremities: HABIT-ILE) Reeducation in Children With Pre-school Bilateral Cerebral Palsy
NCT04362800 ·Status: COMPLETED ·Phase: NA
-
Locomotor Learning in Infants at High Risk for Cerebral Palsy
NCT04561232 ·Status: COMPLETED ·Phase: NA
-
Acute Effects of a Passive Stretching Session on the Mechanical Properties of Medial Gastrocnemius Muscle in Children With Cerebral Palsy
NCT03714269 ·Status: COMPLETED ·Phase: NA
-
Arm More + Camp: An Implementation Study
NCT07065916 ·Status: ENROLLING_BY_INVITATION ·Phase: NA
-
Comparision Of The Effectiveness Of Physiotherapy Methods
NCT04328168 ·Status: COMPLETED ·Phase: NA
-
Effects of DMI vs Bobath on Neuromuscular Development in CP
NCT07238634 ·Status: COMPLETED ·Phase: NA
-
Actigraphs for Detection of Asymmetries
NCT03054441 ·Status: COMPLETED
-
Mirror Therapy in Children With Hemiparetic Cerebral Palsy
NCT07734376 ·Status: NOT_YET_RECRUITING ·Phase: NA
-
Effectiveness of a Treatment With the Robot - Assisted Gait Training System Walkbot in Patients With Cerebral Palsy
NCT04329793 ·Status: COMPLETED ·Phase: NA
-
Physical Therapy to Prevent Osteopenia in Preterm Infants
NCT04356807 ·Status: COMPLETED ·Phase: NA