Large language model-generated versus teacher-written objective structured clinical examination stations for medical students: a blinded comparative pilot study
- Author:
Piotr SZYCHOWIAK
1
;
Jonathan WONG-SO
;
Hélène MESSET
;
Mélanie FAURE
;
Isaure BRETEAU
;
Simon JAMARD
;
François BARBIER
;
Maxime DESGROUAS
Author Information
- Publication Type:Brief report
- From:Journal of Educational Evaluation for Health Professions 2026;23(1):9-
- CountryRepublic of Korea
- Abstract: Developing objective structured clinical examination (OSCE) stations is time-consuming for medical teachers. We aimed to evaluate the ability of a large language model (LLM) to generate ready-to-use OSCE stations. Five OSCE stations generated by the LLM GPT-4o were evaluated by 7 expert assessors using a 5-point Likert scale and compared with 5 teacher-written stations targeting similar learning objectives. A station was considered to be of good quality if most assessors responded “agree” or “strongly agree” to the statement “The station is good enough to be used by students.” All teacher-written stations were rated as being of good quality, compared with only one GPT-4o-generated station. The LLM produced adequate clinical scenarios when reference knowledge was provided and tasks were clearly ordered, but it failed to generate reliable assessment grids. Careful review by teachers remained essential. GPT-4o failed to consistently produce fully ready-to-use OSCE stations.
