1.Evaluating the Accuracy and Diagnostic Reasoning of Multimodal Large Language Models in Interpreting Neuroradiology Cases From RadioGraphics
Pae Sun SUH ; Ji Su KO ; Woo Hyun SHIM ; Hwon HEO ; Chang-Yun WOO ; Hyungjun PARK ; Chong Hyun SUH
Korean Journal of Radiology 2026;27(3):214-226
Objective:
To evaluate the accuracy and reasoning capabilities of large multimodal language models compared with those of neuroradiology subspecialty-trained radiologists in neuroradiology case interpretation.
Materials and Methods:
This experimental study used custom-made 401 radiologic quizzes derived from articles published in RadioGraphics covering neuroradiology and head and neck topics (October 2020 to February 2024). We prompted the GPT-4 Turbo with Vision (GPT-4V), GPT-4 Omni, Gemini Flash, and Claude models to provide the top three differential diagnoses with a rationale and describe examination characteristics such as imaging modality, sequence, use of contrast, image plane, and body part. The temperature was adjusted to 0 and 1 (T1). Two neuroradiologists answered the same questions.The accuracies of the large language models (LLMs) and the neuroradiologists were compared using generalized estimating equations. Three neuroradiologists assessed the rationale provided by the LLMs for their differential diagnoses using four-point scales, separately for specific lesion locations and imaging findings, and evaluated the presence of hallucinations and the overall acceptability of the responses.
Results:
Top-3 accuracy (i.e., correct answers present among top-3 differential diagnoses) of LLMs ranged from 29.9% (120 of 401) to 49.4% (198 of 401, obtained with GPT-4V in the T1 setting), while radiologists achieved 80.3% (322 of 401) and 68.3% (274 of 401), respectively (P < 0.001). Regarding the rationale for differential diagnoses, GPT-4V (T1) accurately identified both the specific lesion location and imaging findings in 30.7% (123 of 401) and 12.9% (16 of 124) of cases without textual clinical history. Hallucinations occurred in 4.5% (18 of 401), and only 29.4% (118 of 401) of the LLM-generated analyses were deemed acceptable. GPT-4V (T1) demonstrated high accuracy in identifying the imaging modality (97.4% [800 of 821]) and scanned body parts (92.2% [756 of 820]).
Conclusion
LLMs remarkably underperformed compared with neuroradiologists and showed unsatisfactory reasoning for their differential diagnoses, with performance declining further in cases without textual input of clinical history. These findings highlight the limitations of current multimodal LLMs in neuroradiological interpretation and their reliance on text input.
3.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
4.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
5.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
6.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
7.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
8.Comparative Analysis Between Single and Double 3-Dimensional Printed Titanium Cages: 1-Year Outcomes After Unilateral Biportal Endoscopic Transforaminal Lumbar Interbody Fusion
Ji Yeon KIM ; Hyun Jin HONG ; Su Yong CHOI ; Dong Chan LEE ; Hyeun Sung KIM ; Dong Hwa HEO
Journal of Minimally Invasive Spine Surgery and Technique 2025;10(Suppl 2):S225-S234
Objective:
This study aimed to illustrate the techniques of unilateral biportal endoscopic transforaminal lumbar interbody fusion (UBE-TLIF) using double cages and compare surgical outcomes with those using a single cage.
Methods:
We retrospectively analyzed 62 patients who underwent single-level UBE-TLIF using 3-dimensional (3D)-printed titanium cages (29 with a single cage and 33 with double cages). Radiological parameters, including Bridwell fusion and subsidence grading, were assessed via x-ray and computed tomography at 6 months and 1 year. Clinical outcomes were measured using the visual analogue scale (VAS) for lower back and leg pain, as well as the Oswestry Disability Index (ODI).
Results:
At the 6-month follow-up, the overall fusion rate (grades I + II) was significantly higher in the double-cage group (86.2% vs. 100%, p=0.04); however, no significant difference was noted at the 1-year follow-up (89.6% vs. 100%). The double-cage group showed lower cage subsidence rates at both follow-ups (3% vs. 31%, p=0.01; 12.1% vs. 38%, p=0.03). The double-cage group exhibited significantly greater improvements in leg pain (VAS: 7.6 to 1.7 vs. 7.2 to 1.4, p=0.03) and ODI (28.6 to 10.8 vs. 28.3 to 9.4, p=0.01) at 1 year.
Conclusion
UBE-TLIF with 3D-printed titanium cages achieved a favorable fusion rate at the 1-year follow-up, regardless of whether single or double cages were used. However, double cages accelerated fusion at the 6-month follow-up and significantly reduced cage subsidence at both the 6-month and 1-year follow-ups. These benefits contributed to improved VAS scores for leg pain and ODI at the final follow-up.
9.Dietary interventions to reduce heavy metal exposure in antepartum and postpartum women: a systematic review
Su Ji HEO ; Nalae MOON ; Ju Hee KIM
Women’s Health Nursing 2024;30(4):265-276
Heavy metals, which are persistent in the environment and toxic, can accumulate in the body and cause organ damage, which may further negatively affect perinatal women and their fetuses. Therefore, this systematic review was conducted to evaluate the effectiveness of dietary interventions to reduce heavy metal exposure in antepartum and postpartum women. Methods: We searched five databases (PubMed, Embase, Scopus, Web of Science, and Cochrane Library) for randomized controlled trials that provided dietary interventions for antepartum and postpartum women. Quality assessments were conducted independently by two reviewers using the Cochrane Risk-of-Bias tool, a quality assessment tool for randomized controlled trials. Results: A total of seven studies were included. The studies were conducted in six countries, with interventions categorized into “nutritional supplements,” “food supply,” and “educational” strategies. Interventions involving nutritional supplements, such as calcium and probiotics, primarily reduced heavy metal levels in the blood and minimized toxicity. Food-based interventions, including specific fruit consumption, decreased heavy metal concentrations in breast milk. Educational interventions effectively promoted behavioral changes, such as adopting diets low in mercury. The studies demonstrated a low overall risk of bias, supporting the reliability of the findings. These strategies underscore the effectiveness of dietary approaches in mitigating heavy metal exposure and improving maternal and child health. Conclusion: The main findings underscore the importance of dietary interventions in reducing heavy metal exposure. This emphasizes the critical role of nursing in guiding dietary strategies to minimize exposure risks, ultimately supporting maternal and fetal health during pregnancy.
10.Prevalence and Associated Factors of Depression and Anxiety Among Healthcare Workers During the Coronavirus Disease 2019 Pandemic:A Nationwide Study in Korea
Shinwon LEE ; Soyoon HWANG ; Ki Tae KWON ; EunKyung NAM ; Un Sun CHUNG ; Shin-Woo KIM ; Hyun-Ha CHANG ; Yoonjung KIM ; Sohyun BAE ; Ji-Yeon SHIN ; Sang-geun BAE ; Hyun Wook RYOO ; Juhwan JEONG ; NamHee OH ; So Hee LEE ; Yeonjae KIM ; Chang Kyung KANG ; Hye Yoon PARK ; Jiho PARK ; Se Yoon PARK ; Bongyoung KIM ; Hae Suk CHEONG ; Ji Woong SON ; Su Jin LIM ; Seongcheol YUN ; Won Sup OH ; Kyung-Hwa PARK ; Ju-Yeon LEE ; Sang Taek HEO ; Ji-yeon LEE
Journal of Korean Medical Science 2024;39(13):e120-
Background:
A healthcare system’s collapse due to a pandemic, such as the coronavirus disease 2019 (COVID-19), can expose healthcare workers (HCWs) to various mental health problems. This study aimed to investigate the impact of the COVID-19 pandemic on the depression and anxiety of HCWs.
Methods:
A nationwide questionnaire-based survey was conducted on HCWs who worked in healthcare facilities and public health centers in Korea in December 2020. Patient Health Questionnaire-9 (PHQ-9) and Generalized Anxiety Disorder-7 (GAD-7) were used to measure depression and anxiety. To investigate factors associated with depression and anxiety, stepwise multiple logistic regression analysis was performed.
Results:
A total of 1,425 participating HCWs were included. The mean depression score (PHQ-9) of HCWs before and after COVID-19 increased from 2.37 to 5.39, and the mean anxiety score (GAD-7) increased from 1.41 to 3.41. The proportion of HCWs with moderate to severe depression (PHQ-9 ≥ 10) increased from 3.8% before COVID-19 to 19.5% after COVID-19, whereas that of HCWs with moderate to severe anxiety (GAD-7 ≥ 10) increased from 2.0% to 10.1%. In our study, insomnia, chronic fatigue symptoms and physical symptoms after COVID-19, anxiety score (GAD-7) after COVID-19, living alone, and exhaustion were positively correlated with depression. Furthermore, post-traumatic stress symptoms, stress score (Global Assessment of Recent Stress), depression score (PHQ-9) after COVID-19, and exhaustion were positively correlated with anxiety.
Conclusion
In Korea, during the COVID-19 pandemic, HCWs commonly suffered from mental health problems, including depression and anxiety. Regularly checking the physical and mental health problems of HCWs during the COVID-19 pandemic is crucial, and social support and strategy are needed to reduce the heavy workload and psychological distress of HCWs.

Result Analysis
Print
Save
E-mail