1.Diagnostic Accuracy of Serological Tests for Mycoplasma pneumoniae Infections in Children with Pneumonia, Based on Symptom Onset
Gahee KIM ; Ki Wook YUN ; Dayun KANG ; Taek Jin LEE ; Byung Wook EUN ; Hyunju LEE ; Yae-Jean KIM ; Doo Ri KIM ; Areum SHIN ; Hyun Mi KANG ; Ye Ji KIM ; Byung Ok KWAK ; Younghee LEE ; Ye Kyung KIM ; Young June CHOE ; Woosuck SUH ; Kyo Jin JO ; Kyung-Ran KIM ; Eun Young CHO ; Kyung Min KIM ; Joon Kee LEE ; Su Eun PARK
Annals of Laboratory Medicine 2026;46(2):162-170
Background:
Mycoplasma pneumoniae is a major cause of community-acquired pneumonia (CAP) in children, with a rising incidence of macrolide resistance. Early diagnosis is crucial for reducing the disease burden; however, current diagnostic tools have limitations.We evaluated the diagnostic accuracy of serological assays and their performance based on symptom onset in children with CAP.
Methods:
From September 2023 to September 2024, we prospectively enrolled children with CAP, classified as M. pneumoniae pneumonia (MPP) or non-MPP, from 16 hospitals in Korea. Serological testing included chemiluminescence immunoassay (CLIA) and ELISA for detecting IgM and IgG, along with particle agglutination (PA) for total antibody measurements. Serological responses were analyzed at different times after symptom onset (0–4, 5–9, and 10–21 days).
Results:
Among 472 children with CAP (362 MPP, 110 non-MPP), 138 (29.2%) underwent PA testing, and 334 (70.8%) underwent IgM testing. PA at a 1:640 cutoff showed 48.0% sensitivity and 100% specificity. CLIA and ELISA showed comparable sensitivities (69.1% vs. 69.2%) and specificities (76.9% vs. 66.7%) for IgM testing. Seropositivity increased significantly with time since symptom onset (P for trend < 0.001), reaching 97.9% for IgM, 62.5% for IgG, and 94.7% for PA at 10–21 days.
Conclusions
The time post-symptom onset significantly influenced the diagnostic utility of serological tests for pediatric MPP, which showed limited value during the early stage of illness. These findings emphasize the importance of symptom onset-based interpretation of serological test results and their utility in complementing PCR when optimizing MPP diagnosis in children.
2.Early prediction of transient versus permanent congenital hypothyroidism: a retrospective cohort study
Myung Ji YOO ; Ji-Eun LEE ; Eun Young JOO ; Jisun PARK ; Young Ju SUH ; Su Jin KIM
Annals of Pediatric Endocrinology & Metabolism 2026;31(1):38-44
Purpose:
Early differentiation between transient congenital hypothyroidism (TCH) and permanent congenital hypothyroidism (PCH) is crucial for optimizing the duration of treatment. This retrospective cohort study aimed to evaluate whether levothyroxine (LT4) dose requirements over time can predict TCH and guide earlier discontinuation of treatment.
Methods:
We retrospectively analyzed 105 infants with congenital hypothyroidism and normal thyroid glands confirmed by imaging at a single tertiary care center (Inha University Hospital) between January 2013 and December 2022. Patients were classified into TCH (n=70) or PCH (n=35) based on thyroid function after LT4 withdrawal at 3 years of age. LT4 dose/kg at 6, 12, and 24 months, along with clinical and biochemical parameters, were compared between the 2 groups. Receiver operating characteristic (ROC) curve analysis was used to assess the predictive performance of LT4 dose thresholds.
Results:
The LT4 dose was significantly lower in the TCH group at 6 (3.16±0.83 μg/kg vs. 3.75±0.99 μg/kg, P=0.005), 12 (2.51±0.82 μg/kg vs. 3.37±1.17 μg/kg, P<0.001), and 24 months (2.02±0.61 μg/kg vs. 3.09±1.19 μg/kg, P<0.001). ROC curve analysis showed an area under the curve (AUC) of 0.649, 0.746, and 0.794 at 6, 12, and 24 months, respectively. A logistic regression model incorporating LT4 dose, birth weight, and thyroid-stimulating hormone (TSH) levels improved prediction accuracy (AUC: 0.740, 0.782, 0.833 at 6, 12, and 24 months, respectively).
Conclusion
LT4 dose requirements at 6, 12, and 24 months serve as useful indicators for differentiating TCH from PCH. A combined predictive model incorporating LT4 dose, birth weight, and TSH levels may improve diagnostic accuracy, supporting earlier discontinuation of treatment.
3.Myopia Management Consensus Statement in South Korean Children 2025 by the Korean Myopia Society for the Korean Association for Pediatric Ophthalmology and Strabismus
Yeon-Hee LEE ; Jae Yun SUNG ; Sun Young SHIN ; Young-Woo SUH ; Ungsoo Samuel KIM ; Hyunkyung KIM ; Kyung-Ah PARK ; Su Jin KIM ; MiRae KIM ; Hyun Jin SHIN ; Kyeong Wook LEE ; Haeng-Jin LEE ; So Young HAN ; Jinu HAN ; Eun Hee HONG ; Seung-Hee Hannah BAEK ; Hae Jung PAIK ;
Korean Journal of Ophthalmology 2026;40(2):185-205
Myopia, particularly high myopia, is a significant risk factor for several ocular pathologies including cataract, glaucoma, and retinal detachment. Excessive axial elongation associated with high myopia can induce biomechanical stretching, increasing the risk of serious complications like posterior staphyloma and myopic maculopathy. Global meta-analyses estimate that approximately 10 million people were visually impaired due to myopic maculopathy in 2015, with 3 million being blind. Recent nationwide surveys in South Korea revealed a prevalence of 65.4% for myopia and 6.9% for high myopia in children and adolescents, highlighting the urgent need for effective management. Delaying the onset and slowing the progression of myopia during childhood and adolescence is crucial for reducing the potential lifetime risk of these complications. This consensus statement, prepared by the Korean Myopia Society for the Korean Association for Pediatric Ophthalmology and Strabismus (KAPOS), reviews the current evidence for myopia control interventions and provides management strategies applicable to the South Korean clinical setting. Key interventions covered include lifestyle modifications (outdoor time, near work adjustment), optical methods (myopia-control spectacle lenses, dual-focus soft contact lenses, orthokeratology), and pharmacologic treatment (low-concentration atropine), as well as combination therapies. The statement also addresses patient selection, treatment outcome evaluation using spherical equivalent and axial length changes, and the crucial aspects related to treatment cessation and the rebound effect.
4.Evaluating the Accuracy and Diagnostic Reasoning of Multimodal Large Language Models in Interpreting Neuroradiology Cases From RadioGraphics
Pae Sun SUH ; Ji Su KO ; Woo Hyun SHIM ; Hwon HEO ; Chang-Yun WOO ; Hyungjun PARK ; Chong Hyun SUH
Korean Journal of Radiology 2026;27(3):214-226
Objective:
To evaluate the accuracy and reasoning capabilities of large multimodal language models compared with those of neuroradiology subspecialty-trained radiologists in neuroradiology case interpretation.
Materials and Methods:
This experimental study used custom-made 401 radiologic quizzes derived from articles published in RadioGraphics covering neuroradiology and head and neck topics (October 2020 to February 2024). We prompted the GPT-4 Turbo with Vision (GPT-4V), GPT-4 Omni, Gemini Flash, and Claude models to provide the top three differential diagnoses with a rationale and describe examination characteristics such as imaging modality, sequence, use of contrast, image plane, and body part. The temperature was adjusted to 0 and 1 (T1). Two neuroradiologists answered the same questions.The accuracies of the large language models (LLMs) and the neuroradiologists were compared using generalized estimating equations. Three neuroradiologists assessed the rationale provided by the LLMs for their differential diagnoses using four-point scales, separately for specific lesion locations and imaging findings, and evaluated the presence of hallucinations and the overall acceptability of the responses.
Results:
Top-3 accuracy (i.e., correct answers present among top-3 differential diagnoses) of LLMs ranged from 29.9% (120 of 401) to 49.4% (198 of 401, obtained with GPT-4V in the T1 setting), while radiologists achieved 80.3% (322 of 401) and 68.3% (274 of 401), respectively (P < 0.001). Regarding the rationale for differential diagnoses, GPT-4V (T1) accurately identified both the specific lesion location and imaging findings in 30.7% (123 of 401) and 12.9% (16 of 124) of cases without textual clinical history. Hallucinations occurred in 4.5% (18 of 401), and only 29.4% (118 of 401) of the LLM-generated analyses were deemed acceptable. GPT-4V (T1) demonstrated high accuracy in identifying the imaging modality (97.4% [800 of 821]) and scanned body parts (92.2% [756 of 820]).
Conclusion
LLMs remarkably underperformed compared with neuroradiologists and showed unsatisfactory reasoning for their differential diagnoses, with performance declining further in cases without textual input of clinical history. These findings highlight the limitations of current multimodal LLMs in neuroradiological interpretation and their reliance on text input.
6.Korean colorectal cancer screening guidelines for asymptomatic, average-risk adults: the 2025 revision
EunKyo KANG ; Jae Myung CHA ; Seo Young KANG ; Kiheon LEE ; Su Young KIM ; Younghoon KIM ; An Na SEO ; Hyo-Jin KANG ; Jong Keon JANG ; Kwang-Pil KO ; Aesun SHIN ; Dae Kyung SOHN ; Youngki HONG ; Eun-Jung CHO ; Minje HAN ; Soo Young KIM ; Hyeon Ji LEE ; Chang Kyun CHOI ; Mina SUH
Journal of the Korean Medical Association 2026;69(3):268-280
Purpose:
To develop the 2025 update to the Korean colorectal cancer (CRC) screening guidelines by systematically evaluating recent evidence, integrating domestic data, and addressing changes since the 2015 guideline revision, thereby providing an evidence-based standard for clinicians and policymakers.
Methods:
A multidisciplinary committee developed the guidelines using the Grading of Recommendations, Assessment, Development and Evaluation (GRADE) methodology. The process included formulation of three key questions addressing screening efficacy, diagnostic accuracy, and optimal screening age and interval. A systematic review of international guidelines and primary literature was conducted, yielding 327 eligible studies. In addition, a utility-based analysis using a Markov model was performed to determine optimal screening ages and intervals.
Results:
The evidence synthesis identified high-certainty evidence supporting the use of the fecal immunochemical test (FIT) for reducing CRC mortality and moderate-certainty evidence for colonoscopy. Evidence for computed tomographic colonography (CTC) and stool DNA testing was rated as very low certainty. Based on the evidence review and cost-utility analysis, the committee conditionally recommends CRC screening for asymptomatic, average-risk adults aged 45–74 years using either colonoscopy every 10 years or FIT every 1–2 years. CTC and stool DNA testing were not recommended owing to insufficient evidence.
Conclusion
The 2025 Korean Guidelines for Colorectal Cancer Screening present updated, evidence-based recommendations tailored to the domestic healthcare context. By conditionally endorsing both colonoscopy and FIT for individuals aged 45–74 years, these guidelines aim to improve population-level screening effectiveness and reduce the burden of CRC in South Korea.
7.Comparing Susceptibility-Weighted Imaging and T2* Gradient-Recalled Echo for Cerebral Microbleeds Detection: A Systematic Review and Meta-Analysis
Su Jeong YANG ; Jae‑Sung LIM ; Yangsean CHOI ; Ho Sung KIM ; Sang Joon KIM ; Jae-Hong LEE ; Chong Hyun SUH
Journal of Clinical Neurology 2026;22(2):193-202
Background:
and Purpose Criteria for amyloid-related imaging abnormalities in anti-amyloid therapy are based on T2* gradient-recalled echo (GRE), but susceptibility-weighted imaging (SWI) is widely used, creating uncertainty. This study quantitatively compared the detectability of SWI and GRE for cerebral microbleeds and established evidence supporting distinct microbleed criteria for each.
Methods:
A systematic review and meta-analysis were conducted following PRISMA guidelines. PubMed and Embase were searched for studies directly comparing SWI and GRE up to August 8, 2024. Study quality was assessed with QUADAS-2. The pooled proportion of microbleed detection and detection ratio were calculated. Subgroup analyses were performed based on magnetic field strength (1.5 T vs. 3 T) and SWI slice thickness (<2 mm vs. ≥2 mm), equipment vendor, and study quality.
Results:
Thirteen studies were included. SWI detected cerebral microbleeds approximately 1.6times more effectively than GRE. At 3.0 T and 1.5 T, SWI exhibited 1.7-fold and 1.5-fold greater detectability, respectively. SWI with thinner slices (<2 mm) showed a 1.9-fold improvement, while thicker slices (≥2 mm) showed a 1.3-fold improvement. Subgroup analyses revealed no significant differences between vendors (0.61 vs. 0.60, p=0.89), or by study quality (0.61 vs. 0.59,p=0.89).
Conclusions
SWI detects cerebral microbleeds about 1.6 times more effectively than GRE, highlighting important differences between the two techniques. Cautious exploration of adjusted thresholds may be needed, and prospective validation in therapy-specific cohorts will be essential before clinical application.
8.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
9.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
10.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.

Result Analysis
Print
Save
E-mail