1.Defect Size-Based Comparative Analysis of Treatment Modalities for Esophagojejunal Anastomotic Leakage Following Gastrectomy
Ba Ool SEONG ; Ji Yong AHN ; Juno YOO ; Chang Seok KO ; Sa-Hong MIN ; Chung Sik GONG ; Beom Su KIM ; Moon-Won YOO ; Jeong Hwan YOOK ; Hee Jin CHOI ; In-Seob LEE
Journal of Gastric Cancer 2026;26(2):295-306
Purpose:
Esophagojejunal anastomotic leakage (EJAL) represents a severe postoperative complication following total or proximal gastrectomy. Treatment strategies include conservative management, endoscopic interventions, and surgery; however, comparative data remain limited. This study aimed to compare clinical outcomes of different strategies to identify the optimal approach based on anastomotic defect size.
Materials and Methods:
This retrospective study reviewed 100 patients diagnosed with EJAL between January 2015 and October 2024. Patients were categorized into four groups:conservative management, endoscopic vacuum-assisted closure (E-VAC), other endoscopic treatments, and surgery. The primary outcomes were leakage duration and length of hospital stay after EJAL diagnosis, whereas the secondary outcome was time to C-reactive protein normalization. Subgroup analyses were performed according to defect size.
Results:
Among the 100 patients, 76 were male and 24 were female, with a mean age of 65.7 years. Conservative treatment was the most common modality (53%), followed by other endoscopic treatments (19%), E-VAC (14%), and surgery (14%). In patients with a defect size <1 cm, conservative treatment was associated with significantly shorter leakage duration (P=0.035) and earlier resumption of diet (P=0.029) compared with endoscopic treatment.Among those with defects ≥2 cm, E-VAC demonstrated the most favorable median outcomes across all variables; however, statistical significance was not achieved because of the small sample size.
Conclusions
Conservative treatment appears to be the most effective treatment strategy for EJAL with anastomotic defects <1 cm. For larger defects (≥2 cm), E-VAC may offer clinical benefit, although further studies are needed to confirm its efficacy. These findings highlight the importance of individualized treatment selection based on defect size.
2.Evaluating the Accuracy and Diagnostic Reasoning of Multimodal Large Language Models in Interpreting Neuroradiology Cases From RadioGraphics
Pae Sun SUH ; Ji Su KO ; Woo Hyun SHIM ; Hwon HEO ; Chang-Yun WOO ; Hyungjun PARK ; Chong Hyun SUH
Korean Journal of Radiology 2026;27(3):214-226
Objective:
To evaluate the accuracy and reasoning capabilities of large multimodal language models compared with those of neuroradiology subspecialty-trained radiologists in neuroradiology case interpretation.
Materials and Methods:
This experimental study used custom-made 401 radiologic quizzes derived from articles published in RadioGraphics covering neuroradiology and head and neck topics (October 2020 to February 2024). We prompted the GPT-4 Turbo with Vision (GPT-4V), GPT-4 Omni, Gemini Flash, and Claude models to provide the top three differential diagnoses with a rationale and describe examination characteristics such as imaging modality, sequence, use of contrast, image plane, and body part. The temperature was adjusted to 0 and 1 (T1). Two neuroradiologists answered the same questions.The accuracies of the large language models (LLMs) and the neuroradiologists were compared using generalized estimating equations. Three neuroradiologists assessed the rationale provided by the LLMs for their differential diagnoses using four-point scales, separately for specific lesion locations and imaging findings, and evaluated the presence of hallucinations and the overall acceptability of the responses.
Results:
Top-3 accuracy (i.e., correct answers present among top-3 differential diagnoses) of LLMs ranged from 29.9% (120 of 401) to 49.4% (198 of 401, obtained with GPT-4V in the T1 setting), while radiologists achieved 80.3% (322 of 401) and 68.3% (274 of 401), respectively (P < 0.001). Regarding the rationale for differential diagnoses, GPT-4V (T1) accurately identified both the specific lesion location and imaging findings in 30.7% (123 of 401) and 12.9% (16 of 124) of cases without textual clinical history. Hallucinations occurred in 4.5% (18 of 401), and only 29.4% (118 of 401) of the LLM-generated analyses were deemed acceptable. GPT-4V (T1) demonstrated high accuracy in identifying the imaging modality (97.4% [800 of 821]) and scanned body parts (92.2% [756 of 820]).
Conclusion
LLMs remarkably underperformed compared with neuroradiologists and showed unsatisfactory reasoning for their differential diagnoses, with performance declining further in cases without textual input of clinical history. These findings highlight the limitations of current multimodal LLMs in neuroradiological interpretation and their reliance on text input.
4.Korean colorectal cancer screening guidelines for asymptomatic, average-risk adults: the 2025 revision
EunKyo KANG ; Jae Myung CHA ; Seo Young KANG ; Kiheon LEE ; Su Young KIM ; Younghoon KIM ; An Na SEO ; Hyo-Jin KANG ; Jong Keon JANG ; Kwang-Pil KO ; Aesun SHIN ; Dae Kyung SOHN ; Youngki HONG ; Eun-Jung CHO ; Minje HAN ; Soo Young KIM ; Hyeon Ji LEE ; Chang Kyun CHOI ; Mina SUH
Journal of the Korean Medical Association 2026;69(3):268-280
Purpose:
To develop the 2025 update to the Korean colorectal cancer (CRC) screening guidelines by systematically evaluating recent evidence, integrating domestic data, and addressing changes since the 2015 guideline revision, thereby providing an evidence-based standard for clinicians and policymakers.
Methods:
A multidisciplinary committee developed the guidelines using the Grading of Recommendations, Assessment, Development and Evaluation (GRADE) methodology. The process included formulation of three key questions addressing screening efficacy, diagnostic accuracy, and optimal screening age and interval. A systematic review of international guidelines and primary literature was conducted, yielding 327 eligible studies. In addition, a utility-based analysis using a Markov model was performed to determine optimal screening ages and intervals.
Results:
The evidence synthesis identified high-certainty evidence supporting the use of the fecal immunochemical test (FIT) for reducing CRC mortality and moderate-certainty evidence for colonoscopy. Evidence for computed tomographic colonography (CTC) and stool DNA testing was rated as very low certainty. Based on the evidence review and cost-utility analysis, the committee conditionally recommends CRC screening for asymptomatic, average-risk adults aged 45–74 years using either colonoscopy every 10 years or FIT every 1–2 years. CTC and stool DNA testing were not recommended owing to insufficient evidence.
Conclusion
The 2025 Korean Guidelines for Colorectal Cancer Screening present updated, evidence-based recommendations tailored to the domestic healthcare context. By conditionally endorsing both colonoscopy and FIT for individuals aged 45–74 years, these guidelines aim to improve population-level screening effectiveness and reduce the burden of CRC in South Korea.
5.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
6.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
7.Adherence of Studies on Large Language Models for Medical Applications Published in Leading Medical Journals According to the MI-CLEAR-LLM Checklist
Ji Su KO ; Hwon HEO ; Chong Hyun SUH ; Jeho YI ; Woo Hyun SHIM
Korean Journal of Radiology 2025;26(4):304-312
Objective:
To evaluate the adherence of large language model (LLM)-based healthcare research to the Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM) checklist, a framework designed to enhance the transparency and reproducibility of studies on the accuracy of LLMs for medical applications.
Materials and Methods:
A systematic PubMed search was conducted to identify articles on LLM performance published in high-ranking clinical medicine journals (the top 10% in each of the 59 specialties according to the 2023 Journal Impact Factor) from November 30, 2022, through June 25, 2024. Data on the six MI-CLEAR-LLM checklist items: 1) identification and specification of the LLM used, 2) stochasticity handling, 3) prompt wording and syntax, 4) prompt structuring, 5) prompt testing and optimization, and 6) independence of the test data—were independently extracted by two reviewers, and adherence was calculated for each item.
Results:
Of 159 studies, 100% (159/159) reported the name of the LLM, 96.9% (154/159) reported the version, and 91.8% (146/159) reported the manufacturer. However, only 54.1% (86/159) reported the training data cutoff date, 6.3% (10/159) documented access to web-based information, and 50.9% (81/159) provided the date of the query attempts. Clear documentation regarding stochasticity management was provided in 15.1% (24/159) of the studies. Regarding prompt details, 49.1% (78/159) provided exact prompt wording and syntax but only 34.0% (54/159) documented prompt-structuring practices. While 46.5% (74/159) of the studies detailed prompt testing, only 15.7% (25/159) explained the rationale for specific word choices. Test data independence was reported for only 13.2% (21/159) of the studies, and 56.6% (43/76) provided URLs for internet-sourced test data.
Conclusion
Although basic LLM identification details were relatively well reported, other key aspects, including stochasticity, prompts, and test data, were frequently underreported. Enhancing adherence to the MI-CLEAR-LLM checklist will allow LLM research to achieve greater transparency and will foster more credible and reliable future studies.
8.Proposal of age definition for early-onset gastric cancer based on the Korean Gastric Cancer Association nationwide survey data: a retrospective observational study
Seong-A JEONG ; Ji Sung LEE ; Ba Ool SEONG ; Seul-gi OH ; Chang Seok KO ; Sa-Hong MIN ; Chung Sik GONG ; Beom Su KIM ; Moon-Won YOO ; Jeong Hwan YOOK ; In-Seob LEE ;
Annals of Surgical Treatment and Research 2025;108(4):245-255
Purpose:
This study aimed to define an optimal age cutoff for early-onset gastric cancer (EOGC) and compare its characteristics with those of late-onset gastric cancer (LOGC) using nationwide survey data.
Methods:
Using data from a nationwide survey, this comprehensive population-based study analyzed data spanning 3 years (2009, 2014, and 2019). The joinpoint analysis and interrupted time series (ITS) methodology were employed to identify age cutoffs for EOGC based on the sex ratio and tumor histology. Clinicopathologic characteristics and surgical outcomes were compared between the EOGC and LOGC groups.
Results:
The age cutoff for defining EOGC was suggested to be 50 years, supported by joinpoint and ITS analyses. Early gastric cancer was predominantly present in the EOGC and LOGC groups. Patients with EOGC comprised 20.3% of the total study cohort and demonstrated a more advanced disease stage compared to patients with LOGC. However, patients with EOGC underwent more minimally invasive surgeries, experienced shorter hospital stays, and had lower postoperative morbidity and mortality rates.
Conclusion
This study proposes an age of ≤50 years as a criterion for defining EOGC and highlights its features compared to LOGC. Further research using this criterion should guide tailored treatment strategies and improve outcomes for young patients with gastric cancer.
9.CORRIGENDUM: Proposal of age definition for early-onset gastric cancer based on the Korean Gastric Cancer Association nationwide survey data: a retrospective observational study
Seong-A JEONG ; Ji Sung LEE ; Ba Ool SEONG ; Seul-gi OH ; Chang Seok KO ; Sa-Hong MIN ; Chung Sik GONG ; Beom Su KIM ; Moon-Won YOO ; Jeong Hwan YOOK ; In-Seob LEE ;
Annals of Surgical Treatment and Research 2025;108(5):331-331
10.Characteristics and Prevalence of Sequelae after COVID-19: A Longitudinal Cohort Study
Se Ju LEE ; Yae Jee BAEK ; Su Hwan LEE ; Jung Ho KIM ; Jin Young AHN ; Jooyun KIM ; Ji Hoon JEON ; Hyeri SEOK ; Won Suk CHOI ; Dae Won PARK ; Yunsang CHOI ; Kyoung-Ho SONG ; Eu Suk KIM ; Hong Bin KIM ; Jae-Hoon KO ; Kyong Ran PECK ; Jae-Phil CHOI ; Jun Hyoung KIM ; Hee-Sung KIM ; Hye Won JEONG ; Jun Yong CHOI
Infection and Chemotherapy 2025;57(1):72-80
Background:
The World Health Organization has declared the end of the coronavirus disease 2019 (COVID-19) public health emergency. However, this did not indicate the end of COVID-19. Several months after the infection, numerous patients complain of respiratory or nonspecific symptoms; this condition is called long COVID. Even patients with mild COVID-19 can experience long COVID, thus the burden of long COVID remains considerable. Therefore, we conducted this study to comprehensively analyze the effects of long COVID using multi-faceted assessments.
Materials and Methods:
We conducted a prospective cohort study involving patients diagnosed with COVID-19 between February 2020 and September 2021 in six tertiary hospitals in Korea. Patients were followed up at 1, 3, 6, 12, 18, and 24 months after discharge. Long COVID was defined as the persistence of three or more COVID-19-related symptoms. The primary outcome of this study was the prevalence of long COVID after the period of COVID-19.
Results:
During the study period, 290 patients were enrolled. Among them, 54.5 and 34.6% experienced long COVID within 6 months and after more than 18 months, respectively. Several patients showed abnormal results when tested for post-traumatic stress disorder (17.4%) and anxiety (31.9%) after 18 months. In patients who underwent follow-up chest computed tomography 18 months after COVID-19, abnormal findings remained at 51.9%. Males (odds ratio [OR], 0.17; 95% confidence interval [CI], 0.05–0.53; P=0.004) and elderly (OR, 1.04; 95% CI, 1.00–1.09; P=0.04) showed a significant association with long COVID after 12–18 months in a multivariable logistic regression analysis.
Conclusion
Many patients still showed long COVID after 18 months post SARS-CoV-2 infection. When managing these patients, the assessment of multiple aspects is necessary.

Result Analysis
Print
Save
E-mail