1.Benchmarking Open-Source Vision Language Models in Orthopedic In-Training Examination:A Comparison with Residents, Domain-Specific Evaluation, and Parameter Scaling
Sunho KO ; Jaewook LEE ; Kyunga KO ; Jihyeung KIM
Clinics in Orthopedic Surgery 2026;18(1):159-166
Background:
Advancing orthopedic care through large language models requires both multimodal processing capabilities for medical images and open-source deployment options for secure in-house operations, yet these remain underexplored in current literature. This study aims to benchmark open-source vision-language models (VLMs) against orthopedic residents using the Orthopedic In-Training Examination (OITE), assess domain-specific performance across orthopedic subspecialties, and investigate the relationship between model parameter size and performance.
Methods:
Six open-source VLMs of varying sizes (Alibaba Qwen2.5-VL-72B-Instruct, Alibaba Qwen2.5-VL-32B-Instruct, Alibaba Qwen2.5-VL-7B-Instruct, Alibaba Qwen2.5-VL-3B-Instruct, Meta Llama-3.2-90B-Vision-Instruct, Meta Llama-3.2-11B-Vision-Instruct) were evaluated using the 2023 OITE (210 questions; 111 with images). Model performance was compared to resident scores from the 2023 OITE technical report. Pearson correlation coefficient was used to assess the association between model size and performance.
Results:
The 2 largest open-source models, Qwen2.5-VL-72B and Llama-3.2-90B, demonstrated performance levels comparable to those of second-year orthopedic residents on the OITE examination. A mid-sized model, Qwen-32B, slightly outscored first-year residents. In contrast, small-sized models (under 11 billion parameters) performed worse than first-year residents. Qwen2.5-VL-72B performed best in foot & ankle and sports medicine topics, while Llama-3.2-90B was strongest in basic science and hand & wrist.All models had the most difficulty with spine and pediatric questions. Overall, model accuracy increased steadily with model size up to 72 billion parameters, but larger sizes showed little additional improvement.
Conclusions
Smaller models offer reduced accuracy in exchange for lower hardware requirements. Spine and pediatric domains remain consistently areas of underperformance across all models. Model selection should be based on domain-specific benchmark results to balance clinical needs with hardware limitations. While promising, open-source VLMs currently require further refinement and validation before they can be reliably applied in clinical or educational settings.
2.Prospective Multicenter Study Comparing Magnetic Resonance Imaging and Ultrasonography for Second Breast Cancer Surveillance in Women With Prior Breast Cancer and Dense Breasts: KBCSG-27 Trial
Yun-Woo CHANG ; Young Mi PARK ; Kyunga KIM ; Min-Ji KIM ; Myoung Kyoung KIM ; Jonghan YU ; Eun Sook KO
Journal of Breast Cancer 2025;28(6):427-436
Purpose:
Surveillance guidelines following breast cancer surgery recommend mammography as the sole imaging modality. However, the accuracy of mammography is low in younger women and in those with dense breast tissue. Additional imaging modalities, such as ultrasonography and magnetic resonance imaging (MRI), may offer diagnostic benefits.This prospective, multicenter study (KBCSG-27) aims to compare the diagnostic performances of mammography, ultrasonography, and MRI for detecting second breast cancer (SBC) in women with a personal history of breast cancer (PHBC) and dense breasts.
Methods
This study will recruit approximately 1,756 women, aged 20–75 years, who were treated for stage 0–III breast cancer and have dense breast tissue on mammography.Participants will undergo two annual breast screenings, each consisting of mammography, ultrasonography, and MRI. MRI will be performed using either abbreviated magnetic resonance imaging (AB-MRI) or full-protocol magnetic resonance imaging (FP-MRI), which will be randomly assigned such that each participant receives both protocols alternately.Radiologists will independently interpret all images. A combination of pathology results and 12-month follow-up will serve as the reference standard. A patient-reported outcome (PRO) tool will be used to assess patients’ experiences and preferences between AB-MRI and FPMRI. The primary objective is to compare the cancer detection rates of ultrasonography versus AB-MRI and ultrasonography versus FP-MRI. Secondary outcomes include comparisons of the invasive cancer detection rates, abnormal interpretation rates, sensitivity, specificity, positive and negative predictive values, accuracy, and interval cancer rates. Subgroup analyses will be conducted based on age, menopausal status, mammographic breast density, and molecular subtype. Additionally, PRO results of AB-MRI and FP-MRI will be compared.Discussion: This ongoing, prospective, multicenter study aims to evaluate the performance of ultrasonography, AB-MRI, and FP-MRI in SBC surveillance in women with PHBC and dense breasts. Enrollment is expected to be completed by 2025, with results anticipated after 2028.
3.Total Knee Arthroplasty: Is It Safe? A Single-Center Study of 4,124 Patients in South Korea
Kyunga KO ; Kee Hyun KIM ; Sunho KO ; Changwung JO ; Hyuk-Soo HAN ; Myung Chul LEE ; Du Hyun RO
Clinics in Orthopedic Surgery 2023;15(6):935-941
Background:
Although total knee arthroplasty (TKA) is considered an effective treatment for knee osteoarthritis, it carries risks of complications. With a growing number of TKAs performed on older patients, understanding the cause of mortality is crucial to enhance the safety of TKA. This study aimed to identify the major causes of short- and long-term mortality after TKA and report mortality trends for major causes of death.
Methods:
A total of 4,124 patients who underwent TKA were analyzed. The average age at surgery was 70.7 years. The average follow-up time was 73.5 months. The causes of death were retrospectively collected through Korean Statistical Information Service and classified into 13 subgroups based on the International Classification of Diseases-10 code. The short- and long-term causes of death were identified within the time-to-death intervals of 30, 60, 90, 180, 180 days, and > 180 days. Standard mortality ratios (SMRs) and cumulative incidence of deaths were computed to examine mortality trends after TKA.
Results:
The short-term mortality rate was 0.07% for 30 days, 0.1% for 60 days, 0.2% for 90 days, and 0.2% for 180 days. Malignant neoplasm and cardiovascular disease were the main short-term causes of death. The long-term (> 180 days) mortality rate was 6.2%. Malignant neoplasm (35%), others (11.7%), and respiratory disease (10.1%) were the major long-term causes of death.Men had a higher cumulative risk of death for respiratory, metabolic, and cardiovascular diseases. Age-adjusted mortality was significantly higher in TKA patients aged 70 years (SMR, 4.3; 95% confidence interval [CI], 3.3–5.4) and between 70 and 79 years (SMR 2.9; 95% CI, 2.5–3.5) than that in the general population.
Conclusions
The short-term mortality rate after TKA was low, and most of the causes were unrelated to TKA. The major causes of long-term death were consistent with previous findings. Our findings can be used as counseling data to understand the survival and mortality of TKA patients.

Result Analysis
Print
Save
E-mail