1.Combined Transarterial Chemoembolization and External Beam Radiotherapy for Identifying Surgical Candidates for Hepatocellular Carcinoma with Macroscopic Vascular Invasion: A Propensity Score–Weighted Analysis
Sumin LEE ; Jinhong JUNG ; Jonggi CHOI ; So Yeon KIM ; Jin Hyoung KIM ; Danbi LEE ; Ju Hyun SHIM ; Kang Mo KIM ; Young-Suk LIM ; Han Chu LEE ; Gi-Won SONG ; Jin-hong PARK ; Sang Min YOON
Cancer Research and Treatment 2026;58(1):275-283
Purpose:
This study aimed to evaluate the role of hepatic resection in patients with objective responses after combined transarterial chemoembolization (TACE) and radiotherapy (RT) for hepatocellular carcinoma (HCC) with macroscopic vascular invasion (MVI).
Materials and Methods:
We retrospectively reviewed the patients treated with combined TACE and RT for HCC with MVI between 2010 and 2015. Some of the patients with objective responses underwent hepatic resection or liver transplantation; to investigate the impact of surgery, patients with objective responses who did not undergo surgery were selected as the control group. Survival outcomes were compared using a propensity score–based stabilized inverse probability of treatment weighting method.
Results:
Out of the 170 patients with objective responses after combined TACE and RT, 41 patients underwent surgery, including eight liver transplantations. The unweighted surgery group was younger and had a higher proportion of solitary tumors and unilateral vascular involvement. After adjustment, the 3-year overall survival (OS) rates were 61.0% and 28.6% in the surgery and non-surgery groups, respectively. The most important prognostic factor for OS was surgery (adjusted Cox hazard ratio [HR], 0.28; 95% confidence interval [CI], 0.17 to 0.46; p < 0.001). Complete response after TACE and RT (vs. partial response) was also a significant prognostic factor for OS (adjusted HR, 0.41; 95% CI, 0.27 to 0.61; p < 0.001). There was no surgical mortality. Four patients (9.8%) required additional surgery due to bleeding or graft failure.
Conclusion
Hepatic resection was significantly associated with improved OS in patients who showed objective responses after receiving combined TACE and RT for HCC with MVI.
2.Enhancing Identification of High-Risk cN0 Lung Adenocarcinoma Patients Using MRI-Based Radiomic Features
Harim KIM ; Jonghoon KIM ; Soohyun HWANG ; You Jin OH ; Joong Hyun AHN ; Min-Ji KIM ; Tae Hee HONG ; Sung Goo PARK ; Joon Young CHOI ; Hong Kwan KIM ; Jhingook KIM ; Sumin SHIN ; Ho Yun LEE
Cancer Research and Treatment 2025;57(1):57-69
Purpose:
This study aimed to develop a magnetic resonance imaging (MRI)–based radiomics model to predict high-risk pathologic features for lung adenocarcinoma: micropapillary and solid pattern (MPsol), spread through air space, and poorly differentiated patterns.
Materials and Methods:
As a prospective study, we screened clinical N0 lung cancer patients who were surgical candidates and had undergone both 18F-fluorodeoxyglucose (FDG) positron emission tomography–computed tomography (PET/CT) and chest CT from August 2018 to January 2020. We recruited patients meeting our proposed imaging criteria indicating high-risk, that is, poorer prognosis of lung adenocarcinoma, using CT and FDG PET/CT. If possible, these patients underwent an MRI examination from which we extracted 77 radiomics features from T1-contrast-enhanced and T2-weighted images. Additionally, patient demographics, maximum standardized uptake value on FDG PET/CT, and the mean apparent diffusion coefficient value on diffusion-weighted image, were considered together to build prediction models for high-risk pathologic features.
Results:
Among 616 patients, 72 patients met the imaging criteria for high-risk lung cancer and underwent lung MRI. The magnetic resonance (MR)–eligible group showed a higher prevalence of nodal upstaging (29.2% vs. 4.2%, p < 0.001), vascular invasion (6.5% vs. 2.1%, p=0.011), high-grade pathologic features (p < 0.001), worse 4-year disease-free survival (p < 0.001) compared with non-MR-eligible group. The prediction power for MR-based radiomics model predicting high-risk pathologic features was good, with mean area under the receiver operating curve (AUC) value measuring 0.751-0.886 in test sets. Adding clinical variables increased the predictive performance for MPsol and the poorly differentiated pattern using the 2021 grading system (AUC, 0.860 and 0.907, respectively).
Conclusion
Our imaging criteria can effectively screen high-risk lung cancer patients and predict high-risk pathologic features by our MR-based prediction model using radiomics.
3.Enhancing Identification of High-Risk cN0 Lung Adenocarcinoma Patients Using MRI-Based Radiomic Features
Harim KIM ; Jonghoon KIM ; Soohyun HWANG ; You Jin OH ; Joong Hyun AHN ; Min-Ji KIM ; Tae Hee HONG ; Sung Goo PARK ; Joon Young CHOI ; Hong Kwan KIM ; Jhingook KIM ; Sumin SHIN ; Ho Yun LEE
Cancer Research and Treatment 2025;57(1):57-69
Purpose:
This study aimed to develop a magnetic resonance imaging (MRI)–based radiomics model to predict high-risk pathologic features for lung adenocarcinoma: micropapillary and solid pattern (MPsol), spread through air space, and poorly differentiated patterns.
Materials and Methods:
As a prospective study, we screened clinical N0 lung cancer patients who were surgical candidates and had undergone both 18F-fluorodeoxyglucose (FDG) positron emission tomography–computed tomography (PET/CT) and chest CT from August 2018 to January 2020. We recruited patients meeting our proposed imaging criteria indicating high-risk, that is, poorer prognosis of lung adenocarcinoma, using CT and FDG PET/CT. If possible, these patients underwent an MRI examination from which we extracted 77 radiomics features from T1-contrast-enhanced and T2-weighted images. Additionally, patient demographics, maximum standardized uptake value on FDG PET/CT, and the mean apparent diffusion coefficient value on diffusion-weighted image, were considered together to build prediction models for high-risk pathologic features.
Results:
Among 616 patients, 72 patients met the imaging criteria for high-risk lung cancer and underwent lung MRI. The magnetic resonance (MR)–eligible group showed a higher prevalence of nodal upstaging (29.2% vs. 4.2%, p < 0.001), vascular invasion (6.5% vs. 2.1%, p=0.011), high-grade pathologic features (p < 0.001), worse 4-year disease-free survival (p < 0.001) compared with non-MR-eligible group. The prediction power for MR-based radiomics model predicting high-risk pathologic features was good, with mean area under the receiver operating curve (AUC) value measuring 0.751-0.886 in test sets. Adding clinical variables increased the predictive performance for MPsol and the poorly differentiated pattern using the 2021 grading system (AUC, 0.860 and 0.907, respectively).
Conclusion
Our imaging criteria can effectively screen high-risk lung cancer patients and predict high-risk pathologic features by our MR-based prediction model using radiomics.
4.Advancing Korean Medical Large Language Models: Automated Pipeline for Korean Medical Preference Dataset Construction
Jean SEO ; Sumin PARK ; Sungjoo BYUN ; Jinwook CHOI ; Jinho CHOI ; Hyopil SHIN
Healthcare Informatics Research 2025;31(2):166-174
Objectives:
Developing large language models (LLMs) in biomedicine requires access to high-quality training and alignment tuning datasets. However, publicly available Korean medical preference datasets are scarce, hindering the advancement of Korean medical LLMs. This study constructs and evaluates the efficacy of the Korean Medical Preference Dataset (KoMeP), an alignment tuning dataset constructed with an automated pipeline, minimizing the high costs of human annotation.
Methods:
KoMeP was generated using the DAHL score, an automated hallucination evaluation metric. Five LLMs (Dolly-v2-3B, MPT-7B, GPT-4o, Qwen-2-7B, Llama-3-8B) produced responses to 8,573 biomedical examination questions, from which 5,551 preference pairs were extracted. Each pair consisted of a “chosen” response and a “rejected” response, as determined by their DAHL scores. The dataset was evaluated when trained through two different alignment tuning methods, direct preference optimization (DPO) and odds ratio preference optimization (ORPO) respectively across five different models. The KorMedMCQA benchmark was employed to assess the effectiveness of alignment tuning.
Results:
Models trained with DPO consistently improved KorMedMCQA performance; notably, Llama-3.1-8B showed a 43.96% increase. In contrast, ORPO training produced inconsistent results. Additionally, English-to-Korean transfer learning proved effective, particularly for English-centric models like Gemma-2, whereas Korean-to-English transfer learning achieved limited success. Instruction tuning with KoMeP yielded mixed outcomes, which suggests challenges in dataset formatting.
Conclusions
KoMeP is the first publicly available Korean medical preference dataset and significantly improves alignment tuning performance in LLMs. The DPO method outperforms ORPO in alignment tuning. Future work should focus on expanding KoMeP, developing a Korean-native dataset, and refining alignment tuning methods to produce safer and more reliable Korean medical LLMs.
7.Advancing Korean Medical Large Language Models: Automated Pipeline for Korean Medical Preference Dataset Construction
Jean SEO ; Sumin PARK ; Sungjoo BYUN ; Jinwook CHOI ; Jinho CHOI ; Hyopil SHIN
Healthcare Informatics Research 2025;31(2):166-174
Objectives:
Developing large language models (LLMs) in biomedicine requires access to high-quality training and alignment tuning datasets. However, publicly available Korean medical preference datasets are scarce, hindering the advancement of Korean medical LLMs. This study constructs and evaluates the efficacy of the Korean Medical Preference Dataset (KoMeP), an alignment tuning dataset constructed with an automated pipeline, minimizing the high costs of human annotation.
Methods:
KoMeP was generated using the DAHL score, an automated hallucination evaluation metric. Five LLMs (Dolly-v2-3B, MPT-7B, GPT-4o, Qwen-2-7B, Llama-3-8B) produced responses to 8,573 biomedical examination questions, from which 5,551 preference pairs were extracted. Each pair consisted of a “chosen” response and a “rejected” response, as determined by their DAHL scores. The dataset was evaluated when trained through two different alignment tuning methods, direct preference optimization (DPO) and odds ratio preference optimization (ORPO) respectively across five different models. The KorMedMCQA benchmark was employed to assess the effectiveness of alignment tuning.
Results:
Models trained with DPO consistently improved KorMedMCQA performance; notably, Llama-3.1-8B showed a 43.96% increase. In contrast, ORPO training produced inconsistent results. Additionally, English-to-Korean transfer learning proved effective, particularly for English-centric models like Gemma-2, whereas Korean-to-English transfer learning achieved limited success. Instruction tuning with KoMeP yielded mixed outcomes, which suggests challenges in dataset formatting.
Conclusions
KoMeP is the first publicly available Korean medical preference dataset and significantly improves alignment tuning performance in LLMs. The DPO method outperforms ORPO in alignment tuning. Future work should focus on expanding KoMeP, developing a Korean-native dataset, and refining alignment tuning methods to produce safer and more reliable Korean medical LLMs.
9.Advancing Korean Medical Large Language Models: Automated Pipeline for Korean Medical Preference Dataset Construction
Jean SEO ; Sumin PARK ; Sungjoo BYUN ; Jinwook CHOI ; Jinho CHOI ; Hyopil SHIN
Healthcare Informatics Research 2025;31(2):166-174
Objectives:
Developing large language models (LLMs) in biomedicine requires access to high-quality training and alignment tuning datasets. However, publicly available Korean medical preference datasets are scarce, hindering the advancement of Korean medical LLMs. This study constructs and evaluates the efficacy of the Korean Medical Preference Dataset (KoMeP), an alignment tuning dataset constructed with an automated pipeline, minimizing the high costs of human annotation.
Methods:
KoMeP was generated using the DAHL score, an automated hallucination evaluation metric. Five LLMs (Dolly-v2-3B, MPT-7B, GPT-4o, Qwen-2-7B, Llama-3-8B) produced responses to 8,573 biomedical examination questions, from which 5,551 preference pairs were extracted. Each pair consisted of a “chosen” response and a “rejected” response, as determined by their DAHL scores. The dataset was evaluated when trained through two different alignment tuning methods, direct preference optimization (DPO) and odds ratio preference optimization (ORPO) respectively across five different models. The KorMedMCQA benchmark was employed to assess the effectiveness of alignment tuning.
Results:
Models trained with DPO consistently improved KorMedMCQA performance; notably, Llama-3.1-8B showed a 43.96% increase. In contrast, ORPO training produced inconsistent results. Additionally, English-to-Korean transfer learning proved effective, particularly for English-centric models like Gemma-2, whereas Korean-to-English transfer learning achieved limited success. Instruction tuning with KoMeP yielded mixed outcomes, which suggests challenges in dataset formatting.
Conclusions
KoMeP is the first publicly available Korean medical preference dataset and significantly improves alignment tuning performance in LLMs. The DPO method outperforms ORPO in alignment tuning. Future work should focus on expanding KoMeP, developing a Korean-native dataset, and refining alignment tuning methods to produce safer and more reliable Korean medical LLMs.
10.Enhancing Identification of High-Risk cN0 Lung Adenocarcinoma Patients Using MRI-Based Radiomic Features
Harim KIM ; Jonghoon KIM ; Soohyun HWANG ; You Jin OH ; Joong Hyun AHN ; Min-Ji KIM ; Tae Hee HONG ; Sung Goo PARK ; Joon Young CHOI ; Hong Kwan KIM ; Jhingook KIM ; Sumin SHIN ; Ho Yun LEE
Cancer Research and Treatment 2025;57(1):57-69
Purpose:
This study aimed to develop a magnetic resonance imaging (MRI)–based radiomics model to predict high-risk pathologic features for lung adenocarcinoma: micropapillary and solid pattern (MPsol), spread through air space, and poorly differentiated patterns.
Materials and Methods:
As a prospective study, we screened clinical N0 lung cancer patients who were surgical candidates and had undergone both 18F-fluorodeoxyglucose (FDG) positron emission tomography–computed tomography (PET/CT) and chest CT from August 2018 to January 2020. We recruited patients meeting our proposed imaging criteria indicating high-risk, that is, poorer prognosis of lung adenocarcinoma, using CT and FDG PET/CT. If possible, these patients underwent an MRI examination from which we extracted 77 radiomics features from T1-contrast-enhanced and T2-weighted images. Additionally, patient demographics, maximum standardized uptake value on FDG PET/CT, and the mean apparent diffusion coefficient value on diffusion-weighted image, were considered together to build prediction models for high-risk pathologic features.
Results:
Among 616 patients, 72 patients met the imaging criteria for high-risk lung cancer and underwent lung MRI. The magnetic resonance (MR)–eligible group showed a higher prevalence of nodal upstaging (29.2% vs. 4.2%, p < 0.001), vascular invasion (6.5% vs. 2.1%, p=0.011), high-grade pathologic features (p < 0.001), worse 4-year disease-free survival (p < 0.001) compared with non-MR-eligible group. The prediction power for MR-based radiomics model predicting high-risk pathologic features was good, with mean area under the receiver operating curve (AUC) value measuring 0.751-0.886 in test sets. Adding clinical variables increased the predictive performance for MPsol and the poorly differentiated pattern using the 2021 grading system (AUC, 0.860 and 0.907, respectively).
Conclusion
Our imaging criteria can effectively screen high-risk lung cancer patients and predict high-risk pathologic features by our MR-based prediction model using radiomics.

Result Analysis
Print
Save
E-mail