Construction and validation of a machine learning-based risk assessment model for post-transplant diabetes mellitus
- VernacularTitle:基于机器学习的肝移植术后糖尿病风险评估模型的构建与验证
- Author:
Hao WANG
1
;
Yongqiang FAN
1
;
Zhiyong SHI
1
;
Rui ZHANG
1
;
Yan WANG
1
;
Jun XU
1
Author Information
- Publication Type:Journal Article
- Keywords: Liver Transplantation; Diabetes Mellitus; Machine Learning; Models, Statistical
- From: Journal of Clinical Hepatology 2026;42(8):1894-1901
- CountryChina
- Language:Chinese
- Abstract: ObjectiveTo construct and validate a risk assessment model for post-transplant diabetes mellitus (PTDM) using multiple machine learning algorithms, and to realize the early identification of PTDM. MethodsA retrospective analysis was performed for the clinical data of the patients who underwent allogeneic liver transplantation in The First Hospital of Shanxi Medical University from April 1, 2020 to December 31, 2024, and they were randomly divided into a training set and a validation set at a ratio of 7∶3. The LASSO regression analysis combined with 5-fold cross-validation was used for feature selection. Seven machine learning models were developed in the training set, i.e., logistic regression (LR), decision tree (DT), Naive Bayes (NB), random forest (RF), K-nearest neighbor (KNN), extreme gradient boosting (XGBoost), and adaptive boosting (AdaBoost). In the validation set, various methods were used to assess the predictive performance of each model, such as accuracy, precision, recall rate, specificity, F1 score, area under the receiver operating characteristic curve (AUROC), and area under the precision-recall curve (AUPRC). The Brier score and decision curve analysis were used to assess the calibration and clinical practicability of the models, and the SHAP method was used to analyze feature importance. The independent-samples t test or the Wilcoxon rank-sum test was used for comparison of continuous data between two groups, and the chi-square test or the Fisher’s exact test was used for comparison of categorical data between two groups. ResultsA total of 135 liver transplant recipients were enrolled, among whom 26 developed PTDM, and there were 94 patients in the training set and 41 in the validation set. Feature extraction and screening identified 7 key features of sex, overweight or obesity, anhepatic phase, time of operation, length of hospital stay, early postoperative hypomagnesemia, and impaired fasting glucose (IFG). In the validation set, the XGBoost model showed the best predictive performance, with an AUROC of 0.907 (95% confidence interval [CI]: 0.807 — 0.989), an AUPRC of 0.649 (95%CI: 0.339 — 0.955), an accuracy of 0.878, a precision of 0.667, a recall rate of 0.750, an F1-score of 0.706, a specificity of 0.909, and a Brier score of 0.104. The decision curve analysis showed that when the threshold probability was below 0.667, application of the XGBoost model in clinical decision-making provided relatively high net benefit. The SHAP analysis showed that the length of hospital stay ranked first in terms of feature importance, followed by time of operation, overweight or obesity, sex, IFG, anhepatic phase, and early postoperative hypomagnesemia. ConclusionThe risk assessment model for PTDM in liver transplant recipients based on XGBoost algorithm has excellent performance and can effectively identify high-risk individuals; however, multicenter large-sample data are needed for further validation.
