Risk prediction models of early diagnosis of prostate cancer based on machine learning algorithms
10.12483/j.issn.1009-8291.2026.03.009
- VernacularTitle:基于机器学习算法构建前列腺癌早期诊断的风险预测模型
- Author:
Wuxue LI
1
;
Tianhe ZHANG
1
;
Xinghua ZHAO
1
;
Changbao XU
1
;
Haiyang WEI
1
;
Zixu ZHANG
1
Author Information
1. Department of Urology, Second Affiliated Hospital of Zhengzhou University, Zhengzhou 450000, China
- Publication Type:Journal Article
- Keywords:
prostate cancer;
machine learning;
prostate-specific antigen;
prostate biopsy;
GBM model
- From:
Journal of Modern Urology
2026;31(3):249-257
- CountryChina
- Language:Chinese
-
Abstract:
Objective To construct prostate cancer(PCa)prediction models based on machine learning algorithms, so as to improve the accuracy of early diagnosis of PCa. Methods A retrospective analysis was performed on the clinical data of 504 patients who underwent prostate biopsy at our hospital during Jan. 2020 and Nov. 2024. Patients' age, body mass index(BMI), history of hypertension, diabetes and smoking, total prostate-specific antigen(tPSA), free prostate-specific antigen(fPSA), f/tPSA, prostate volume(PV), neutrophil count, lymphocyte count, neutrophil-to-lymphocyte ratio(NLR), Prostate Imaging Reporting and Data System(PI-RADS)score, digital rectal examination(DRE)results, and pathological findings were collected. The patients were divided into the training and testing sets at a ratio of 7:3. Ten early diagnosis prediction models of PCa were constructed using 10 supervised machine learning algorithms. Model performance was evaluated and validated using metrics including area under the receiver operating characteristic curve(AUC), accuracy, sensitivity, specificity, calibration curves, and decision curve analysis(DCA). SHAP analysis was used to interpret the models, and to clarify the importance of each feature and the basis for model decisions. Results All models showed good predictive value, with the gradient boosting machine(GBM)model performing the best(AUC=0.905, accuracy=84.1%, sensitivity=90.2%, specificity=80.0%). Calibration curves indicated good calibration and fitting of the GBM model, while DCA demonstrated favorable clinical net benefits. SHAP analysis identified the most significant features affecting PCa occurrence in descending order:tPSA, f/tPSA, PIRADS score, PV, age, and DRE. Conclusion The GBM model exhibits the optimal performance among the 10 models. The importance of features for predicting the occurrence of PCa, from the highest to the lowest, is as follows:tPSA, f/tPSA, PIRADS score, PV, age, and DRE.