1.Reassessing the gold standard: The role of AI-powered urinalysis in diagnosing urinary tract infections.
Philippine Journal of Pathology 2026;11(1):37-44
Urinary tract infections (UTIs) are among the most common bacterial infections worldwide, requiring timely and accurate diagnosis to guide appropriate therapy and reduce antimicrobial resistance. Although urine culture remains the diagnostic gold standard, its prolonged turnaround time and susceptibility to pre-analytical variability limit its clinical efficiency. Recent advances in artificial intelligence (AI) have positioned urinalysis as a promising alternative diagnostic approach, utilizing machine learning and deep learning algorithms for automated analysis and prediction. This review synthesizes current evidence on AI applications in urinalysis for UTI diagnosis, examining computational techniques, diagnostic performance, clinical integration, limitations, and future directions. The literature demonstrates that AI-powered urinalysis can achieve diagnostic accuracy comparable to urine culture, with high sensitivity and specificity while reducing diagnostic time. Integration of AI into clinical workflows has the potential to enhance decision-making, streamline laboratory processes, and support antimicrobial stewardship. However, challenges related to data heterogeneity, algorithm interpretability, validation, and regulatory requirements remain significant barriers to widespread adoption. Overall, AI-driven urinalysis represents a transformative opportunity to complement the existing diagnostic standard and advance more rapid, efficient, and personalized approaches to UTI management.
Human ; Artificial Intelligence ; Urinalysis ; Urinary Tract Infections ; Machine Learning ; Deep Learning
2.Identification of high-risk preoperative blood indicators and baseline characteristics for multiple postoperative complications in rheumatoid arthritis patients undergoing total knee arthroplasty: a multi-machine learning feature contribution analysis.
Kejia ZHU ; Zhiyang HUANG ; Biao WANG ; Hang LI ; Yuangang WU ; Bin SHEN ; Yong NIE
Chinese Journal of Reparative and Reconstructive Surgery 2025;39(12):1532-1542
OBJECTIVE:
To explore, identify, and develop novel blood-based indicators using machine learning algorithms for accurate preoperative assessment and effective prediction of postoperative complication risks in patients with rheumatoid arthritis (RA) undergoing total knee arthroplasty (TKA).
METHODS:
A retrospective cohort study was conducted including RA patients who underwent unilateral TKA between January 2019 and December 2024. Inpatient and 30-day postoperative outpatient follow-up data were collected. Six machine learning algorithms, including decision tree, random forest, logistic regression, support vector machine, extreme gradient boosting, and light gradient boosting machine, were used to construct predictive models. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), F1-score, accuracy, precision, and recall. SHapley Additive exPlanations (SHAP) values were employed to interpret and rank the importance of individual variables.
RESULTS:
According to the inclusion criteria, a total of 1 548 patients were enrolled. Ultimately, 18 preoperative indicators were identified as effective predictive features, and 8 postoperative complications were defined as prediction labels for inclusion in the study. Within 30 days after surgery, 453 patients (29.2%) developed one or more complications. Considering overall accuracy, precision, recall, and F1-score, the random forest model [AUC=0.930, 95% CI (0.910, 0.950)] and the extreme gradient boosting model [AUC=0.909, 95% CI (0.880, 0.938)] demonstrated the best predictive performance. SHAP analysis revealed that anti-cyclic citrullinated peptide antibody, C-reactive protein, rheumatoid factor, interleukin-6, body mass index, age, and smoking status made significant contributions to the overall prediction of postoperative complications.
CONCLUSION
Machine learning-based models enable accurate prediction of postoperative complication risks among RA patients undergoing TKA. Inflammatory and immune-related blood biomarkers, such as anti-cyclic citrullinated peptide antibody, C-reactive protein, and rheumatoid factor, interleukin-6, play key predictive roles, highlighting their potential value in perioperative risk stratification and individualized management.
Humans
;
Arthroplasty, Replacement, Knee/adverse effects*
;
Arthritis, Rheumatoid/blood*
;
Machine Learning
;
Postoperative Complications/blood*
;
Female
;
Male
;
Retrospective Studies
;
Middle Aged
;
Aged
;
Risk Factors
;
Preoperative Period
;
C-Reactive Protein/analysis*
;
Risk Assessment
3.Machine learning models established to distinguish OA and RA based on immune factors in the knee joint fluid.
Qin LIANG ; Lingzhi ZHAO ; Yan LU ; Rui ZHANG ; Qiaolin YANG ; Hui FU ; Haiping LIU ; Lei ZHANG ; Guoduo LI
Chinese Journal of Cellular and Molecular Immunology 2025;41(4):331-338
Objective Based on 25 indicators including immune factors, cell count classification, and smear results of the knee joint fluid, machine learning models were established to distinguish between osteoarthritis (OA) and rheumatoid arthritis (RA). Methods 100 OA and 40 RA patients scheduled for total knee arthroplasty were enrolled respectively. Each patient's knee joint fluid was collected preoperatively. Nucleated cells were counted and classified. The expression levels of immune factors, including tumor necrosis factor alpha (TNF-α), interleukin-1 beta (IL-1β), IL-6, IL-8, IL-15, matrix metalloproteinase 3 (MMP3), MMP9, MMP13, rheumatoid factor (RF), serum amyloid A (SAA), C-reactive protein (CRP), and others were measured. Smears and microscopic classification of all the immune factors were performed. Independent influencing factors for OA or RA were identified using univariate binary logistic regression, Lasso regression, and multivariate binary logistic regression. Based on the independent influencing factors, three machine learning models were constructed which are logistic regression, random forest, and support vector machine. Receiver operating characteristic curve (ROC), calibration curve and decision curve analysis (DCA) were used to evaluate and compare the models. Results A total of 5 indicators in the knee joint fluid were screened out to distinguish OA and RA, which were IL-1β(odds ratio(OR)=10.512, 95× confidence interval (95×CI) was 1.048-105.42, P=0.045), IL-6 (OR=1.007, 95×CI was 1.001-1.014, P=0.022), MMP9 (OR=3.202, 95×CI was 1.235-8.305, P=0.017), MMP13 (OR=1.002, 95× CI was 1-1.004, P=0.049), and RF (OR=1.091, 95×CI was 1.01-1.179, P=0.026). According to the results of ROC, calibration curve and DCA, the accuracy (0.979), sensitivity (0.98) and area under the curve (AUC, 0.996, 95×CI was 0.991-1) of the random forest model were the highest. It has good validity and feasibility, and its distinguishing ability is better than the other two models. Conclusion The machine learning model based on immune factors in the knee joint fluid holds significant value in distinguishing OA and RA. It provides an important reference for the clinical early differential diagnosis, prevention and treatment of OA and RA.
Humans
;
Arthritis, Rheumatoid/metabolism*
;
Machine Learning
;
Male
;
Female
;
Middle Aged
;
Aged
;
Synovial Fluid/immunology*
;
Osteoarthritis, Knee/metabolism*
;
Knee Joint/metabolism*
;
ROC Curve
;
Diagnosis, Differential
4.Value of biomarkers related to routine blood tests in early diagnosis of allergic rhinitis in children.
Jinjie LI ; Xiaoyan HAO ; Yijuan XIN ; Rui LI ; Lin ZHU ; Xiaoli CHENG ; Liu YANG ; Jiayun LIU
Chinese Journal of Cellular and Molecular Immunology 2025;41(4):339-347
Objective To mine and analyze the routine blood test data of children with allergic rhinitis (AR), identify routine blood parameters related to childhood allergic rhinitis, establish an effective diagnostic model, and evaluate the performance of the model. Methods This study was a retrospective study of clinical cases. The experimental group comprised a total of 1110 children diagnosed with AR at the First Affiliated Hospital of Air Force Medical University during the period from December 12, 2020 to December 12, 2021, while the control group included 1109 children without a history of allergic rhinitis or other allergic diseases who underwent routine physical examinations during the same period. Information such as age, sex and routine blood test results was collected for all subjects. The levels of routine blood test indicators were compared between AR children and healthy children using comprehensive intelligent baseline analysis, with indicators of P≥0.05 excluded; variables were screened by Lasso regression. Binary Logistic regression was used to further evaluate the influence of multiple routine blood indexes on the results. Five kinds of machine model algorithms were used, namely extreme value gradient lift (XGBoost), logistic regression (LR), gradient lift decision tree (LGBMC), Random forest (RF) and adaptive lift algorithm (AdaBoost), to establish the diagnostic models. The receiver operating characteristic (ROC) curve was used to screen the optimal model. The best LightGBM algorithm was used to build an online patient risk assessment tool for clinical application. Results Statistically significant differences were observed between the AR group and the control group in the following routine blood test indicators: mean cellular hemoglobin concentration (MCHC), hemoglobin (HGB), absolute value of basophils (BASO), absolute value of eosinophils (EOS), large platelet ratio (P-LCR), mean platelet volume (MPV), platelet distribution width (PDW), platelet count (PLT), absolute values of leukocyte neutrophil (W-LCC), leukocyte monocyte (W-MCC), leukocyte lymphocyte (W-SCC), and age. Lasso regression identified these variables as important predictors, and binary Logistic regression further analyzed the significant influence of these variables on the results. The optimal machine learning algorithm LightGBM was used to establish a multi-index joint detection model. The model showed robust prediction performance in the training set, with AUC values of 0.8512 and 0.8103 in the internal validation set. Conclusion The identified routine blood parameters can be used as potential biomarkers for early diagnosis and risk assessment of AR, which can improve the accuracy and efficiency of diagnosis. The established model provides scientific basis for more accurate diagnostic tools and personalized prevention strategies. Future studies should prospectively validate these findings and explore their applicability in other related diseases.
Humans
;
Male
;
Female
;
Rhinitis, Allergic/blood*
;
Child
;
Biomarkers/blood*
;
Retrospective Studies
;
Early Diagnosis
;
Child, Preschool
;
ROC Curve
;
Logistic Models
;
Hematologic Tests
;
Algorithms
;
Adolescent
;
Machine Learning
5.Explainable machine learning model for predicting septic shock in critically sepsis patients based on coagulation indexes: A multicenter cohort study.
Qing-Bo ZENG ; En-Lan PENG ; Ye ZHOU ; Qing-Wei LIN ; Lin-Cui ZHONG ; Long-Ping HE ; Nian-Qing ZHANG ; Jing-Chun SONG
Chinese Journal of Traumatology 2025;28(6):404-411
PURPOSE:
Septic shock is associated with high mortality and poor outcomes among sepsis patients with coagulopathy. Although traditional statistical methods or machine learning (ML) algorithms have been proposed to predict septic shock, these potential approaches have never been systematically compared. The present work aimed to develop and compare models to predict septic shock among patients with sepsis.
METHODS:
It is a retrospective cohort study based on 484 patients with sepsis who were admitted to our intensive care units between May 2018 and November 2022. Patients from the 908th Hospital of Chinese PLA Logistical Support Force and Nanchang Hongdu Hospital of Traditional Chinese Medicine were respectively allocated to training (n=311) and validation (n=173) sets. All clinical and laboratory data of sepsis patients characterized by comprehensive coagulation indexes were collected. We developed 5 models based on ML algorithms and 1 model based on a traditional statistical method to predict septic shock in the training cohort. The performance of all models was assessed using the area under the receiver operating characteristic curve and calibration plots. Decision curve analysis was used to evaluate the net benefit of the models. The validation set was applied to verify the predictive accuracy of the models. This study also used Shapley additive explanations method to assess variable importance and explain the prediction made by a ML algorithm.
RESULTS:
Among all patients, 37.2% experienced septic shock. The characteristic curves of the 6 models ranged from 0.833 to 0.962 and 0.630 to 0.744 in the training and validation sets, respectively. The model with the best prediction performance was based on the support vector machine (SVM) algorithm, which was constructed by age, tissue plasminogen activator-inhibitor complex, prothrombin time, international normalized ratio, white blood cells, and platelet counts. The SVM model showed good calibration and discrimination and a greater net benefit in decision curve analysis.
CONCLUSION
The SVM algorithm may be superior to other ML and traditional statistical algorithms for predicting septic shock. Physicians can better understand the reliability of the predictive model by Shapley additive explanations value analysis.
Humans
;
Shock, Septic/blood*
;
Machine Learning
;
Male
;
Female
;
Retrospective Studies
;
Middle Aged
;
Aged
;
Sepsis/complications*
;
ROC Curve
;
Cohort Studies
;
Adult
;
Intensive Care Units
;
Algorithms
;
Blood Coagulation
;
Critical Illness
6.Early prediction and warning of MODS following major trauma via identification of cytokine storm: A prospective cohort study.
Panpan CHANG ; Rui LI ; Jiahe WEN ; Guanjun LIU ; Feifei JIN ; Yongpei YU ; Yongzheng LI ; Guang ZHANG ; Tianbing WANG
Chinese Journal of Traumatology 2025;28(6):391-398
PURPOSE:
Early mortality in major trauma has decreased, but MODS remains a leading cause of poor outcomes, driven by trauma-induced cytokine storms that exacerbate injuries and organ damage.
METHODS:
This prospective cohort study included 79 major trauma patients (ISS >15) treated in the National Center for Trauma Medicine, Peking University People's Hospital, from September 1, 2021, to July 31, 2023. Patients (1) with ISS >15 (according to AIS 2015), (2) aged 15-80 years, (3) admitted within 6 h of injury, (4) having no prior treatment before admission, were included. Exclusion criteria were (1) GCS score <9 or AIS score ≥3 for TBI, (2) confirmed infection, infectious disease, or high infection risk, (3) pregnancy, (4) severe primary diseases affecting survival, (5) recent use of immunosuppressive or cytotoxic drugs within the past 6 months, (6) psychiatric patients, (7) participation in other clinical trials within the past 30 days, (8) patients with incomplete data or missing blood samples. Admission serum inflammatory cytokines and pathophysiological data were analyzed to develop machine learning models predicting MODS within 7 days. LR, DR, RF, SVM, NB, and XGBoost were evaluated based on the area under the AUROC. The SHAP method was used to interpret results.
RESULTS:
This study enrolled 79 patients with major trauma, and the median (Q1, Q3) age was 51 (35, 59) years (52 males, 65.8%). The inflammatory cytokine data were collected for all participants. Among these patients, 35 (44.3%) developed MODS, and 44 (55.7%) did not. Additionally, 2 patients (2.5%) from the MODS group succumbed. The logistic regression model showed strong performance in predicting MODS. Ten key cytokines, IL-18, Eotaxin, MCP-4, IP-10, CXCL12, MIP-3α, MCP-1, IL-1RA, Cystatin C, and MRP8/14 were identified as critical to the trauma-induced cytokine storm and MODS development. Early elevation of these cytokines achieved high predictive accuracy, with an AUROC of 0.887 (95% CI 0.813-0.976).
CONCLUSION
Trauma-induced cytokine storms are strongly associated with MODS. Early identification of inflammatory cytokine changes enables better prediction and timely interventions to improve outcomes.
Humans
;
Prospective Studies
;
Middle Aged
;
Male
;
Female
;
Adult
;
Aged
;
Cytokine Release Syndrome/etiology*
;
Adolescent
;
Young Adult
;
Aged, 80 and over
;
Wounds and Injuries/complications*
;
Cytokines/blood*
;
Multiple Organ Failure/diagnosis*
;
Machine Learning
7.Predictability of varicocele repair success: preliminary results of a machine learning-based approach.
Andrea CRAFA ; Marco RUSSO ; Rossella CANNARELLA ; Murat GÜL ; Michele COMPAGNONE ; Laura M MONGIOÌ ; Vittorio CANNARELLA ; Rosita A CONDORELLI ; Sandro La VIGNERA ; Aldo E CALOGERO
Asian Journal of Andrology 2025;27(1):52-58
Varicocele is a prevalent condition in the infertile male population. However, to date, which patients may benefit most from varicocele repair is still a matter of debate. The purpose of this study was to evaluate whether certain preintervention sperm parameters are predictive of successful varicocele repair, defined as an improvement in total motile sperm count (TMSC). We performed a retrospective study on 111 patients with varicocele who had undergone varicocele repair, collected from the Department of Endocrinology, Metabolic Diseases and Nutrition, University of Catania (Catania, Italy), and the Unit of Urology at the Selcuk University School of Medicine (Konya, Türkiye). The predictive analysis was conducted through the use of the Brain Project, an innovative tool that allows a complete and totally unbiased search of mathematical expressions that relate the object of study to the various parameters available. Varicocele repair was considered successful when TMSC increased by at least 50% of the preintervention value. For patients with preintervention TMSC below 5 × 10 6 , improvement was considered clinically relevant when the increase exceeded 50% and the absolute TMSC value was >5 × 10 6 . From the preintervention TMSC alone, we found a model that predicts patients who appear to benefit little from varicocele repair with a sensitivity of 50.0% and a specificity of 81.8%. Varicocele grade and serum follicle-stimulating hormone (FSH) levels did not play a predictive role, but it should be noted that all patients enrolled in this study were selected with intermediate- or high-grade varicocele and normal FSH levels. In conclusion, preintervention TMSC is predictive of the success of varicocele repair in terms of TMSC improvement in patients with intermediate- or high-grade varicoceles and normal FSH levels.
Humans
;
Varicocele/complications*
;
Male
;
Retrospective Studies
;
Machine Learning
;
Adult
;
Treatment Outcome
;
Sperm Count
;
Infertility, Male/etiology*
;
Sperm Motility
;
Follicle Stimulating Hormone/blood*
;
Young Adult
8.The application of machine learning in the auxiliary diagnosis of specific learning disorder.
Hao ZHAO ; Shu-Lan MEI ; Jing-Yu WANG ; Xia CHI
Chinese Journal of Contemporary Pediatrics 2025;27(11):1420-1425
Specific learning disorder (SLD) is a common neurodevelopmental disorder in children that significantly affects academic performance and quality of life. At present, diagnosis mainly relies on standardized tests and professional evaluations, a process that is complex and time-consuming. Multiple studies have shown that machine learning can analyze diverse data, including test scores, handwriting samples, eye movement data, neuroimaging data, and genetic data, to automatically learn the relationships between input features and output labels and achieve efficient prediction. It shows great potential for early screening, auxiliary diagnosis, and research on underlying mechanisms in SLD. This article reviews the applications of machine learning in the auxiliary diagnosis of SLD and discusses its performance when handling different data types.
Humans
;
Machine Learning
;
Specific Learning Disorder/diagnosis*
;
Child
9.Exploration of the Predictive Value of Peripheral Blood-related Indicators for EGFR Mutations and Prognosis in Non-small Cell Lung Cancer Using Machine Learning.
Shulei FU ; Shaodi WEN ; Jiaqiang ZHANG ; Xiaoyue DU ; Ru LI ; Bo SHEN
Chinese Journal of Lung Cancer 2025;28(2):105-113
BACKGROUND:
Epidermal growth factor receptor (EGFR) sensitive mutation is one of the effective targets of targeted therapy for non-small cell lung cancer (NSCLC). However, due to the difficulty of obtaining some primary tissues and the economic factors in some underdeveloped areas, some patients cannot undergo traditional genetic testing. The aim of this study is to establish a machine learning (ML) model using non-invasive peripheral blood markers to explore the biomarkers closely related to EGFR mutation status in NSCLC and evaluate their potential prognostic value.
METHODS:
2642 lung cancer patients who visited Jiangsu Cancer Hospital from November 2016 to May 2023 were retrospectively enrolled and finally 175 NSCLC patients with complete follow-up data were included in the study. The ML model was constructed based on peripheral blood indicators and divided into training set and test set according to the ratio of 8:2. Unsupervised learning algorithms were used for clustering blood features and mutual information method for feature selection, and an ensemble learning algorithm based on Shapley value was designed to calculate the contribution of each feature to the model prediction result. The receiver operating characteristic (ROC) curve was used to evaluate the predictive ability of the model.
RESULTS:
Through the feature extraction and contribution analysis of the predictive results of the interpretable ML model based on the Shapley value, the top ten indicators with the highest contribution were: pathological type, phosphorus, eosinophils, monocyte count, activated partial thromboplastin time, potassium, total bilirubin, sodium, eosinophil percentage, and total cholesterol. The area under the curve (AUC) of the model was 0.80. In addition, patients with hyponatremia and squamous cell carcinoma group had a poor prognosis (P<0.05).
CONCLUSIONS
The interpretable model constructed in this study provides a new approach for the prediction of EGFR mutation status in NSCLC patients, which provides a scientific basis for the diagnosis and treatment of patients who cannot undergo genetic testing.
Humans
;
Carcinoma, Non-Small-Cell Lung/diagnosis*
;
Machine Learning
;
Lung Neoplasms/diagnosis*
;
Male
;
Female
;
Mutation
;
Middle Aged
;
ErbB Receptors/genetics*
;
Prognosis
;
Aged
;
Retrospective Studies
;
Adult
;
Biomarkers, Tumor/genetics*
10.Application of machine learning algorithms in predicting new onset hypertension: a study based on the China Health and Nutrition Survey.
Manhui ZHANG ; Xian XIA ; Qiqi WANG ; Yue PAN ; Guanyi ZHANG ; Zhigang WANG
Environmental Health and Preventive Medicine 2025;30():3-3
BACKGROUND:
Hypertension is a serious chronic disease that can significantly lead to various cardiovascular diseases, affecting vital organs such as the heart, brain, and kidneys. Our goal is to predict the risk of new onset hypertension using machine learning algorithms and identify the characteristics of patients with new onset hypertension.
METHODS:
We analyzed data from the 2011 China Health and Nutrition Survey cohort of individuals who were not hypertensive at baseline and had follow-up results available for prediction by 2015. We tested and evaluated the performance of four traditional machine learning algorithms commonly used in epidemiological studies: Logistic Regression, Support Vector Machine, XGBoost, LightGBM, and two deep learning algorithms: TabNet and AMFormer model. We modeled using 16 and 29 features, respectively. SHAP values were applied to select key features associated with new onset hypertension.
RESULTS:
A total of 4,982 participants were included in the analysis, of whom 1,017 developed hypertension during the 4-year follow-up. Among the 16-feature models, Logistic Regression had the highest AUC of 0.784(0.775∼0.806). In the 29-feature prediction models, AMFormer performed the best with an AUC of 0.802(0.795∼0.820), and also scored the highest in MCC (0.417, 95%CI: 0.400∼0.434) and F1 (0.503, 95%CI: 0.484∼0.505) metrics, demonstrating superior overall performance compared to the other models. Additionally, key features selected based on the AMFormer, such as age, province, waist circumference, urban or rural location, education level, employment status, weight, WHR, and BMI, played significant roles.
CONCLUSION
We used the AMFormer model for the first time in predicting new onset hypertension and achieved the best results among the six algorithms tested. Key features associated with new onset hypertension can be determined through this algorithm. The practice of machine learning algorithms can further enhance the predictive efficacy of diseases and identify risk factors for diseases.
Humans
;
China/epidemiology*
;
Hypertension/diagnosis*
;
Machine Learning
;
Male
;
Female
;
Middle Aged
;
Adult
;
Nutrition Surveys
;
Algorithms
;
Aged
;
Risk Factors


Result Analysis
Print
Save
E-mail