1.Reassessing the gold standard: The role of AI-powered urinalysis in diagnosing urinary tract infections.
Philippine Journal of Pathology 2026;11(1):37-44
Urinary tract infections (UTIs) are among the most common bacterial infections worldwide, requiring timely and accurate diagnosis to guide appropriate therapy and reduce antimicrobial resistance. Although urine culture remains the diagnostic gold standard, its prolonged turnaround time and susceptibility to pre-analytical variability limit its clinical efficiency. Recent advances in artificial intelligence (AI) have positioned urinalysis as a promising alternative diagnostic approach, utilizing machine learning and deep learning algorithms for automated analysis and prediction. This review synthesizes current evidence on AI applications in urinalysis for UTI diagnosis, examining computational techniques, diagnostic performance, clinical integration, limitations, and future directions. The literature demonstrates that AI-powered urinalysis can achieve diagnostic accuracy comparable to urine culture, with high sensitivity and specificity while reducing diagnostic time. Integration of AI into clinical workflows has the potential to enhance decision-making, streamline laboratory processes, and support antimicrobial stewardship. However, challenges related to data heterogeneity, algorithm interpretability, validation, and regulatory requirements remain significant barriers to widespread adoption. Overall, AI-driven urinalysis represents a transformative opportunity to complement the existing diagnostic standard and advance more rapid, efficient, and personalized approaches to UTI management.
Human ; Artificial Intelligence ; Urinalysis ; Urinary Tract Infections ; Machine Learning ; Deep Learning
2.Artificial intelligence-assisted design, mining, and modification of CRISPR-Cas systems.
Yufeng MAO ; Guangyun CHU ; Qingling LIANG ; Ye LIU ; Yi YANG ; Xiaoping LIAO ; Meng WANG
Chinese Journal of Biotechnology 2025;41(3):949-967
With the rapid advancement of synthetic biology, CRISPR-Cas systems have emerged as a powerful tool for gene editing, demonstrating significant potential in various fields, including medicine, agriculture, and industrial biotechnology. This review comprehensively summarizes the significant progress in applying artificial intelligence (AI) technologies to the design, mining, and modification of CRISPR-Cas systems. AI technologies, especially machine learning, have revolutionized sgRNA design by analyzing high-throughput sequencing data, thereby improving the editing efficiency and predicting off-target effects with high accuracy. Furthermore, this paper explores the role of AI in sgRNA design and evaluation, highlighting its contributions to the annotation and mining of CRISPR arrays and Cas proteins, as well as its potential for modifying key proteins involved in gene editing. These advancements have not only improved the efficiency and precision of gene editing but also expanded the horizons of genome engineering, paving the way for intelligent and precise genome editing.
CRISPR-Cas Systems/genetics*
;
Artificial Intelligence
;
Gene Editing/methods*
;
RNA, Guide, CRISPR-Cas Systems/genetics*
;
Machine Learning
;
Humans
;
Genetic Engineering/methods*
;
Synthetic Biology
3.Intelligent mining, engineering, and de novo design of proteins.
Cui LIU ; Zhenkun SHI ; Hongwu MA ; Xiaoping LIAO
Chinese Journal of Biotechnology 2025;41(3):993-1010
Natural components serve the survival instincts of cells that are obtained through long-term evolution, while they often fail to meet the demands of engineered cells for efficiently performing biological functions in special industrial environments. Enzymes, as biological catalysts, play a key role in biosynthetic pathways, significantly enhancing the rate and selectivity of biochemical reactions. However, the catalytic efficiency, stability, substrate specificity, and tolerance of natural enzymes often fall short of industrial production requirements. Therefore, exploring and modifying enzymes to suit specific biomanufacturing processes has become crucial. In recent years, artificial intelligence (AI) has played an increasingly important role in the discovery, evaluation, engineering, and de novo design of proteins. AI can accelerate the discovery and optimization of proteins by analyzing large amounts of bioinformatics data and predicting protein functions and characteristics by machine learning and deep learning algorithms. Moreover, AI can assist researchers in designing new protein structures by simulating and predicting their performance under different conditions, providing guidance for protein design. This paper reviews the latest research advances in protein discovery, evaluation, engineering, and de novo design for biomanufacturing and explores the hot topics, challenges, and emerging technical methods in this field, aiming to provide guidance and inspiration for researchers in related fields.
Protein Engineering/methods*
;
Artificial Intelligence
;
Proteins/genetics*
;
Computational Biology
;
Machine Learning
;
Data Mining
;
Algorithms
;
Deep Learning
4.Intelligent design of transcription factor-based biosensors.
Chaoning LIANG ; La XIANG ; Shuangyan TANG
Chinese Journal of Biotechnology 2025;41(3):1011-1022
Transcription factor (TF)-based biosensors have been widely applied in metabolic engineering, synthetic biology, metabolites monitoring, etc. These biosensors are praised for the high orthogonality, modularity, and operability. However, most natural TFs with weak responses and low specificity still demand optimization for desired performance in applications. Herein, we comprehensively summarize the recent advances in the engineering and optimization of TF-based biosensors with the assistance of computational simulation and artificial intelligence. This review includes the regulatory protein engineering aided by protein structure prediction and ligand binding simulation and the regulatory protein responses predicted by a mathematical model obtained from machine learning of mutagenesis data. In comparison with conventional tools, computational simulation and artificial intelligence enable more accurate and rapid design and construction of biosensors. Thus, these technologies will greatly promote the development of novel biosensors for applications.
Biosensing Techniques/methods*
;
Transcription Factors/metabolism*
;
Artificial Intelligence
;
Protein Engineering/methods*
;
Computer Simulation
;
Synthetic Biology
;
Machine Learning
5.Machine learning-aided design of synthetic biological parts and circuits.
Chinese Journal of Biotechnology 2025;41(3):1023-1051
Synthetic biology is an emerging interdisciplinary field at the convergence of biology, engineering, and computer science. It employs a bottom-up approach to progressively design biological parts, devices, and circuits, aiming to create artificial biological systems not found in nature or to redesign existing biological systems for specific purposes. With the rapid development of the synthetic biology industry, there is an increasing demand for large complex genetic circuits. However, the traditional trial-and-error methods, heavily reliant on empirical knowledge, have limited efficiency and success rates of parts/circuits construction, thereby impeding the innovation and technology translation for synthetic biology. These limitations have prompted a paradigm shift from labor-intensive, experience-driven trial-and-error models towards standardized, intelligent engineering approaches. Machine learning, capable of uncovering hidden structures and relationships within biological data, offers robust support for the intelligent design of synthetic biological parts and genetic circuits. Here, we review commonly used machine learning algorithms and analyze their typical applications in designing biological parts (e.g., synthetic promoters, RNA regulatory elements, and transcription factors) and simple genetic circuits. Additionally, we discuss the primary challenges in machine learning-aided design and propose potential solutions. Lastly, we envision the future trend of integrating machine learning with synthetic biological system design, highlighting the importance of interdisciplinary collaboration.
Synthetic Biology/methods*
;
Machine Learning
;
Gene Regulatory Networks
;
Algorithms
6.Serum proteomics and machine learning unveil new diagnostic biomarkers for tuberculosis in adolescents and young adults.
Yu CHEN ; Hongxiang XU ; Yao TIAN ; Qian HE ; Xiaoyun ZHAO ; Guobin ZHANG ; Jianping XIE
Chinese Journal of Biotechnology 2025;41(4):1478-1489
Adolescents and young adults (AYAs) are one of the major populations susceptible to tuberculosis. However, little is known about the unique characteristics and diagnostic biomarkers of tuberculosis in this population. In this study, 81 AYAs were recruited, and the high-quality serum proteome of the AYAs with tuberculosis was profiled by quantitative proteomics. The data of serum proteomics indicated that the relative abundance of hemoglobin and apolipoprotein was significantly reduced in the patients with active tuberculosis (ATB). The pathway enrichment analysis showed that the downregulated proteins in the ATB group were mainly involved in the antioxidant and cell detoxification pathways, indicating extensive oxidative stress damage. Random forest (RF) and extreme gradient boosting (XGBoost) were employed to evaluate protein importance, which yielded a set of candidate proteins that can distinguish between ATB and non-ATB. The analysis with the support vector machine algorithm (recursive feature elimination) suggested that the combination of apolipoprotein A-I (APOA1), hemoglobin subunit beta (HBB), and hemoglobin subunit alpha-1 (HBA1) had the highest accuracy and sensitivity in diagnosing ATB. Meanwhile, the levels of hemoglobin (HGB) and albumin (ALB) can be used as blood biochemical indicators to evaluate changes in the protein levels of APOA1 and HBB. This study established the serum proteome landscape of AYAs with tuberculosis and identified new biomarkers for the diagnosis of tuberculosis in this population.
Humans
;
Proteomics/methods*
;
Biomarkers/blood*
;
Adolescent
;
Young Adult
;
Apolipoprotein A-I/blood*
;
Machine Learning
;
Tuberculosis/blood*
;
Proteome/analysis*
;
Male
;
Hemoglobins/analysis*
;
Female
;
Blood Proteins/analysis*
;
Adult
7.pLM4ACP: a model for predicting anticancer peptides based on machine learning and protein language models.
Yitong LIU ; Wenxin CHEN ; Juanjuan LI ; Xue CHI ; Xiang MA ; Yanqiong TANG ; Hong LI
Chinese Journal of Biotechnology 2025;41(8):3252-3261
Cancer is a serious global health problem and a major cause of human death. Conventional cancer treatments often run the risk of impairing vital organ functions. Anticancer peptides (ACPs) are considered to be one of the most promising therapeutic agents against common human cancers due to their small sizes, high specificity, and low toxicity. Since ACP recognition is highly limited to the laboratory, expensive, and time-consuming, we proposed pLM4ACP, a model for predicting ACPs based on machine learning and protein language models. In this model, the protein language model ProtT5 was used to extract the features of ACPs, and the extracted features were input into the support vector machine (SVM) classification algorithm for optimization and performance evaluation. The model showcased significantly higher accuracy than other methods, with the overall accuracy of 0.763, F1-score of 0.767, Matthews correlation coefficient of 0.527, and area under the curve of 0.827 on the independent test set. This study constructs an efficient anticancer peptide prediction model based on protein language models, further advancing the application of artificial intelligence in the biomedical field and promoting the development of precision medicine and computational biology.
Machine Learning
;
Antineoplastic Agents/chemistry*
;
Humans
;
Peptides/chemistry*
;
Support Vector Machine
;
Algorithms
;
Computational Biology/methods*
;
Neoplasms/drug therapy*
8.Exploration of the Predictive Value of Peripheral Blood-related Indicators for EGFR Mutations and Prognosis in Non-small Cell Lung Cancer Using Machine Learning.
Shulei FU ; Shaodi WEN ; Jiaqiang ZHANG ; Xiaoyue DU ; Ru LI ; Bo SHEN
Chinese Journal of Lung Cancer 2025;28(2):105-113
BACKGROUND:
Epidermal growth factor receptor (EGFR) sensitive mutation is one of the effective targets of targeted therapy for non-small cell lung cancer (NSCLC). However, due to the difficulty of obtaining some primary tissues and the economic factors in some underdeveloped areas, some patients cannot undergo traditional genetic testing. The aim of this study is to establish a machine learning (ML) model using non-invasive peripheral blood markers to explore the biomarkers closely related to EGFR mutation status in NSCLC and evaluate their potential prognostic value.
METHODS:
2642 lung cancer patients who visited Jiangsu Cancer Hospital from November 2016 to May 2023 were retrospectively enrolled and finally 175 NSCLC patients with complete follow-up data were included in the study. The ML model was constructed based on peripheral blood indicators and divided into training set and test set according to the ratio of 8:2. Unsupervised learning algorithms were used for clustering blood features and mutual information method for feature selection, and an ensemble learning algorithm based on Shapley value was designed to calculate the contribution of each feature to the model prediction result. The receiver operating characteristic (ROC) curve was used to evaluate the predictive ability of the model.
RESULTS:
Through the feature extraction and contribution analysis of the predictive results of the interpretable ML model based on the Shapley value, the top ten indicators with the highest contribution were: pathological type, phosphorus, eosinophils, monocyte count, activated partial thromboplastin time, potassium, total bilirubin, sodium, eosinophil percentage, and total cholesterol. The area under the curve (AUC) of the model was 0.80. In addition, patients with hyponatremia and squamous cell carcinoma group had a poor prognosis (P<0.05).
CONCLUSIONS
The interpretable model constructed in this study provides a new approach for the prediction of EGFR mutation status in NSCLC patients, which provides a scientific basis for the diagnosis and treatment of patients who cannot undergo genetic testing.
Humans
;
Carcinoma, Non-Small-Cell Lung/diagnosis*
;
Machine Learning
;
Lung Neoplasms/diagnosis*
;
Male
;
Female
;
Mutation
;
Middle Aged
;
ErbB Receptors/genetics*
;
Prognosis
;
Aged
;
Retrospective Studies
;
Adult
;
Biomarkers, Tumor/genetics*
9.Application of machine learning algorithms in predicting new onset hypertension: a study based on the China Health and Nutrition Survey.
Manhui ZHANG ; Xian XIA ; Qiqi WANG ; Yue PAN ; Guanyi ZHANG ; Zhigang WANG
Environmental Health and Preventive Medicine 2025;30():3-3
BACKGROUND:
Hypertension is a serious chronic disease that can significantly lead to various cardiovascular diseases, affecting vital organs such as the heart, brain, and kidneys. Our goal is to predict the risk of new onset hypertension using machine learning algorithms and identify the characteristics of patients with new onset hypertension.
METHODS:
We analyzed data from the 2011 China Health and Nutrition Survey cohort of individuals who were not hypertensive at baseline and had follow-up results available for prediction by 2015. We tested and evaluated the performance of four traditional machine learning algorithms commonly used in epidemiological studies: Logistic Regression, Support Vector Machine, XGBoost, LightGBM, and two deep learning algorithms: TabNet and AMFormer model. We modeled using 16 and 29 features, respectively. SHAP values were applied to select key features associated with new onset hypertension.
RESULTS:
A total of 4,982 participants were included in the analysis, of whom 1,017 developed hypertension during the 4-year follow-up. Among the 16-feature models, Logistic Regression had the highest AUC of 0.784(0.775∼0.806). In the 29-feature prediction models, AMFormer performed the best with an AUC of 0.802(0.795∼0.820), and also scored the highest in MCC (0.417, 95%CI: 0.400∼0.434) and F1 (0.503, 95%CI: 0.484∼0.505) metrics, demonstrating superior overall performance compared to the other models. Additionally, key features selected based on the AMFormer, such as age, province, waist circumference, urban or rural location, education level, employment status, weight, WHR, and BMI, played significant roles.
CONCLUSION
We used the AMFormer model for the first time in predicting new onset hypertension and achieved the best results among the six algorithms tested. Key features associated with new onset hypertension can be determined through this algorithm. The practice of machine learning algorithms can further enhance the predictive efficacy of diseases and identify risk factors for diseases.
Humans
;
China/epidemiology*
;
Hypertension/diagnosis*
;
Machine Learning
;
Male
;
Female
;
Middle Aged
;
Adult
;
Nutrition Surveys
;
Algorithms
;
Aged
;
Risk Factors
10.Artificial intelligence in stomatology: Innovations in clinical practice, research, education, and healthcare management.
Xuliang DENG ; Mingming XU ; Chenlin DU
Journal of Peking University(Health Sciences) 2025;57(5):821-826
In recent years, China has continued to face a high prevalence of oral diseases, along with uneven access to high-quality dental care. Against this backdrop, artificial intelligence (AI), as a data-driven, algorithm-supported, and model-centered technology system, has rapidly expanded its role in transforming the landscape of stomatology. This review summarizes recent advances in the application of AI in stomatology across clinical care, biomedical and materials research, education, and hospital management. In clinical settings, AI has improved diagnostic accuracy, streamlined treatment planning, and enhanced surgical precision and efficiency. In research, machine learning has accelerated the identification of disease biomarkers, deepened insights into the oral microbiome, and supported the development of novel biomaterials. In education, AI has enabled the construction of knowledge graphs, facilitated personalized learning, and powered simulation-based training, driving innovation in teaching methodologies. Meanwhile, in hospital operations, intelligent agents based on large language models (LLMs) have been widely deployed for intelligent triage, structured pre-consultations, automated clinical documentation, and quality control, contributing to more standardized and efficient healthcare delivery. Building on these foundations, a multi-agent collaborative framework centered around an AI assistant for stomatology is gradually emerging, integrating task-specific agents for imaging, treatment planning, surgical navigation, follow-up prediction, patient communication, and administrative coordination. Through shared interfaces and unified knowledge systems, these agents support seamless human-AI collaboration across the full continuum of care. Despite these achievements, the broader deployment of AI still faces challenges including data privacy, model robustness, cross-institution generalization, and interpretability. Addressing these issues will require the development of federated learning frameworks, multi-center validation, causal reasoning approaches, and strong ethical governance. With these foundations in place, AI is poised to move from a supportive tool to a trusted partner in advancing accessible, efficient, and high-quality stomatology services in China.
Artificial Intelligence
;
Humans
;
Oral Medicine/trends*
;
China
;
Delivery of Health Care
;
Machine Learning


Result Analysis
Print
Save
E-mail