1.Databases, knowledge bases, and large models for biomanufacturing.
Zhitao MAO ; Xiaoping LIAO ; Hongwu MA
Chinese Journal of Biotechnology 2025;41(3):901-916
Biomanufacturing is an advanced manufacturing method that integrates biology, chemistry, and engineering. It utilizes renewable biomass and biological organisms as production media to scale up the production of target products through fermentation. Compared with petrochemical routes, biomanufacturing offers significant advantages in reducing CO2 emissions, lowering energy consumption, and cutting costs. With the development of systems biology and synthetic biology and the accumulation of bioinformatics data, the integration of information technologies such as artificial intelligence, large models, and high-performance computing with biotechnology is propelling biomanufacturing into a data-driven era. This paper reviews the latest research progress on databases, knowledge bases, and large language models for biomanufacturing. It explores the development directions, challenges, and emerging technical methods in this field, aiming to provide guidance and inspiration for scientific research in related areas.
Biotechnology/methods*
;
Knowledge Bases
;
Synthetic Biology
;
Databases, Factual
;
Artificial Intelligence
;
Systems Biology
;
Computational Biology
;
Fermentation
2.Artificial intelligence-enhanced physics-based computational modeling technologies for proteins.
Baoyan LIU ; Shuai LI ; Hao SU ; Xiang SHENG
Chinese Journal of Biotechnology 2025;41(3):917-933
Computational modeling is an invaluable tool for mechanism analysis, directed engineering, and rational design of biological parts, metabolic networks, and even cellular systems. It can provide new technological solutions to address biological challenges at different levels and has become a central focus of research in biomanufacturing. In the computational modeling of proteins, which are the key parts in biological systems, the traditional physics-based methods (computer software and mathematical model) have been widely used to study the physical and chemical processes in the functioning of proteins, and have thus been recognized as a powerful tool for understanding complex biological systems and guiding experimental designs. As the scale of computational modeling continues to expand, traditional modeling techniques face difficulties in balancing computational accuracy and speed. In recent years, the explosive growth of biological data has made it possible to construct high-performance artificial intelligence (AI) models, which brings new opportunities to the computational modeling of proteins, and the AI-enhanced physics-based computational modeling technologies have emerged. This combined strategy not only incorporates the chemical knowledge and established physical principles but also is powerful in data processing and pattern recognition, which greatly improves the computational efficiency and prediction accuracy, as well as possesses stronger interpretation ability, transferability, and robustness. The AI-enhanced physics-based computational modeling technologies have already shown great potential and value in biocatalysis, paving a new way for the future development of biomanufacturing.
Artificial Intelligence
;
Proteins/chemistry*
;
Computer Simulation
;
Software
;
Computational Biology/methods*
3.Research progress in mutation effect prediction based on protein language models.
Liang ZHANG ; Pan TAN ; Liang HONG
Chinese Journal of Biotechnology 2025;41(3):934-948
Predicting protein mutation effects is a key challenge in bioinformatics and protein engineering. Recent advancements in deep learning, particularly the development of protein language models (PLMs), have brought new opportunities to this field. This review summarizes the application of PLMs in predicting protein mutation effects, focusing on three main types of models: sequence-based models, structure-based models, and models that combine sequence and structural information. We analyze in detail the principles, advantages, and limitations of these models and discuss the application of unsupervised and supervised learning in model training. Furthermore, this paper discusses the main challenges currently faced, including the acquisition of high-quality datasets and the handling of data noise. Finally, we look ahead to future research directions, including the application prospects of emerging technologies such as multimodal fusion and few-shot learning. This review aims to provide researchers with a comprehensive perspective to further advance the prediction of protein mutation effects.
Mutation
;
Proteins/chemistry*
;
Computational Biology/methods*
;
Deep Learning
;
Protein Engineering
4.Intelligent mining, engineering, and de novo design of proteins.
Cui LIU ; Zhenkun SHI ; Hongwu MA ; Xiaoping LIAO
Chinese Journal of Biotechnology 2025;41(3):993-1010
Natural components serve the survival instincts of cells that are obtained through long-term evolution, while they often fail to meet the demands of engineered cells for efficiently performing biological functions in special industrial environments. Enzymes, as biological catalysts, play a key role in biosynthetic pathways, significantly enhancing the rate and selectivity of biochemical reactions. However, the catalytic efficiency, stability, substrate specificity, and tolerance of natural enzymes often fall short of industrial production requirements. Therefore, exploring and modifying enzymes to suit specific biomanufacturing processes has become crucial. In recent years, artificial intelligence (AI) has played an increasingly important role in the discovery, evaluation, engineering, and de novo design of proteins. AI can accelerate the discovery and optimization of proteins by analyzing large amounts of bioinformatics data and predicting protein functions and characteristics by machine learning and deep learning algorithms. Moreover, AI can assist researchers in designing new protein structures by simulating and predicting their performance under different conditions, providing guidance for protein design. This paper reviews the latest research advances in protein discovery, evaluation, engineering, and de novo design for biomanufacturing and explores the hot topics, challenges, and emerging technical methods in this field, aiming to provide guidance and inspiration for researchers in related fields.
Protein Engineering/methods*
;
Artificial Intelligence
;
Proteins/genetics*
;
Computational Biology
;
Machine Learning
;
Data Mining
;
Algorithms
;
Deep Learning
5.pLM4ACP: a model for predicting anticancer peptides based on machine learning and protein language models.
Yitong LIU ; Wenxin CHEN ; Juanjuan LI ; Xue CHI ; Xiang MA ; Yanqiong TANG ; Hong LI
Chinese Journal of Biotechnology 2025;41(8):3252-3261
Cancer is a serious global health problem and a major cause of human death. Conventional cancer treatments often run the risk of impairing vital organ functions. Anticancer peptides (ACPs) are considered to be one of the most promising therapeutic agents against common human cancers due to their small sizes, high specificity, and low toxicity. Since ACP recognition is highly limited to the laboratory, expensive, and time-consuming, we proposed pLM4ACP, a model for predicting ACPs based on machine learning and protein language models. In this model, the protein language model ProtT5 was used to extract the features of ACPs, and the extracted features were input into the support vector machine (SVM) classification algorithm for optimization and performance evaluation. The model showcased significantly higher accuracy than other methods, with the overall accuracy of 0.763, F1-score of 0.767, Matthews correlation coefficient of 0.527, and area under the curve of 0.827 on the independent test set. This study constructs an efficient anticancer peptide prediction model based on protein language models, further advancing the application of artificial intelligence in the biomedical field and promoting the development of precision medicine and computational biology.
Machine Learning
;
Antineoplastic Agents/chemistry*
;
Humans
;
Peptides/chemistry*
;
Support Vector Machine
;
Algorithms
;
Computational Biology/methods*
;
Neoplasms/drug therapy*
6.Identification of prognosis-related key genes in hepatocellular carcinoma based on bioinformatics analysis.
Qian XIE ; Yingshan ZHU ; Ge HUANG ; Yue ZHAO
Journal of Central South University(Medical Sciences) 2025;50(2):167-180
OBJECTIVES:
Hepatocellular carcinoma is one of the most common primary malignant tumors with the third highest mortality rate worldwide. This study aims to identify key genes associated with hepatocellular carcinoma prognosis using the Gene Expression Omnibus (GEO) database and provide a theoretical basis for discovering novel prognostic biomarkers for hepatocellular carcinoma.
METHODS:
Hepatocellular carcinoma-related datasets were retrieved from the GEO database. Differentially expressed genes (DEGs) were identified using the GEO2R tool. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed using the Database for Annotation, Visualization, and Integrated Discovery (DAVID). A protein-protein interaction (PPI) network was constructed using the Search Tool for the Retrieval of Interacting Genes/Proteins (STRING), and key genes were identified using Cytoscape software. The University of Alabama at Birmingham Cancer Data Analysis Resource (UALCAN) was used to analyze the expression levels of key genes in normal and hepatocellular carcinoma tissues, as well as their associations with pathological grade, clinical stage, and patient survival. The Human Protein Atlas (THPA) was used to further validate the impact of key genes on overall survival. Expression levels of key genes in the blood of hepatocellular carcinoma patients were evaluated using the expression atlas of blood-based biomarkers in the early diagnosis of cancers (BBCancer).
RESULTS:
A total of 78 DEGs were identified from the GEO database. GO and KEGG analyses indicated that these genes may contribute to hepatocellular carcinoma progression by promoting cell division and regulating protein kinase activity. Sixteen key genes were screened via Cytoscape and validated using UALCAN and THPA. These genes were overexpressed in hepatocellular carcinoma tissues and were associated with disease progression and poor prognosis. Finally, BBCancer analysis showed that ASPM and NCAPG were also elevated in the blood of hepatocellular carcinoma patients.
CONCLUSIONS
This study identified 16 key genes as potential prognostic biomarkers for hepatocellular carcinoma, among which ASPM and NCAPG may serve as promising blood-based markers for hepatocellular carcinoma.
Humans
;
Carcinoma, Hepatocellular/mortality*
;
Liver Neoplasms/pathology*
;
Prognosis
;
Computational Biology/methods*
;
Protein Interaction Maps/genetics*
;
Biomarkers, Tumor/genetics*
;
Gene Expression Regulation, Neoplastic
;
Gene Expression Profiling
;
Gene Ontology
;
Databases, Genetic
7.Identification of shared key genes and pathways in osteoarthritis and sarcopenia patients based on bioinformatics analysis.
Yuyan SUN ; Ziyu LUO ; Huixian LING ; Sha WU ; Hongwei SHEN ; Yuanyuan FU ; Thainamanh NGO ; Wen WANG ; Ying KONG
Journal of Central South University(Medical Sciences) 2025;50(3):430-446
OBJECTIVES:
Osteoarthritis (OA) and sarcopenia are significant health concerns in the elderly, substantially impacting their daily activities and quality of life. However, the relationship between them remains poorly understood. This study aims to uncover common biomarkers and pathways associated with both OA and sarcopenia.
METHODS:
Gene expression profiles related to OA and sarcopenia were retrieved from the Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) between disease and control groups were identified using R software. Common DEGs were extracted via Venn diagram analysis. Gene ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were conducted to identify biological processes and pathways associated with shared DEGs. Protein-protein interaction (PPI) networks were constructed, and candidate hub genes were ranked using the maximal clique centrality (MCC) algorithm. Further validation of hub gene expression was performed using 2 independent datasets. Receiver operating characteristic (ROC) curve analysis was used to evaluate the predictive value of key genes for OA and sarcopenia. Mouse models of OA and sarcopenia were established. Hematoxylin-eosin and Safranin O/Fast Green staining were used to validate the OA model. The sarcopenia model was validated via rotarod testing and quadriceps muscle mass measurement. Real-time reverse transcription PCR (real-time RT-PCR) was employed to assess the mRNA expression levels of candidate key genes in both models. Gene set enrichment analysis (GSEA) was conducted to identify pathways associated with the selected shared key genes in both diseases.
RESULTS:
A total of 89 common DEGs were identified in the gene expression profiles of OA and sarcopenia, including 76 upregulated and 13 downregulated genes. These 89 DEGs were significantly enriched in protein digestion and absorption, the PI3K-Akt signaling pathway, and extracellular matrix-receptor interaction. PPI network analysis and MCC algorithm analysis of the 89 common DEGs identified the top 17 candidate hub genes. Based on the differential expression analysis of these 17 candidate hub genes in the validation datasets, AEBP1 and COL8A2 were ultimately selected as the common key genes for both diseases, both of which showed a significant upregulation trend in the disease groups (all P<0.05). The value of area under the curve (AUC) for AEBP1 and COL8A2 in the OA and sarcopenia datasets were all greater than 0.7, indicating that both genes have potential value in predicting OA and sarcopenia. Real-time RT-PCR results showed that the mRNA expression levels of AEBP1 and COL8A2 were significantly upregulated in the disease groups (all P<0.05), consistent with the results observed in the bioinformatics analysis. GSEA revealed that AEBP1 and COL8A2 were closely related to extracellular matrix-receptor interaction, ribosome, and oxidative phosphorylation in OA and sarcopenia.
CONCLUSIONS
AEBP1 and COL8A2 have the potential to serve as common biomarkers for OA and sarcopenia. The extracellular matrix-receptor interaction pathway may represent a potential target for the prevention and treatment of both OA and sarcopenia.
Sarcopenia/genetics*
;
Osteoarthritis/genetics*
;
Computational Biology/methods*
;
Humans
;
Protein Interaction Maps/genetics*
;
Animals
;
Mice
;
Gene Expression Profiling
;
Gene Ontology
;
Transcriptome
;
Male
;
Signal Transduction/genetics*
;
Gene Regulatory Networks
8.Construction of a treatment response prediction model for multiple myeloma based on multi-omics and machine learning.
Xionghui ZHOU ; Rong GUI ; Jing LIU ; Meng GAO
Journal of Central South University(Medical Sciences) 2025;50(4):531-544
OBJECTIVES:
Multiple myeloma (MM) is a hematologic malignancy characterized by clonal proliferation of plasma cells and remains incurable. Patients with primary refractory multiple myeloma (PRMM) show poor response to initial induction therapy. This study aims to develop a machine learning-based model to predict treatment response in newly diagnosed multiple myeloma (NDMM) patients, in order to optimize therapeutic strategies.
METHODS:
NDMM and post-treatment MM patients hospitalized in the Department of Hematology, Third Xiangya Hospital, Central South University, between August 2022 and July 2023 were enrolled. Post-treatment MM patients were categorized into PRMM patients and treatment-responsive MM (TRMM) patients based on therapeutic efficacy. Serum metabolites were detected and analyzed via metabolomics. Based on the metabolomics analysis results and combined with transcriptomic sequencing data of NDMM patients from databases, differentially expressed amino acid metabolism-related genes (AAMGs) among post-treatment NDMM patients with varying therapeutic outcomes were screened. Using bioinformatics analyses and machine learning algorithms, a predictive model for treatment response in NDMM was constructed and used to identify patients at risk for PRMM.
RESULTS:
A total of 61 patients were included: 22 NDMM, 23 TRMM, and 16 PRMM patients. Significant differences in metabolite levels were observed among the 3 groups, with differential metabolites mainly enriched in amino acid metabolism pathways. Follow-up data were available for 16 of the 22 NDMM patients, including 12 treatment responders (ND_TR group) and 4 with PRMM (ND_PR group). A total of 23 differential metabolites were identified between these 2 groups: 6 metabolites (e.g., tryptophan) were upregulated and 17 (e.g., citric acid) were downregulated in the ND_TR group. Transcriptomic data from 108 TRMM and 77 PRMM patients were analyzed to identify differentially expressed AAMGs, which were then used to construct a prediction model. The area under the receiver operating characteristic curve (AUC) for the model exceeded 0.8, and AUC values in 3 external validation cohorts were all above 0.7.
CONCLUSIONS
This study delineated the metabolic alterations in MM patients with different treatment response, suggesting that dysregulated amino acid metabolism may be associated with poor treatment response in PRMM. By integrating metabolomics and transcriptomics, a machine learning-based predictive model was successfully established to forecast treatment response in NDMM patients.
Humans
;
Multiple Myeloma/drug therapy*
;
Machine Learning
;
Male
;
Female
;
Metabolomics/methods*
;
Middle Aged
;
Aged
;
Treatment Outcome
;
Transcriptome
;
Computational Biology
;
Adult
;
Multiomics
9.Hub biomarkers and their clinical relevance in glycometabolic disorders: A comprehensive bioinformatics and machine learning approach.
Liping XIANG ; Bing ZHOU ; Yunchen LUO ; Hanqi BI ; Yan LU ; Jian ZHOU
Chinese Medical Journal 2025;138(16):2016-2027
BACKGROUND:
Gluconeogenesis is a critical metabolic pathway for maintaining glucose homeostasis, and its dysregulation can lead to glycometabolic disorders. This study aimed to identify hub biomarkers of these disorders to provide a theoretical foundation for enhancing diagnosis and treatment.
METHODS:
Gene expression profiles from liver tissues of three well-characterized gluconeogenesis mouse models were analyzed to identify commonly differentially expressed genes (DEGs). Weighted gene co-expression network analysis (WGCNA), machine learning techniques, and diagnostic tests on transcriptome data from publicly available datasets of type 2 diabetes mellitus (T2DM) patients were employed to assess the clinical relevance of these DEGs. Subsequently, we identified hub biomarkers associated with gluconeogenesis-related glycometabolic disorders, investigated potential correlations with immune cell types, and validated expression using quantitative polymerase chain reaction in the mouse models.
RESULTS:
Only a few common DEGs were observed in gluconeogenesis-related glycometabolic disorders across different contributing factors. However, these DEGs were consistently associated with cytokine regulation and oxidative stress (OS). Enrichment analysis highlighted significant alterations in terms related to cytokines and OS. Importantly, osteomodulin ( OMD ), apolipoprotein A4 ( APOA4 ), and insulin like growth factor binding protein 6 ( IGFBP6 ) were identified with potential clinical significance in T2DM patients. These genes demonstrated robust diagnostic performance in T2DM cohorts and were positively correlated with resting dendritic cells.
CONCLUSIONS
Gluconeogenesis-related glycometabolic disorders exhibit considerable heterogeneity, yet changes in cytokine regulation and OS are universally present. OMD , APOA4 , and IGFBP6 may serve as hub biomarkers for gluconeogenesis-related glycometabolic disorders.
Machine Learning
;
Humans
;
Computational Biology/methods*
;
Biomarkers/metabolism*
;
Diabetes Mellitus, Type 2/genetics*
;
Animals
;
Mice
;
Gluconeogenesis/physiology*
;
Gene Expression Profiling
;
Transcriptome/genetics*
;
Gene Regulatory Networks/genetics*
;
Clinical Relevance
10.Computational pathology in precision oncology: Evolution from task-specific models to foundation models.
Yuhao WANG ; Yunjie GU ; Xueyuan ZHANG ; Baizhi WANG ; Rundong WANG ; Xiaolong LI ; Yudong LIU ; Fengmei QU ; Fei REN ; Rui YAN ; S Kevin ZHOU
Chinese Medical Journal 2025;138(22):2868-2878
With the rapid development of artificial intelligence, computational pathology has been seamlessly integrated into the entire clinical workflow, which encompasses diagnosis, treatment, prognosis, and biomarker discovery. This integration has significantly enhanced clinical accuracy and efficiency while reducing the workload for clinicians. Traditionally, research in this field has depended on the collection and labeling of large datasets for specific tasks, followed by the development of task-specific computational pathology models. However, this approach is labor intensive and does not scale efficiently for open-set identification or rare diseases. Given the diversity of clinical tasks, training individual models from scratch to address the whole spectrum of clinical tasks in the pathology workflow is impractical, which highlights the urgent need to transition from task-specific models to foundation models (FMs). In recent years, pathological FMs have proliferated. These FMs can be classified into three categories, namely, pathology image FMs, pathology image-text FMs, and pathology image-gene FMs, each of which results in distinct functionalities and application scenarios. This review provides an overview of the latest research advancements in pathological FMs, with a particular emphasis on their applications in oncology. The key challenges and opportunities presented by pathological FMs in precision oncology are also explored.
Humans
;
Precision Medicine/methods*
;
Medical Oncology/methods*
;
Artificial Intelligence
;
Neoplasms/pathology*
;
Computational Biology/methods*

Result Analysis
Print
Save
E-mail