1.TCM Data Hub: A traditional Chinese medicine data platform powered by YiYuan large language models
Chongyun ZHOU ; Qin LI ; Tangming CUI ; Chaohui CUI ; Peiyu WANG ; Meiling SUN ; Ying NIE ; Yichen BAI ; Haiyan LI
Science of Traditional Chinese Medicine 2026;4(2):140-151
The digitization of traditional Chinese medicine (TCM) has generated vast amounts of data. However, these data are characterized by significant heterogeneity and complex semantic structures, posing substantial challenges for systematic integration and intelligent analysis, and limiting its potential for modern clinical and computational research. To address the challenges posed by the high heterogeneity and complex structure in TCM data, we designed and developed the TCM Data Hub platform, which is powered by the YiYuan large language models (LLMs). This platform aims to enhance intelligent data processing capabilities and unlock the potential for clinical application of TCM data through systematic integration and efficient utilization, thereby bridging the gap between traditional knowledge and modern computational research. This study first analyzed the heterogeneity and complexity of TCM information with respect to data types, structures, and semantics. A standardized data framework was constructed to enhance data integration and interoperability. Based on the TCM Intelligent Computing Platform of the China Academy of Chinese Medical Sciences, we trained the YiYuan LLMs to acquire domain-specific semantic understanding of TCM, thereby improving the platform’s comprehension of specialized terminology and knowledge systems. Leveraging the natural language processing capabilities of the LLMs, we developed a human-in-the-loop data processing system to enable efficient extraction, cleansing, and structured organization of TCM data. In addition, utilizing Vue and Java technologies, we developed multiple LLM-powered intelligent agents and systems, including a human-in-the-loop data processing system, as well as automated prescription mining and network pharmacology analysis agents. Task-specific agents tailored to TCM data processing were developed to enhance the model’s effectiveness in clinical knowledge discovery. System functionality and platform infrastructure were implemented using Java and Vue technologies.The TCM Data Hub platform has completed system construction and core functionality implementation. It supported integrated management and efficient access to 8 key types of TCM data: prescriptions, materia medica, ingredients, targets, diseases (Western medicine), diseases (TCM), syndromes, and therapeutic methods. The human-in-the-loop data processing system achieved an accuracy of 95.34% in structuring TCM data and supported annotation for data requiring manual labeling. The intelligent agent-driven big-data analytics module enabled 1-click, end-to-end workflows for TCM prescription mining, herb-syndrome association analysis, network pharmacology, and molecular biology research, completing a full data mining task in approximately 30 minutes. Users can interact with and manipulate data through a visual front-end interface. The system demonstrated stable performance, strong scalability, and a user-friendly experience. Empowered by the YiYuan LLMs, the TCM Data Hub platform significantly improves the accessibility, usability, and intelligence of TCM data. It effectively bridges traditional TCM knowledge with modern intelligent technologies, providing robust data support and intelligent tools for TCM research and clinical applications.
2.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
3.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
4.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
5.Construction of a Diagnostic Model for Traditional Chinese Medicine Syndromes of Chronic Cough Based on the Voting Ensemble Machine Learning Algorithm
Yichen BAI ; Suyang QIN ; Chongyun ZHOU ; Liqing SHI ; Kun JI ; Chuchu ZHANG ; Panfei LI ; Tangming CUI ; Haiyan LI
Journal of Traditional Chinese Medicine 2025;66(11):1119-1127
ObjectiveTo explore the construction of a machine learning model for the diagnosis of traditional Chinese medicine (TCM) syndromes in chronic cough and the optimization of this model using the Voting ensemble algorithm. MethodsA retrospective analysis was conducted using clinical data from 921 patients with chronic cough treated at the Respiratory Department of Dongfang Hospital, Beijing University of Chinese Medicine. After standardized processing, 84 clinical features were extracted to determine TCM syndrome types. A specialized dataset for TCM syndrome diagnosis in chronic cough was formed by selecting syndrome types with more than 50 cases. The synthetic minority over-sampling technique (SMOTE) was employed to balance the dataset. Four base models, logistic regression (LR), decision tree (dt), multilayer perceptron (MLP), and Bagging, were constructed and integrated using a hard voting strategy to form a Voting ensemble model. Model performance was evaluated using accuracy, recall, precision, F1-score, receiver operating characteristic (ROC) curve, area under the curve (AUC), and confusion matrix. ResultsAmong the 921 cases, six syndrome types had over 50 cases each, phlegm-heat obstructing the lung (294 cases), wind pathogen latent in the lung (103 cases), cold-phlegm obstructing the lung (102 cases), damp-heat stagnating in the lung (64 cases), lung yang deficiency (54 cases), and phlegm-damp obstructing the lung (53 cases), yielding a total of 670 cases in the specialized dataset. High-frequency symptoms among these patients included cough, expectoration, odor-induced cough, throat itchiness, itch-induced cough, and cough triggered by cold wind. Among the four base models, the MLP model showed the best diagnostic performance (test accuracy: 0.9104; AUC: 0.9828). Compared with the base models, the Voting ensemble model achieved superior performance with an accuracy of 0.9289 on the training set and 0.9253 on the test set, showing a minimal overfitting gap of 0.0036. It also achieved the highest AUC (0.9836) in the test set, outperforming all base models. The model exhi-bited especially strong diagnostic performance for damp-heat stagnating in the lung (AUC: 0.9984) and wind pathogen latent in the lung (AUC: 0.9970). ConclusionThe Voting ensemble algorithm effectively integrates the strengths of multiple machine learning models, resulting in an optimized diagnostic model for TCM syndromes in chronic cough with high accuracy and enhanced generalization ability.
6.Metagenomics reveals an increased proportion of an Escherichia coli-dominated enterotype in elderly Chinese people.
Jinyou LI ; Yue WU ; Yichen YANG ; Lufang CHEN ; Caihong HE ; Shixian ZHOU ; Shunmei HUANG ; Xia ZHANG ; Yuming WANG ; Qifeng GUI ; Haifeng LU ; Qin ZHANG ; Yunmei YANG
Journal of Zhejiang University. Science. B 2025;26(5):477-492
Gut microbial communities are likely remodeled in tandem with accumulated physiological decline during aging, yet there is limited understanding of gut microbiome variation in advanced age. Here, we performed a metagenomics-based enterotype analysis in a geographically homogeneous cohort of 367 enrolled Chinese individuals between the ages of 60 and 94 years, with the goal of characterizing the gut microbiome of elderly individuals and identifying factors linked to enterotype variations. In addition to two adult-like enterotypes dominated by Bacteroides (ET-Bacteroides) and Prevotella (ET-Prevotella), we identified a novel enterotype dominated by Escherichia (ET-Escherichia), whose prevalence increased in advanced age. Our data demonstrated that age explained more of the variance in the gut microbiome than previously identified factors such as type 2 diabetes mellitus (T2DM) or diet. We characterized the distinct taxonomic and functional profiles of ET-Escherichia, and found the strongest cohesion and highest robustness of the microbial co-occurrence network in this enterotype, as well as the lowest species diversity. In addition, we carried out a series of correlation analyses and co-abundance network analyses, which showed that several factors were likely linked to the overabundance of Escherichia members, including advanced age, vegetable intake, and fruit intake. Overall, our data revealed an enterotype variation characterized by Escherichia enrichment in the elderly population. Considering the different age distribution of each enterotype, these findings provide new insights into the changes that occur in the gut microbiome with age and highlight the importance of microbiome-based stratification of elderly individuals.
Aged
;
Aged, 80 and over
;
Female
;
Humans
;
Male
;
Middle Aged
;
Bacteroides
;
China
;
Diabetes Mellitus, Type 2/microbiology*
;
Escherichia coli/classification*
;
Gastrointestinal Microbiome/genetics*
;
Metagenomics
;
East Asian People
7.Perceived stress and occupational burnout among hospital staff in Guangzhou tertiary hospitals
Wenli ZHOU ; Xiaoyi WU ; Yichen YE ; Liman WU ; Biyun CHEN ; Yi SHEN
Journal of Environmental and Occupational Medicine 2025;42(3):354-359
Background Staff in tertiary hospitals are a high-risk group for occupational burnout. Timely identification and precise intervention are crucial for improving healthcare service quality. However, comparative studies on perceived stress and occupational burnout among hospital staff in different positions are lacking. Objective To describe the status of perceived stress and occupational burnout among hospital staff in different positions and compare the differences, explore the relationship between perceived stress and occupational burnout, and identify the influencing factors of occupational burnout. Methods In May 2022,
8.Identification of potential therapeutic targets for hair color and hair shaft abnormalities by integrating human plasma proteomics
Guangdi LI ; Guiwen ZHOU ; Yichen WANG ; Xiao XU ; Minliang CHEN
Chinese Journal of Plastic Surgery 2025;41(5):482-491
Objective:To identify potential therapeutic targets for the treatment of hair color and hair shaft abnormalities, providing novel insights and approaches for managing related conditions.Methods:Using the pQTL data from large-scale proteomics studies, a two-sample Mendelian randomization (TwoSampleMR) was conducted to preliminarily identify potential drug therapeutic targets. Subsequently, sensitivity analysis was performed to evaluate potential confounding factors such as heterogeneity and horizontal pleiotropy, and a leave-one-out sensitivity analysis was conducted by eliminating each single nucleotide polymorphism (SNP) one by one. Finally, co-localization analysis was carried out to explore whether there are shared genetic variants between the identified plasma proteins and traits.Results:The proteomic data used in this study included 4, 907 pQTLs, while the genetic data related to hair pigmentation and shaft abnormalities comprised 124 cases and 432, 686 controls. After Mendelian randomization screening, six candidate protein genes were identified: HLA-DQA2, CTSB, KIR2DS2, SVEP1, HOMER2, and HOMER1. Sensitivity analyses revealed no evidence of heterogeneity or horizontal pleiotropy among these proteins. Leave-one-out sensitivity analysis indicated that no single SNP significantly influenced the overall result. Notably, HOMER2 was supported by colocalization analysis with strong evidence, suggesting its potential role in regulating genetic mechanisms underlying hair pigmentation and shaft health.Conclusions:This study identified six potential therapeutic targets for hair pigmentation and shaft abnormalities, with HOMER2 showing stronger evidence. These findings provide novel directions for the treatment of hair color and hair shaft abnormalities.
9.Hotspots and trends in research field of inflammation in polycystic ovary syndrome:a bibliometric visualization analysis
Yichen ZHOU ; Xiuqi YIN ; Bingyi YANG ; Qitian LU ; Weian YUAN
Journal of Clinical Medicine in Practice 2025;29(10):89-96
Objective To analyze the hotspots and trends in the research field of inflammation in polycystic ovary syndrome(PCOS).Methods This study retrieved data from the Web of Science Core Collection database,with the retrieval period spanning from January 2000 to December 2023.The litera-ture types included were papers and review papers in English.The search was conducted using"Poly-cystic Ovarian Syndrome"and its synonyms,along with"inflammation"as subject terms.Visualization analysis was performed using software such as CiteSpace,VOSviewer,and SCImago Graphica.Results A total of 1,803 articles were included in this study.The number of publications in this field had rapidly increased after 2018.China had the highest number of publications,and Tehran University of Medical Sciences was the institution with the most publications.Key scholars included Asemi Z,Gonzalez AM,Escobar-Morreale HF,etc.,and a core group of authors had been formed.The main research hotspots centered around insulin resistance,obesity,oxidative stress,etc.,and five clusters were formed.Highly cited and burst-cited literature were mainly review papers.Conclusion The re-search field of inflammation in PCOS is currently in a stage of rapid development,with research content covering PCOS inflammation markers,the impact of inflammation on PCOS pathophysiology,inflamma-tion-based PCOS treatment and prognosis,etc.The debate on whether the chronic low-grade inflammatory state in PC OS patients is related to obesity or not remains a focal point of contention.
10.Identification of potential therapeutic targets for hair color and hair shaft abnormalities by integrating human plasma proteomics
Guangdi LI ; Guiwen ZHOU ; Yichen WANG ; Xiao XU ; Minliang CHEN
Chinese Journal of Plastic Surgery 2025;41(7):734-743
Objective:To identify potential therapeutic targets for the treatment of hair color and hair shaft abnormalities, thereby offering innovative insights and strategies for the management of the associated conditions.Methods:Using the protein quantitative trait locus(pQTL) data derived from extensive proteomics studies, a two-sample Mendelian randomization was conducted to preliminarily identify potential drug therapeutic targets. Following this, sensitivity analyses were performed to evaluate potential confounding factors, including heterogeneity and horizontal pleiotropy. A leave-one-out sensitivity analysis was also conducted, systematically excluding each single nucleotide polymorphism (SNP) to evaluate their individual impact. Additionally, co-localization analysis was carried out to determine the presence of shared genetic variants between the identified plasma proteins and the relevant traits.Results:The proteomic data analyzed in this investigation encompassed 4 907 pQTLs, while the genetic data pertaining to hair pigmentation and shaft abnormalities included 124 cases and 432 686 controls. Through the Mendelian randomization screening, six candidate protein genes were identified: HLA-DQA2, CTSB, KIR2DS2, SVEP1, HOMER2, and HOMER1. Sensitivity analyses revealed no evidence of heterogeneity or horizontal pleiotropy among these proteins. The leave-one-out sensitivity analysis demonstrated that no single SNP significantly affected the overall findings. Notably, HOMER2 was substantiated by co-localization analysis, which provided robust evidence of its potential role in regulating the genetic mechanisms associated with hair pigmentation and shaft integrity.Conclusion:This study successfully identified six potential therapeutic targets for hair pigmentation and shaft abnormalities, with HOMER2 exhibiting the most compelling evidence. These findings pave the way for novel therapeutic approaches in the treatment of hair color and hair shaft disorders.

Result Analysis
Print
Save
E-mail