1.TCM Data Hub: A traditional Chinese medicine data platform powered by YiYuan large language models
Chongyun ZHOU ; Qin LI ; Tangming CUI ; Chaohui CUI ; Peiyu WANG ; Meiling SUN ; Ying NIE ; Yichen BAI ; Haiyan LI
Science of Traditional Chinese Medicine 2026;4(2):140-151
The digitization of traditional Chinese medicine (TCM) has generated vast amounts of data. However, these data are characterized by significant heterogeneity and complex semantic structures, posing substantial challenges for systematic integration and intelligent analysis, and limiting its potential for modern clinical and computational research. To address the challenges posed by the high heterogeneity and complex structure in TCM data, we designed and developed the TCM Data Hub platform, which is powered by the YiYuan large language models (LLMs). This platform aims to enhance intelligent data processing capabilities and unlock the potential for clinical application of TCM data through systematic integration and efficient utilization, thereby bridging the gap between traditional knowledge and modern computational research. This study first analyzed the heterogeneity and complexity of TCM information with respect to data types, structures, and semantics. A standardized data framework was constructed to enhance data integration and interoperability. Based on the TCM Intelligent Computing Platform of the China Academy of Chinese Medical Sciences, we trained the YiYuan LLMs to acquire domain-specific semantic understanding of TCM, thereby improving the platform’s comprehension of specialized terminology and knowledge systems. Leveraging the natural language processing capabilities of the LLMs, we developed a human-in-the-loop data processing system to enable efficient extraction, cleansing, and structured organization of TCM data. In addition, utilizing Vue and Java technologies, we developed multiple LLM-powered intelligent agents and systems, including a human-in-the-loop data processing system, as well as automated prescription mining and network pharmacology analysis agents. Task-specific agents tailored to TCM data processing were developed to enhance the model’s effectiveness in clinical knowledge discovery. System functionality and platform infrastructure were implemented using Java and Vue technologies.The TCM Data Hub platform has completed system construction and core functionality implementation. It supported integrated management and efficient access to 8 key types of TCM data: prescriptions, materia medica, ingredients, targets, diseases (Western medicine), diseases (TCM), syndromes, and therapeutic methods. The human-in-the-loop data processing system achieved an accuracy of 95.34% in structuring TCM data and supported annotation for data requiring manual labeling. The intelligent agent-driven big-data analytics module enabled 1-click, end-to-end workflows for TCM prescription mining, herb-syndrome association analysis, network pharmacology, and molecular biology research, completing a full data mining task in approximately 30 minutes. Users can interact with and manipulate data through a visual front-end interface. The system demonstrated stable performance, strong scalability, and a user-friendly experience. Empowered by the YiYuan LLMs, the TCM Data Hub platform significantly improves the accessibility, usability, and intelligence of TCM data. It effectively bridges traditional TCM knowledge with modern intelligent technologies, providing robust data support and intelligent tools for TCM research and clinical applications.
2.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
3.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
4.Rapid quality control method for automated organ-at-risk contouring in head and neck radiotherapy
Xiaoyu YANG ; Kaining YAO ; Yichen PU ; Shun ZHOU ; Ruoxi WANG ; Haizhen YUE ; Hao WU
Chinese Journal of Radiological Health 2026;35(2):193-199
Objective To address the need for rapid review of automated organ-at-risk (OAR) contouring for head-and-neck cancer in online adaptive radiotherapy, this study used dosimetric indices as the reference standard to evaluate the suitability of geometric similarity metrics and to develop a rapid screening model that balances the risk of missed errors with review efficiency. Methods This retrospective study included 29 patients with head-and-neck radiotherapy, yielding 243 pairs of OAR contours (auto-generated vs. manual). Geometric similarity metrics, including the Dice similarity coefficient (Dice), the 95th percentile Hausdorff distance (HD95), and the maximum Hausdorff distance (HD100), were computed between the two contour sets. Using the dose distributions from clinical treatment plans, dosimetric differences between the two contour sets were calculated for key indices, including mean dose (Dmean) and maximum dose (Dmax). Correlations between geometric metrics and dosimetric differences were analyzed. Dosimetric discrepancy events were defined as Y3Gy using a threshold of 3 Gy. Within a univariable logistic regression triage framework, each geometric metric was used as an independent variable to estimate the probability of Y3Gy. Discriminative performance was assessed using receiver operating characteristic (ROC) curves and precision-recall (PR) curves, with the area under the curve (AUC) including ROC-AUC and PR-AUC. Practical operating thresholds were determined via threshold sweeping. Results Geometric similarity metrics showed weak-to-moderate correlations with dosimetric differences; compared with maximum dose, correlations with mean dose differences were more consistent. Among the evaluated metrics, HD95 achieved the best classification performance for Y3Gy. Threshold sweeping suggested that an HD95 threshold in the range of 4-9 mm can balance the risk of missed discrepancies with the efficiency of automated contour quality assurance (QA). Conclusion Several geometric similarity metrics demonstrated only weak-to-moderate associations with dosimetric differences. For head-and-neck OAR contour QA in this study, HD95 provided the best discriminative performance and can support rapid triage of automated contours, with a practical operating threshold range of 4-9 mm, potentially improving the efficiency of adaptive radiotherapy workflows.
5.Metagenomics reveals an increased proportion of an Escherichia coli-dominated enterotype in elderly Chinese people.
Jinyou LI ; Yue WU ; Yichen YANG ; Lufang CHEN ; Caihong HE ; Shixian ZHOU ; Shunmei HUANG ; Xia ZHANG ; Yuming WANG ; Qifeng GUI ; Haifeng LU ; Qin ZHANG ; Yunmei YANG
Journal of Zhejiang University. Science. B 2025;26(5):477-492
Gut microbial communities are likely remodeled in tandem with accumulated physiological decline during aging, yet there is limited understanding of gut microbiome variation in advanced age. Here, we performed a metagenomics-based enterotype analysis in a geographically homogeneous cohort of 367 enrolled Chinese individuals between the ages of 60 and 94 years, with the goal of characterizing the gut microbiome of elderly individuals and identifying factors linked to enterotype variations. In addition to two adult-like enterotypes dominated by Bacteroides (ET-Bacteroides) and Prevotella (ET-Prevotella), we identified a novel enterotype dominated by Escherichia (ET-Escherichia), whose prevalence increased in advanced age. Our data demonstrated that age explained more of the variance in the gut microbiome than previously identified factors such as type 2 diabetes mellitus (T2DM) or diet. We characterized the distinct taxonomic and functional profiles of ET-Escherichia, and found the strongest cohesion and highest robustness of the microbial co-occurrence network in this enterotype, as well as the lowest species diversity. In addition, we carried out a series of correlation analyses and co-abundance network analyses, which showed that several factors were likely linked to the overabundance of Escherichia members, including advanced age, vegetable intake, and fruit intake. Overall, our data revealed an enterotype variation characterized by Escherichia enrichment in the elderly population. Considering the different age distribution of each enterotype, these findings provide new insights into the changes that occur in the gut microbiome with age and highlight the importance of microbiome-based stratification of elderly individuals.
Aged
;
Aged, 80 and over
;
Female
;
Humans
;
Male
;
Middle Aged
;
Bacteroides
;
China
;
Diabetes Mellitus, Type 2/microbiology*
;
Escherichia coli/classification*
;
Gastrointestinal Microbiome/genetics*
;
Metagenomics
;
East Asian People
6.Identification of potential therapeutic targets for hair color and hair shaft abnormalities by integrating human plasma proteomics
Guangdi LI ; Guiwen ZHOU ; Yichen WANG ; Xiao XU ; Minliang CHEN
Chinese Journal of Plastic Surgery 2025;41(7):734-743
Objective:To identify potential therapeutic targets for the treatment of hair color and hair shaft abnormalities, thereby offering innovative insights and strategies for the management of the associated conditions.Methods:Using the protein quantitative trait locus(pQTL) data derived from extensive proteomics studies, a two-sample Mendelian randomization was conducted to preliminarily identify potential drug therapeutic targets. Following this, sensitivity analyses were performed to evaluate potential confounding factors, including heterogeneity and horizontal pleiotropy. A leave-one-out sensitivity analysis was also conducted, systematically excluding each single nucleotide polymorphism (SNP) to evaluate their individual impact. Additionally, co-localization analysis was carried out to determine the presence of shared genetic variants between the identified plasma proteins and the relevant traits.Results:The proteomic data analyzed in this investigation encompassed 4 907 pQTLs, while the genetic data pertaining to hair pigmentation and shaft abnormalities included 124 cases and 432 686 controls. Through the Mendelian randomization screening, six candidate protein genes were identified: HLA-DQA2, CTSB, KIR2DS2, SVEP1, HOMER2, and HOMER1. Sensitivity analyses revealed no evidence of heterogeneity or horizontal pleiotropy among these proteins. The leave-one-out sensitivity analysis demonstrated that no single SNP significantly affected the overall findings. Notably, HOMER2 was substantiated by co-localization analysis, which provided robust evidence of its potential role in regulating the genetic mechanisms associated with hair pigmentation and shaft integrity.Conclusion:This study successfully identified six potential therapeutic targets for hair pigmentation and shaft abnormalities, with HOMER2 exhibiting the most compelling evidence. These findings pave the way for novel therapeutic approaches in the treatment of hair color and hair shaft disorders.
7.Identification of potential therapeutic targets for hair color and hair shaft abnormalities by integrating human plasma proteomics
Guangdi LI ; Guiwen ZHOU ; Yichen WANG ; Xiao XU ; Minliang CHEN
Chinese Journal of Plastic Surgery 2025;41(5):482-491
Objective:To identify potential therapeutic targets for the treatment of hair color and hair shaft abnormalities, providing novel insights and approaches for managing related conditions.Methods:Using the pQTL data from large-scale proteomics studies, a two-sample Mendelian randomization (TwoSampleMR) was conducted to preliminarily identify potential drug therapeutic targets. Subsequently, sensitivity analysis was performed to evaluate potential confounding factors such as heterogeneity and horizontal pleiotropy, and a leave-one-out sensitivity analysis was conducted by eliminating each single nucleotide polymorphism (SNP) one by one. Finally, co-localization analysis was carried out to explore whether there are shared genetic variants between the identified plasma proteins and traits.Results:The proteomic data used in this study included 4, 907 pQTLs, while the genetic data related to hair pigmentation and shaft abnormalities comprised 124 cases and 432, 686 controls. After Mendelian randomization screening, six candidate protein genes were identified: HLA-DQA2, CTSB, KIR2DS2, SVEP1, HOMER2, and HOMER1. Sensitivity analyses revealed no evidence of heterogeneity or horizontal pleiotropy among these proteins. Leave-one-out sensitivity analysis indicated that no single SNP significantly influenced the overall result. Notably, HOMER2 was supported by colocalization analysis with strong evidence, suggesting its potential role in regulating genetic mechanisms underlying hair pigmentation and shaft health.Conclusions:This study identified six potential therapeutic targets for hair pigmentation and shaft abnormalities, with HOMER2 showing stronger evidence. These findings provide novel directions for the treatment of hair color and hair shaft abnormalities.
8.Construction of a Diagnostic Model for Traditional Chinese Medicine Syndromes of Chronic Cough Based on the Voting Ensemble Machine Learning Algorithm
Yichen BAI ; Suyang QIN ; Chongyun ZHOU ; Liqing SHI ; Kun JI ; Chuchu ZHANG ; Panfei LI ; Tangming CUI ; Haiyan LI
Journal of Traditional Chinese Medicine 2025;66(11):1119-1127
ObjectiveTo explore the construction of a machine learning model for the diagnosis of traditional Chinese medicine (TCM) syndromes in chronic cough and the optimization of this model using the Voting ensemble algorithm. MethodsA retrospective analysis was conducted using clinical data from 921 patients with chronic cough treated at the Respiratory Department of Dongfang Hospital, Beijing University of Chinese Medicine. After standardized processing, 84 clinical features were extracted to determine TCM syndrome types. A specialized dataset for TCM syndrome diagnosis in chronic cough was formed by selecting syndrome types with more than 50 cases. The synthetic minority over-sampling technique (SMOTE) was employed to balance the dataset. Four base models, logistic regression (LR), decision tree (dt), multilayer perceptron (MLP), and Bagging, were constructed and integrated using a hard voting strategy to form a Voting ensemble model. Model performance was evaluated using accuracy, recall, precision, F1-score, receiver operating characteristic (ROC) curve, area under the curve (AUC), and confusion matrix. ResultsAmong the 921 cases, six syndrome types had over 50 cases each, phlegm-heat obstructing the lung (294 cases), wind pathogen latent in the lung (103 cases), cold-phlegm obstructing the lung (102 cases), damp-heat stagnating in the lung (64 cases), lung yang deficiency (54 cases), and phlegm-damp obstructing the lung (53 cases), yielding a total of 670 cases in the specialized dataset. High-frequency symptoms among these patients included cough, expectoration, odor-induced cough, throat itchiness, itch-induced cough, and cough triggered by cold wind. Among the four base models, the MLP model showed the best diagnostic performance (test accuracy: 0.9104; AUC: 0.9828). Compared with the base models, the Voting ensemble model achieved superior performance with an accuracy of 0.9289 on the training set and 0.9253 on the test set, showing a minimal overfitting gap of 0.0036. It also achieved the highest AUC (0.9836) in the test set, outperforming all base models. The model exhi-bited especially strong diagnostic performance for damp-heat stagnating in the lung (AUC: 0.9984) and wind pathogen latent in the lung (AUC: 0.9970). ConclusionThe Voting ensemble algorithm effectively integrates the strengths of multiple machine learning models, resulting in an optimized diagnostic model for TCM syndromes in chronic cough with high accuracy and enhanced generalization ability.
9.Perceived stress and occupational burnout among hospital staff in Guangzhou tertiary hospitals
Wenli ZHOU ; Xiaoyi WU ; Yichen YE ; Liman WU ; Biyun CHEN ; Yi SHEN
Journal of Environmental and Occupational Medicine 2025;42(3):354-359
Background Staff in tertiary hospitals are a high-risk group for occupational burnout. Timely identification and precise intervention are crucial for improving healthcare service quality. However, comparative studies on perceived stress and occupational burnout among hospital staff in different positions are lacking. Objective To describe the status of perceived stress and occupational burnout among hospital staff in different positions and compare the differences, explore the relationship between perceived stress and occupational burnout, and identify the influencing factors of occupational burnout. Methods In May 2022,
10.The clinical significance of Th17 cell heterogeneity in myelodysplastic neoplasms
Yichen WANG ; Wenguang ZHOU ; Yanwen YAN ; Fang YI ; Lingsha QIN ; Wei LI ; Yuquan LI ; Xiangzong ZENG
Tianjin Medical Journal 2025;53(9):942-946
Objective To investigate the proportion of Th17 cells,Th1-like Th17 cells and FoxP3+Th17 cells in bone marrow of patients with myelodysplastic syndrome(MDS),the expression of interleukin-17A(IL-17A)in bone marrow supernatant and its clinical significance.Methods Forty MDS patients(MDS group)and 18 patients with nutritional anemia(control group)were selected.MDS patients were classified into the low blast(MDS-LB)group(19 cases)and the increased blast(MDS-IB)group(21 cases,including 11 cases of type IB1 and 10 cases of type IB2)based on morphological definition.The MDS patients were scored according to the revised International Prognostic Scoring System(IPSS-R),with 18 cases in the low-risk group(≤4.5)and 22 cases in the high-risk group(>4.5).Flow cytometry was used to detect the proportion of Th17 cells,Th1-like Th17 cells and FoxP3+Th17 cells in bone marrow of the MDS group and the control group.Enzyme-linked immunosorbent assay(ELISA)was used to detect the level of IL-17A in bone marrow supernatant of the above samples.Results The proportion of Th17 cells and the level of IL-17A were higher in patients of the MDS group than those in the control group(P<0.05).According to the median expression level of IL-17A,the MDS group was divided into the low-expression group(<13.71 ng/L,20 cases)and the high-expression group(≥13.71 ng/L,20 cases).Compared with the low-expression group,there were higher proportion of patients with blast cells<5%and low-risk patients(P<0.05)in the high-expression group.Compared with the IL-17A low-expression group,the IL-17A high-expression group had a higher proportion of patients with blast cells<5%and relatively low-risk patients(P<0.05).Compared with the low-risk patients,high-risk patients had a lower proportion of Th17 cells,IL-17A levels and Th1-like Th17 cells,and a higher proportion of FoxP3+Th17 cells(P<0.05).Compared with the MDS-LB group,the MDS-IB group had a lower proportion of Th17 cells,IL-17A levels and Th1-like Th17 cells,and a higher proportion of FoxP3+Th17 cells(P<0.05).Conclusion The proportion of Th17 cells and the level of IL-17A are significantly increased in MDS patients.The decreased proportion of Th1-like Th17 cells and the increased proportion of FoxP3+Th17 cells may be related to the increased proportion of blast cells and higher risk stratification in patients.

Result Analysis
Print
Save
E-mail