1.TCM Data Hub: A traditional Chinese medicine data platform powered by YiYuan large language models
Chongyun ZHOU ; Qin LI ; Tangming CUI ; Chaohui CUI ; Peiyu WANG ; Meiling SUN ; Ying NIE ; Yichen BAI ; Haiyan LI
Science of Traditional Chinese Medicine 2026;4(2):140-151
The digitization of traditional Chinese medicine (TCM) has generated vast amounts of data. However, these data are characterized by significant heterogeneity and complex semantic structures, posing substantial challenges for systematic integration and intelligent analysis, and limiting its potential for modern clinical and computational research. To address the challenges posed by the high heterogeneity and complex structure in TCM data, we designed and developed the TCM Data Hub platform, which is powered by the YiYuan large language models (LLMs). This platform aims to enhance intelligent data processing capabilities and unlock the potential for clinical application of TCM data through systematic integration and efficient utilization, thereby bridging the gap between traditional knowledge and modern computational research. This study first analyzed the heterogeneity and complexity of TCM information with respect to data types, structures, and semantics. A standardized data framework was constructed to enhance data integration and interoperability. Based on the TCM Intelligent Computing Platform of the China Academy of Chinese Medical Sciences, we trained the YiYuan LLMs to acquire domain-specific semantic understanding of TCM, thereby improving the platform’s comprehension of specialized terminology and knowledge systems. Leveraging the natural language processing capabilities of the LLMs, we developed a human-in-the-loop data processing system to enable efficient extraction, cleansing, and structured organization of TCM data. In addition, utilizing Vue and Java technologies, we developed multiple LLM-powered intelligent agents and systems, including a human-in-the-loop data processing system, as well as automated prescription mining and network pharmacology analysis agents. Task-specific agents tailored to TCM data processing were developed to enhance the model’s effectiveness in clinical knowledge discovery. System functionality and platform infrastructure were implemented using Java and Vue technologies.The TCM Data Hub platform has completed system construction and core functionality implementation. It supported integrated management and efficient access to 8 key types of TCM data: prescriptions, materia medica, ingredients, targets, diseases (Western medicine), diseases (TCM), syndromes, and therapeutic methods. The human-in-the-loop data processing system achieved an accuracy of 95.34% in structuring TCM data and supported annotation for data requiring manual labeling. The intelligent agent-driven big-data analytics module enabled 1-click, end-to-end workflows for TCM prescription mining, herb-syndrome association analysis, network pharmacology, and molecular biology research, completing a full data mining task in approximately 30 minutes. Users can interact with and manipulate data through a visual front-end interface. The system demonstrated stable performance, strong scalability, and a user-friendly experience. Empowered by the YiYuan LLMs, the TCM Data Hub platform significantly improves the accessibility, usability, and intelligence of TCM data. It effectively bridges traditional TCM knowledge with modern intelligent technologies, providing robust data support and intelligent tools for TCM research and clinical applications.
2.Construction of a Diagnostic Model for Traditional Chinese Medicine Syndromes of Chronic Cough Based on the Voting Ensemble Machine Learning Algorithm
Yichen BAI ; Suyang QIN ; Chongyun ZHOU ; Liqing SHI ; Kun JI ; Chuchu ZHANG ; Panfei LI ; Tangming CUI ; Haiyan LI
Journal of Traditional Chinese Medicine 2025;66(11):1119-1127
ObjectiveTo explore the construction of a machine learning model for the diagnosis of traditional Chinese medicine (TCM) syndromes in chronic cough and the optimization of this model using the Voting ensemble algorithm. MethodsA retrospective analysis was conducted using clinical data from 921 patients with chronic cough treated at the Respiratory Department of Dongfang Hospital, Beijing University of Chinese Medicine. After standardized processing, 84 clinical features were extracted to determine TCM syndrome types. A specialized dataset for TCM syndrome diagnosis in chronic cough was formed by selecting syndrome types with more than 50 cases. The synthetic minority over-sampling technique (SMOTE) was employed to balance the dataset. Four base models, logistic regression (LR), decision tree (dt), multilayer perceptron (MLP), and Bagging, were constructed and integrated using a hard voting strategy to form a Voting ensemble model. Model performance was evaluated using accuracy, recall, precision, F1-score, receiver operating characteristic (ROC) curve, area under the curve (AUC), and confusion matrix. ResultsAmong the 921 cases, six syndrome types had over 50 cases each, phlegm-heat obstructing the lung (294 cases), wind pathogen latent in the lung (103 cases), cold-phlegm obstructing the lung (102 cases), damp-heat stagnating in the lung (64 cases), lung yang deficiency (54 cases), and phlegm-damp obstructing the lung (53 cases), yielding a total of 670 cases in the specialized dataset. High-frequency symptoms among these patients included cough, expectoration, odor-induced cough, throat itchiness, itch-induced cough, and cough triggered by cold wind. Among the four base models, the MLP model showed the best diagnostic performance (test accuracy: 0.9104; AUC: 0.9828). Compared with the base models, the Voting ensemble model achieved superior performance with an accuracy of 0.9289 on the training set and 0.9253 on the test set, showing a minimal overfitting gap of 0.0036. It also achieved the highest AUC (0.9836) in the test set, outperforming all base models. The model exhi-bited especially strong diagnostic performance for damp-heat stagnating in the lung (AUC: 0.9984) and wind pathogen latent in the lung (AUC: 0.9970). ConclusionThe Voting ensemble algorithm effectively integrates the strengths of multiple machine learning models, resulting in an optimized diagnostic model for TCM syndromes in chronic cough with high accuracy and enhanced generalization ability.

Result Analysis
Print
Save
E-mail