1.Development and validation of PhenoRAG: A visualization tool for automated human phenotype ontology term annotation based on large language models and retrieval-augmented generation technology.
Wei ZHONG ; Yousheng YAN ; Kai YANG ; Yan LIU ; Xinyu FU ; Zhengyang YAO ; Chenghong YIN
Chinese Journal of Medical Genetics 2026;43(1):36-43
OBJECTIVE:
To develop a user-friendly visualization application for the automatic annotation of Human Phenotype Ontology (HPO) terms based on large language models and retrieval-augmented generation (RAG) technology, and to validate its performance in an authoritative case dataset.
METHODS:
By integrating the domestic open-source large language model DeepSeek-V3 with RAG technology, an interactive web application was deployed on the Streamlit cloud platform. Using only the latest official HPO dataset as the data source, the lightweight sentence-embedding model BAAI/bge-small-en-v1.5 was employed to construct a FAISS vector index. During the online phase, a four-step closed-loop process is automatically completed: multilingual translation, phenotype phrase extraction, RAG candidate retrieval, term mapping, and official database validation. 121 English case reports publicly released by BMJ Case Reports and Oxford Medical Case Reports (with a gold-standard HPO set of 1 794 terms) were selected for application validation. Precision, recall, and F1 score were calculated and compared horizontally with traditional dictionary tools, standalone large language models, and the similar application "RAG-HPO". Finally, replace the model with the more advanced ChatGPT-5 and evaluate its performance on the newly extracted dataset.
RESULTS:
An HPO term automatic annotation visualization application named PhenoRAG, based on large language models and RAG technology, was successfully developed. Users can access it directly via a web link. Across the 112 cases, a total of 2 150 HPO terms were generated; 2,064 (96.0%) were fully validated by the official database, with a hallucination rate of 1.3% and an HPO ID-name mismatch rate of 2.7%. After deduplication, 1,906 terms remained for testing. The overall precision was 63.65%, recall was 67.34%, and F1 was 65.44%, significantly outperforming traditional annotation tools (F1: 0.45-0.49, P < 0.001). Although PhenoRAG's F1 was lower than that of RAG-HPO (F1 = 0.78, P < 0.001), which relies on a manually constructed synonym database of 54 000 entries plus the HPO dataset, it requires no additional dictionary maintenance and can be used without any background in computer programming. Moreover, after switching to the GPT-5 model, PhenoRAG exhibited no hallucination rate on the new dataset, and its F1 score significantly increased (P = 0.038).
CONCLUSION
Without constructing a synonym database, the PhenoRAG achieved high-accuracy automatic mapping from clinical text to standard HPO terms. It features a low usage threshold, free access, and a Chinese-language interface, and can directly serve rare disease diagnosis, genetic counseling, and research scenarios in China and worldwide, warranting further clinical promotion and multicenter validation.
Humans
;
Phenotype
;
Biological Ontologies
;
Language
;
Software
;
Large Language Models
2.External ocular manifestations among patients diagnosed with Coronavirus disease 2019 in a referral center in the Philippines.
Alyssa Louise B. Pejana-Paulino ; Aramis B. Torrefranca Jr. ; Nilo Vincent DG. Florcruz ; Ma. Dominga B. Padilla
Acta Medica Philippina 2026;60(1):69-77
BACKGROUND AND OBJECTIVES
data-mce-style="text-align: justify;">The global pandemic caused by Coronavirus Disease 2019 (COVID-19) has affected millions, with growing evidence of the potential role of ocular tissues in viral transmission. At the time of writing, local data regarding the phenomenon was limited. This study investigated external ocular manifestations in patients with COVID-19 at a referral center in the Philippines, examined correlations between demographics, systemic manifestations, and laboratory results with ocular manifestations, and determined their timing relative to systemic symptoms.
METHODSdata-mce-style="text-align: justify;">This single-center, descriptive cross-sectional study was carried out from December 8 to 18, 2020 at the adult COVID-19 wards of the Philippine General Hospital involving 72 participants. Data collection involved relevant clinical history taking and performing gross eye examination. The prevalence of ocular manifestations was described with 95% confidence intervals. Correlations between ocular manifestations and quantitative variables were analyzed with point-biserial correlation, and associations with qualitative variables were tested using chi-square or Fisher’s exact tests.
RESULTSdata-mce-style="text-align: justify;">Among participants, 31.9% presented with ocular manifestations with foreign body sensation as the most prevalent ocular symptom (11.1%) and conjunctival hyperemia as the most prevalent ocular finding (19.4%). The median age of patients with ocular manifestations was 41 years old with a higher prevalence in the male population (73.9%, CI=95%, p=0.001). No significant correlation was observed between presence of external ocular manifestations and the different systemic and ocular co-morbidities as well as with COVID-19 clinical classification. Among those who experienced symptoms, majority (29.2%) of the patients experienced systemic symptoms prior to the onset of ocular symptoms. Ocular complaints may present as the sole manifestation (13.9%). Several laboratory parameters were measured and only temperature and AST levels showed a low positive correlation with the presence of ocular manifestations.
CONCLUSIONdata-mce-style="text-align: justify;">Ocular manifestations occur in roughly one third of patients with COVID-19 based on this study population. With some individuals presenting with ocular signs or symptoms as the initial and sole manifestation, healthcare practitioners must exercise caution and remain vigilant in managing patients who present as such. At the time of writing, this is the first local study investigating the different external ocular manifestations in patients with COVID-19. There is a need to pursue more robust studies and conduct more local investigations which will guide both ophthalmologists and other practitioners in strengthening existing guidelines regarding precautionary practices, clinical diagnosis, and management of COVID-19 patients.
Human ; Sars-cov-2 ; Covid-19 ; Philippines ; Adult ; Association ; Classification ; Collection ; Confidence Intervals ; Coronavirus ; Cross-sectional Studies ; Data Collection ; Demography ; Diagnosis ; Disease ; Exercise ; Eye ; Foreign Bodies ; History ; Hospitals ; Hospitals, General ; Hyperemia ; Laboratories ; Male ; Morbidity ; Ophthalmologists ; Pandemics ; Patients ; Population ; Prevalence ; Referral And Consultation ; Role ; Sensation ; Temperature ; Time ; Tissues ; Volition ; World Health Organization ; Writing
3.Common detoxification mechanisms in processing of toxic medicinal herbs of the same genus: a case study of Euphorbia pekinensis, E. ebracteolata, and E. fischeriana.
En-Ci JIANG ; Hong-Li YU ; Shu-Rui ZHANG ; Bing-Bing LIU ; Xin-Zhi WANG ; Hao WU
China Journal of Chinese Materia Medica 2025;50(13):3615-3675
Traditional Chinese medicine(TCM) processing is a specialized pharmaceutical technique with the primary objective of reducing the toxicity of medicinal substances. Euphorbia pekinensis, E. ebracteolata, and E. fischeriana, all belonging to Euphorbiaceae, are classified as drastic purgative herbs, traditionally used for eliminating retained water, reducing swelling, resolving toxicity, and dispersing masses. However, these herbs are also associated with adverse effects such as abdominal pain and diarrhea. Accordingly, they are commonly processed with vinegar, milk, or Terminalia chebula decoction to reduce the toxicity. This review summarizes the chemical constituents, pharmacological activities, historical evolution of processing methods, and detoxification mechanisms of the three toxic Euphorbia species. The primary toxic constituents are terpenoids. Specifically, E. ebracteolata and E. fischeriana are rich in diterpenoids, while E. pekinensis contains diterpenoids, triterpenoids, and sesquiterpenoids. Studies have shown that vinegar processing promotes structural transformations of diterpenoids, including ether bond hydrolysis, lactone ring opening, esterification, oxidation, and epoxide ring cleavage, thereby reducing the content and toxicity of these compounds. Milk processing facilitates the dissolution of toxic components into the residual liquid of excipients, leading to decreases in their concentrations in the final decoction pieces. Processing with T. chebula decoction raises the levels of tannin-derived phenolic acids, which antagonize the adverse effects of the intestine. These findings reveal a shared detoxification pattern among the three toxic herbs. Accordingly, this review proposes the concept of a shared detoxification mechanism for toxic herbs belonging to the same family or genus. That is, toxic herbs belonging to the same taxon often exhibit similar toxicological profiles and can undergo detoxification through the same processing methods, reflecting common underlying mechanisms. Investigating such shared mechanisms across multiple species of the same genus offers a promising research strategy. Ultimately, the research into processing-induced detoxification mechanisms provides both theoretical and practical support for ensuring the safety of toxic TCM.
Euphorbia/classification*
;
Drugs, Chinese Herbal/metabolism*
;
Humans
;
Animals
;
Inactivation, Metabolic
;
Medicine, Chinese Traditional
4.Identification and expression analysis of AP2/ERF family members in Lonicera macranthoides.
Si-Min ZHOU ; Mei-Ling QU ; Juan ZENG ; Jia-Wei HE ; Jing-Yu ZHANG ; Zhi-Hui WANG ; Qiao-Zhen TONG ; Ri-Bao ZHOU ; Xiang-Dan LIU
China Journal of Chinese Materia Medica 2025;50(15):4248-4262
The AP2/ERF transcription factor family is a class of transcription factors widely present in plants, playing a crucial role in regulating flowering, flower development, flower opening, and flower senescence. Based on transcriptome data from flower, leaf, and stem samples of two Lonicera macranthoides varieties, 117 L. macranthoides AP2/ERF family members were identified, including 14 AP2 subfamily members, 61 ERF subfamily members, 40 DREB subfamily members, and 2 RAV subfamily members. Bioinformatics and differential gene expression analyses were performed using NCBI, ExPASy, SOMPA, and other platforms, and the expression patterns of L. macranthoides AP2/ERF transcription factors were validated via qRT-PCR. The results indicated that the 117 LmAP2/ERF members exhibited both similarities and variations in protein physicochemical properties, AP2 domains, family evolution, and protein functions. Differential gene expression analysis revealed that AP2/ERF transcription factors were primarily differentially expressed in the flowers of the two L. macranthoides varieties, with the differentially expressed genes mainly belonging to the ERF and DREB subfamilies. Further analysis identified three AP2 subfamily genes and two ERF subfamily genes as potential regulators of flower development, two ERF subfamily genes involved in flower opening, and two ERF subfamily genes along with one DREB subfamily gene involved in flower senescence. Based on family evolution and expression analyses, it is speculated that AP2/ERF transcription factors can regulate flower development, opening, and senescence in L. macranthoides, with ERF subfamily genes potentially serving as key regulators of flowering duration. These findings provide a theoretical foundation for further research into the specific functions of the AP2/ERF transcription factor family in L. macranthoides and offer important theoretical insights into the molecular mechanisms underlying floral phenotypic differences among its varieties.
Plant Proteins/chemistry*
;
Gene Expression Regulation, Plant
;
Transcription Factors/chemistry*
;
Lonicera/classification*
;
Flowers/metabolism*
;
Phylogeny
;
Gene Expression Profiling
;
Multigene Family
5.Textual study of Baihuasheshecao (Hedyotis diffusa).
Dong-Min JIANG ; Chu-Chu ZHONG ; Pang-Chui SHAW ; Bik-San LAU ; Tai-Wai LAU ; Guang-Hao XU ; Ying ZHANG ; Zhi-Guo MA ; Hui CAO ; Meng-Hua WU
China Journal of Chinese Materia Medica 2025;50(15):4386-4396
Baihuasheshecao(Hedyotis diffusa) is a commonly used traditional Chinese medicine derived from the whole herb of H. diffusa and has been widely utilized in folk medicine. It possesses anti-tumor, antibacterial, and anti-inflammatory properties, making it one of the frequently used herbs in TCM clinical practice. However, Shuixiancao(H. corymbosa) and Xianhuaercao(H. tenelliflora), species of the same genus, are often used as substitutes for Baihuasheshecao. To substantiate the medicinal basis of Baihuasheshecao, this study systematically reviewed classical herbal texts and modern literature, examining its nomenclature, botanical origin, harvesting, processing, properties, meridian tropism, pharmacological effects, and clinical applications. The results indicate that Baihuasheshecao was initially recorded as "Shuixiancao" in Preface to the Indexes to the Great Chinese Botany(Zhi Wu Ming Shi Tu Kao). Based on its morphological characteristics and habitat description, it was identified as H. diffusa in the Rubiaceae family. Subsequent records predominantly refer to it as Baihuasheshecao as its official name. In most regions, Baihuasheshecao is recognized as the authentic medicinal material, distinct from Shuixiancao and Xianhuaercao. Baihuasheshecao is harvested in late summer and early autumn, and the dried whole plant, including its roots, is used medicinally. The standard processing method involves cutting. It is known for its effects in clearing heat, removing toxins, reducing swelling and pain, and promoting diuresis to resolve abscesses. Initially, it was mainly used for treating appendicitis, intestinal abscesses, and venomous snake bites, and later, it became a treatment for cancer. The excavation of its clinical value followed a process in which overseas Chinese introduced the herb from Chinese folk medicine to other countries. After its unique anti-cancer effects were recognized abroad, it was reintroduced to China and gradually became a crucial TCM for cancer treatment. The findings of this study help clarify the historical and contemporary uses of Baihuasheshecao, providing literature support and a scientific basis for its rational development and precise clinical application.
Humans
;
China
;
Drugs, Chinese Herbal/chemistry*
;
Hedyotis/classification*
;
Medicine, Chinese Traditional/history*
6.Quantitative analysis of spatial distribution patterns and formation factors of medicinal plant resources in Anhui province.
Yong-Fei YIN ; Ke ZHANG ; Zhi-Xian JING ; Dai-Yin PENG ; Xiao-Bo ZHANG
China Journal of Chinese Materia Medica 2025;50(16):4584-4592
Analyzing the spatial distribution pattern and formation factors of medicinal plant resources can provide a scientific basis for the protection and development of traditional Chinese medicine(TCM) resources. This study is based on the survey data of medicinal plant resources in 104 county-level administrative regions of Anhui province in the Fourth National Survey of TCM Resources. The global spatial autocorrelation analysis, trend surface analysis, local spatial autocorrelation analysis, hotspot analysis, and a geodetector were employed to analyze the spatial distribution pattern of medicinal plant richness, and its relationship with natural factors was explored. The results can provide a basis for the formulation of development strategies such as the protection and utilization of TCM resources, as well as offer a scientific foundation for the establishment of regional planning schemes for TCM resources in Anhui province. The results indicated that the richness of medicinal plant resources in Anhui province had significant spatial heterogeneity, exhibiting highly clustered distribution characteristics. Cold spots and hot spots presented clustered distribution patterns, with cold spots mostly located north of the Huaihe River and hot spots south of the Yangtze River. Overall, the distribution of medicinal plant resources in Anhui province showed an overall trend of high in the south and low in the north, which was consistent with the overall geomorphic trend of this province. In addition, natural factors such as altitude, precipitation, and vegetation type played an important role in the diversity and spatial distribution pattern formation of medicinal plant resources. The extraction and analysis of the spatial distribution characteristics of natural factors in cold and hot spot regions discovered that the heterogeneity of eco-environments constituted a fundamental condition for the formation of species diversity.
Plants, Medicinal/classification*
;
China
;
Spatial Analysis
;
Conservation of Natural Resources
;
Biodiversity
7.Quality evaluation of Xinjiang Rehmannia glutinosa and Rehmannia glutinosa based on fingerprint and multi-component quantification combined with chemical pattern recognition.
Pan-Ying REN ; Wei ZHANG ; Xue LIU ; Juan ZHANG ; Cheng-Fu SU ; Hai-Yan GONG ; Chun-Jing YANG ; Jing-Wei LEI ; Su-Qing ZHI ; Cai-Xia XIE
China Journal of Chinese Materia Medica 2025;50(16):4630-4640
The differences in chemical quality characteristics between Xinjiang Rehmannia glutinosa and R. glutinosa were analyzed to provide a theoretical basis for the introduction and quality control of R. glutinosa. In this study, the high performance liquid chromatography(HPLC) fingerprints of 6 batches of Xinjiang R. glutinosa and 10 batches of R. glutinosa samples were established. The content of iridoid glycosides, phenylethanoid glycosides, monosaccharides, oligosaccharides, and polysaccharides in Xinjiang R. glutinosa and R. glutinosa was determined by high performance liquid chromatography-diode array detection(HPLC-DAD), high performance liquid chromatography-evaporative light scattering detection(HPLC-ELSD), and ultraviolet-visible spectroscopy(UV-Vis). The determination results were analyzed with by chemical pattern recognition and entropy weight TOPSIS method. The results showed that there were 19 common peaks in the HPLC fingerprints of the 16 batches of R. glutinosa, and catalpol, aucubin, rehmannioside D, rehmannioside A, hydroxytyrosol, leonuride, salidroside, cistanoside A, and verbascoside were identified. Hierarchical cluster analysis(HCA) and principal component analysis(PCA) showed that Qinyang R. glutinosa, Mengzhou R. glutinosa, and Xinjiang R. glutinosa were grouped into three different categories, and eight common components causing the chemical quality difference between Xinjiang R. glutinosa and R. glutinosa in Mengzhou and Qinyang of Henan province were screened out by orthogonal partial least squares discriminant analysis(OPLS-DA). The results of content determination showed that there were glucose, sucrose, raffinose, stachyose, polysaccharides, and nine glycosides in Xinjiang R. glutinosa and R. glutinosa samples, and the content of catalpol, rehmannioside A, leonuride, cistanoside A, verbascoside, sucrose, and glucose was significantly different between Xinjiang R. glutinosa and R. glutinosa. The analysis with entropy weight TOPSIS method showed that the comprehensive quality of R. glutinosa in Mengzhou and Qinyang of Henan province was better than that of Xinjiang R. glutinosa. In conclusion, the types of main chemical components of R. glutinosa and Xinjiang R. glutinosa were the same, but their content was different. The chemical quality of R. glutinosa was better than Xinjiang R. glutinosa, and other components in R. glutinosa from two producing areas and their effects need further study.
Rehmannia/classification*
;
Drugs, Chinese Herbal/chemistry*
;
Chromatography, High Pressure Liquid/methods*
;
Quality Control
8.Research on arrhythmia classification algorithm based on adaptive multi-feature fusion network.
Mengmeng HUANG ; Mingfeng JIANG ; Yang LI ; Xiaoyu HE ; Zefeng WANG ; Yongquan WU ; Wei KE
Journal of Biomedical Engineering 2025;42(1):49-56
Deep learning method can be used to automatically analyze electrocardiogram (ECG) data and rapidly implement arrhythmia classification, which provides significant clinical value for the early screening of arrhythmias. How to select arrhythmia features effectively under limited abnormal sample supervision is an urgent issue to address. This paper proposed an arrhythmia classification algorithm based on an adaptive multi-feature fusion network. The algorithm extracted RR interval features from ECG signals, employed one-dimensional convolutional neural network (1D-CNN) to extract time-domain deep features, employed Mel frequency cepstral coefficients (MFCC) and two-dimensional convolutional neural network (2D-CNN) to extract frequency-domain deep features. The features were fused using adaptive weighting strategy for arrhythmia classification. The paper used the arrhythmia database jointly developed by the Massachusetts Institute of Technology and Beth Israel Hospital (MIT-BIH) and evaluated the algorithm under the inter-patient paradigm. Experimental results demonstrated that the proposed algorithm achieved an average precision of 75.2%, an average recall of 70.1% and an average F 1-score of 71.3%, demonstrating high classification accuracy and being able to provide algorithmic support for arrhythmia classification in wearable devices.
Humans
;
Arrhythmias, Cardiac/diagnosis*
;
Algorithms
;
Electrocardiography/methods*
;
Neural Networks, Computer
;
Signal Processing, Computer-Assisted
;
Deep Learning
;
Classification Algorithms
9.A study on heart sound classification algorithm based on improved Mel-frequency cepstrum coefficient feature extraction and deep Transformer.
Journal of Biomedical Engineering 2025;42(5):1012-1020
Heart sounds are critical for early detection of cardiovascular diseases, yet existing studies mostly focus on traditional signal segmentation, feature extraction, and shallow classifiers, which often fail to sufficiently capture the dynamic and nonlinear characteristics of heart sounds, limit recognition of complex heart sound patterns, and are sensitive to data imbalance, resulting in poor classification performance. To address these limitations, this study proposes a novel heart sound classification method that integrates improved Mel-frequency cepstral coefficients (MFCC) for feature extraction with a convolutional neural network (CNN) and a deep Transformer model. In the preprocessing stage, a Butterworth filter is applied for denoising, and continuous heart sound signals are directly processed without segmenting the cardiac cycles, allowing the improved MFCC features to better capture dynamic characteristics. These features are then fed into a CNN for feature learning, followed by global average pooling (GAP) to reduce model complexity and mitigate overfitting. Lastly, a deep Transformer module is employed to further extract and fuse features, completing the heart sound classification. To handle data imbalance, the model uses focal loss as the objective function. Experiments on two public datasets demonstrate that the proposed method performs effectively in both binary and multi-class classification tasks. This approach enables efficient classification of continuous heart sound signals, provides a reference methodology for future heart sound research for disease classification, and supports the development of wearable devices and home monitoring systems.
Heart Sounds/physiology*
;
Humans
;
Algorithms
;
Neural Networks, Computer
;
Signal Processing, Computer-Assisted
;
Deep Learning
;
Cardiovascular Diseases/diagnosis*
;
Classification Algorithms
10.Enrichment Analysis and Deep Learning in Biomedical Ontology: Applications and Advancements.
Hong-Yu FU ; Yang-Yang LIU ; Mei-Yi ZHANG ; Hai-Xiu YANG
Chinese Medical Sciences Journal 2025;40(1):45-56
Biomedical big data, characterized by its massive scale, multi-dimensionality, and heterogeneity, offers novel perspectives for disease research, elucidates biological principles, and simultaneously prompts changes in related research methodologies. Biomedical ontology, as a shared formal conceptual system, not only offers standardized terms for multi-source biomedical data but also provides a solid data foundation and framework for biomedical research. In this review, we summarize enrichment analysis and deep learning for biomedical ontology based on its structure and semantic annotation properties, highlighting how technological advancements are enabling the more comprehensive use of ontology information. Enrichment analysis represents an important application of ontology to elucidate the potential biological significance for a particular molecular list. Deep learning, on the other hand, represents an increasingly powerful analytical tool that can be more widely combined with ontology for analysis and prediction. With the continuous evolution of big data technologies, the integration of these technologies with biomedical ontologies is opening up exciting new possibilities for advancing biomedical research.
Deep Learning
;
Biological Ontologies
;
Humans
;
Big Data
;
Biomedical Research


Result Analysis
Print
Save
E-mail