Comparison of Automatic Evaluation Methods for Pattern Hallucinations of Large Language Models in Traditional Chinese Medicine Syndrome Differentiation
10.13288/j.11-2166/r.2026.17.008
- VernacularTitle:面向中医辨证任务的大语言模型证型幻觉自动评测方法比较研究
- Author:
Qinwei WU
1
;
Yuzhu GAO
1
;
Xingyue GOU
1
;
Junyu YAO
1
;
Chuangan ZHOU
1
;
Zhengchun XUE
1
;
Zhirong XU
1
;
Xinlin CHEN
2
;
Dong CAO
1
Author Information
1. School of Medical Information Engineering,Guangzhou University of Chinese Medicine,Guangzhou,510006
2. School of Basic Medical Sciences,Guangzhou University of Chinese Medicine
- Publication Type:Journal Article
- Keywords:
traditional Chinese medicine syndrome differentiation;
large language models;
syndrome elements;
machine learning;
multilayer perceptron;
text similarity;
Jaccard coefficient
- From:
Journal of Traditional Chinese Medicine
2026;67(17):1845-1852
- CountryChina
- Language:Chinese
-
Abstract:
ObjectiveTo compare the performance of different automatic evaluation methods for detecting pattern hallucinations generated by large language models (LLMs) in traditional Chinese medicine (TCM) syndrome differentiation and to identify a strategy with higher overall discriminative performance against expert manual judgment. MethodsA standardized TCM syndrome-differentiation dataset containing 598 cases was constructed from case reports published in Chinese core journals indexed by Peking University Core Journals or Chinese Science Citation Database. Each of the 598 cases was submitted to five Chinese LLMs, with each model generating one syndrome-pattern output per case, yielding 2990 outputs in total. The 598 outputs generated by Qwen3-Next-80B-A3B-Thinking were independently annotated for syndrome-pattern hallucinations by two licensed TCM physicians with intermediate or higher professional titles. Disagreements were adjudicated by a third licensed TCM physician with a senior associate professional title and more than 10 years of clinical experience. The resulting consensus annotations served as the reference labels for training and validating the automated evaluation method. Based on these reference labels, three categories of automated evaluation methods were developed and compared following a progressive strategy from text-level semantic similarity, to syndrome-element structural consistency, and finally multi-feature fusion. Method A employed sentence-embedding models based on semantic similarity between pattern texts. Method B adopted a rule-based threshold method using weighted Jaccard coefficients of disease-location and disease-nature syndrome elements. Method C utilized a machine-learning classifier integrating syndrome-element matching scores, missing and redundant syndrome-element counts, and multiple semantic-similarity features. Five-fold cross-validation was used to evaluate the discriminative performance of each method against manual judgment. The optimal method was then applied to uniformly assess the pattern hallucination rates of five LLMs. ResultsIn method A, text2vec-large-chinese showed best overall performance, with an area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (PR-AUC), F1-score, sensitivity, and specificity of 0.849, 0.884, 0.794, 0.801, and 0.691, respectively. In method B, the corresponding metrics of the rule-based syndrome-element Jaccard threshold method were 0.848, 0.851, 0.817, 0.845, and 0.684, respectively; those of the LLM-based syndrome-element Jaccard method were 0.872, 0.906, 0.795, 0.764, and 0.780, respectively. In method C, the multilayer perceptron (MLP) classifier showed the best overall performance, with AUROC, PR-AUC, F1-score, sensitivity, and specificity of 0.933, 0.951, 0.871, 0.857, and 0.846, respectively. Under the unified evaluation framework combining DeepSeek-R1 syndrome-element extraction and MLP classifier, the estimated pattern hallucination rates of the five LLMs ranged from 56.0% to 77.4%. ConclusionAmong the three automatic evaluation methods, Method C integrating syndrome-element discrepancy features and text semantic features using an MLP classifier showed the highest overall discriminative performance against manual judgment, and is more suitable for evaluating syndrome hallucinations in LLM-based TCM syndrome differentiation tasks.