A comparative analysis of responses of seven Chinese-developed GenAI models to common pediatric urogenital diseases
10.12483/j.issn.1009-8291.2026.05.002
- VernacularTitle:7款国产GenAI大模型对小儿泌尿生殖系统3种常见疾病应答结果的对比分析
- Author:
Lei KANG
1
;
Tao GUO
1
;
Gaofeng ZHANG
1
;
Guangyu ZHU
1
;
Suhua PENG
1
;
Xi CHEN
1
;
Yuxiang WANG
1
;
Zhuangyu JIANG
1
;
Ming BAI
1
Author Information
1. Department of Urology, Xi'an Children's Hospital, Xi'an 710003, China
- Publication Type:Journal Article
- Keywords:
artificial intelligence;
Chinese-developed GenAI large models;
hypospadias;
hydronephrosis;
cryptorchidism;
medical question answering
- From:
Journal of Modern Urology
2026;31(5):404-412
- CountryChina
- Language:Chinese
-
Abstract:
Objective To compare the performance of 7 Chinese-developed GenAI models (DeepSeek, Wenxin Yiyan, Tongyi, Doubao, Kimi, Hunyuan, and iFLYTEK Spark) in responding to common diseases of the pediatric genitourinary system and to explore their potential applications.Methods A total of 80 questions covering basic disease information, diagnosis, treatment, and prognosis related to “hypospadias, hydronephrosis, and cryptorchidism” were developed as prompts.The models were tested under 3 modes:“quick thinking, ”“deep thinking, ” and “deep thinking + internet access.” The results were evaluated in a blinded review by two pediatric urologists.The responses were graded into 4 levels—A, B, C, and D—based on accuracy, completeness, and relevance.The proportions of responses rated as A and A+B were statistically analyzed and compared across the different models/modes.Results The “deep thinking” mode of DeepSeek demonstrated the best performance, with 88% of responses rated as A and 96% rated as A+B.DeepSeek (deep thinking + internet access), Doubao (quick thinking), Hunyuan (deep thinking), and Doubao (deep thinking + internet access) followed closely, also showing excellent response capabilities.Among the top 5 models/modes, significant differences were observed in the proportion of Alevel responses (P<0.05), while differences in the proportion of A+B-level responses were not statistically significant (P>0.05).Conclusion Certain Chinese-developed GenAI models, represented by DeepSeek (deep thinking), demonstrate significant application potential.However, they still cannot consistently provide recommendations rated at the A+B level (recommended grade or higher) with 100% accuracy.In clinical practice, both doctors and patients should view GenAI as an auxiliary tool for education and information acquisition rather than as a final decision-making authority.All information must be reviewed and confirmed by professional physicians.