Verification of the Validity and Reliability of Therapeutic Communication Responses Generated by Large Language Models (LLMs): A Comparative Study of Prompt-Engineering Strategies
10.12934/jkpmhn.2025.34.S1.23
- Author:
Seong Kwang KIM
1
;
Geun-Myun KIM
;
Sunkyung CHA
;
Miran JUNG
Author Information
1. Postdoctoral Researcher, Department of Nursing, Gangneung-Wonju National University, Wonju, Korea
- Publication Type:Original Article
- From:Journal of Korean Academy of Psychiatric and Mental Health Nursing
2025;34(Special Issue):23-35
- CountryRepublic of Korea
- Language:English
-
Abstract:
Purpose:This study evaluated the validity and reliability of therapeutic communication responses generated by GPT-5 by comparing different prompt-engineering strategies and identifying optimal methods for psychiatric nursing education.
Methods:A methodological design compared four strategies—Zero-shot, Few-shot, Chain-of-Thought (CoT), and Role-prompt—applied to ten psychiatric nursing communication questions. Fourteen experts with over ten years of combined clinical and educational experience rated 40 anonymized responses using a 4-point scale. Inter-rater reliability was assessed with Fleiss' κ and Krippendorff's ⍺ from 400 repeated outputs. Semantic consistency was examined using cosine similarity. Analyses were conducted in Python.
Results:Role-prompt showed the highest content validity (mean=3.54±0.39, S-CVI/Ave=0.95), followed by CoT and Zero-shot. Inter-rater agreement was slight (κ=0.07; ⍺=0.15). Semantic consistency was highest for CoT (0.76±0.12) and Role-prompt (0.75±0.11). A Friedman test indicated significant differences among strategies (x2(3)=11.88, p=.008).
Conclusion:Role-prompt yielded the most valid therapeutic responses, while CoT produced the most consistent outputs. These findings support the use of Role-prompt and CoT strategies to enhance the validity and reliability of therapy.