1.Imaging differentiation of hepatocellular carcinoma, combined hepatocellular-cholangiocarcinoma, and intrahepatic cholangiocarcinoma: pitfalls and advances
Jaeseung SHIN ; Taek CHUNG ; Sang Yun HA ; Hyungjin RHEE
Journal of Liver Cancer 2026;26(1):9-18
Accurate non-invasive differentiation of primary liver cancers, such as hepatocellular carcinoma (HCC), intrahepatic cholangiocarcinoma (iCCA), and combined hepatocellular-cholangiocarcinoma (cHCC-CCA), is crucial for optimal management but challenging due to shared risk factors and overlapping imaging phenotypes. While the Liver Imaging Reporting and Data System category M effectively captures the classic targetoid appearance of large duct type iCCA, the small duct type frequently exhibits HCC-mimicking non-rim arterial phase hyperenhancement and non-peripheral washout, potentially compromising diagnostic specificity. Furthermore, cHCC-CCA presents a formidable diagnostic dilemma, existing on a continuous imaging spectrum that reflects its histologic dominance. This continuous imaging spectrum not only blurs radiologic distinctions but also complicates tissue sampling, limiting the diagnostic accuracy of core needle biopsies and highlighting the risk of misclassification. To enhance diagnostic clarity, this review highlights their key imaging hallmarks: while HCC typically shows non-rim arterial phase hyperenhancement (APHE) and non-peripheral washout, large duct iCCA displays a classic targetoid appearance with rim APHE and progressive central enhancement. Conversely, small duct iCCA often mimics HCC, and cHCC-CCA exhibits a variable spectrum depending on its predominant histologic component. Ultimately, overcoming these diagnostic pitfalls requires a rigorous, multidisciplinary approach that synthesizes imaging findings, serologic tumor markers, and clinical contexts.
2.Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN
Ultrasonography 2025;44(3):220-231
Purpose:
This study aimed to evaluate the diagnostic accuracy of three multimodal large language models (LLMs) in radiological image interpretation and to assess the impact of prompt engineering strategies and input conditions.
Methods:
This study analyzed 67 radiological quiz cases from the Korean Society of Ultrasound in Medicine. Three multimodal LLMs (Claude 3.5 Sonnet, GPT-4o, and Gemini-1.5-Pro-002) were evaluated using six types of prompts (basic [without system prompt], original [specific instructions], chain-of-thought, reflection, multiagent, and artificial intelligence [AI]–generated). Performance was assessed across various factors, including tumor versus non-tumor status, case rarity, difficulty, and knowledge cutoff dates. A subgroup analysis compared diagnostic accuracy between imaging-only inputs and combined imaging-descriptive text inputs.
Results:
With imaging-only inputs, Claude 3.5 Sonnet achieved the highest overall accuracy (46.3%, 186/402), followed by GPT-4o (43.5%, 175/402) and Gemini-1.5-Pro-002 (39.8%, 160/402). AI-generated prompts yielded superior combined accuracy across all three models, with significant improvements over the basic (7.96%, P=0.009), chain-of-thought (6.47%, P=0.029), and multiagent prompts (5.97%, P=0.043). The integration of descriptive text significantly enhanced diagnostic accuracy for Claude 3.5 Sonnet (46.3% to 66.2%, P<0.001), GPT-4o (43.5% to 57.5%, P<0.001), and Gemini-1.5-Pro-002 (39.8% to 60.4%, P<0.001). Model performance was significantly influenced by case rarity for GPT-4o (rare: 6.7% vs. nonrare: 53.9%, P=0.001) and by knowledge cutoff dates for Claude 3.5 Sonnet (post-cutoff: 23.5% vs. pre-cutoff: 64.0%, P=0.005).
Conclusion
Claude 3.5 Sonnet achieved the highest diagnostic accuracy in radiological quiz cases, followed by GPT-4o and Gemini-1.5-Pro-002. The use of AI-generated prompts and the integration of descriptive text inputs enhanced model performance.
3.Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN
Ultrasonography 2025;44(3):220-231
Purpose:
This study aimed to evaluate the diagnostic accuracy of three multimodal large language models (LLMs) in radiological image interpretation and to assess the impact of prompt engineering strategies and input conditions.
Methods:
This study analyzed 67 radiological quiz cases from the Korean Society of Ultrasound in Medicine. Three multimodal LLMs (Claude 3.5 Sonnet, GPT-4o, and Gemini-1.5-Pro-002) were evaluated using six types of prompts (basic [without system prompt], original [specific instructions], chain-of-thought, reflection, multiagent, and artificial intelligence [AI]–generated). Performance was assessed across various factors, including tumor versus non-tumor status, case rarity, difficulty, and knowledge cutoff dates. A subgroup analysis compared diagnostic accuracy between imaging-only inputs and combined imaging-descriptive text inputs.
Results:
With imaging-only inputs, Claude 3.5 Sonnet achieved the highest overall accuracy (46.3%, 186/402), followed by GPT-4o (43.5%, 175/402) and Gemini-1.5-Pro-002 (39.8%, 160/402). AI-generated prompts yielded superior combined accuracy across all three models, with significant improvements over the basic (7.96%, P=0.009), chain-of-thought (6.47%, P=0.029), and multiagent prompts (5.97%, P=0.043). The integration of descriptive text significantly enhanced diagnostic accuracy for Claude 3.5 Sonnet (46.3% to 66.2%, P<0.001), GPT-4o (43.5% to 57.5%, P<0.001), and Gemini-1.5-Pro-002 (39.8% to 60.4%, P<0.001). Model performance was significantly influenced by case rarity for GPT-4o (rare: 6.7% vs. nonrare: 53.9%, P=0.001) and by knowledge cutoff dates for Claude 3.5 Sonnet (post-cutoff: 23.5% vs. pre-cutoff: 64.0%, P=0.005).
Conclusion
Claude 3.5 Sonnet achieved the highest diagnostic accuracy in radiological quiz cases, followed by GPT-4o and Gemini-1.5-Pro-002. The use of AI-generated prompts and the integration of descriptive text inputs enhanced model performance.
4.Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN
Ultrasonography 2025;44(3):220-231
Purpose:
This study aimed to evaluate the diagnostic accuracy of three multimodal large language models (LLMs) in radiological image interpretation and to assess the impact of prompt engineering strategies and input conditions.
Methods:
This study analyzed 67 radiological quiz cases from the Korean Society of Ultrasound in Medicine. Three multimodal LLMs (Claude 3.5 Sonnet, GPT-4o, and Gemini-1.5-Pro-002) were evaluated using six types of prompts (basic [without system prompt], original [specific instructions], chain-of-thought, reflection, multiagent, and artificial intelligence [AI]–generated). Performance was assessed across various factors, including tumor versus non-tumor status, case rarity, difficulty, and knowledge cutoff dates. A subgroup analysis compared diagnostic accuracy between imaging-only inputs and combined imaging-descriptive text inputs.
Results:
With imaging-only inputs, Claude 3.5 Sonnet achieved the highest overall accuracy (46.3%, 186/402), followed by GPT-4o (43.5%, 175/402) and Gemini-1.5-Pro-002 (39.8%, 160/402). AI-generated prompts yielded superior combined accuracy across all three models, with significant improvements over the basic (7.96%, P=0.009), chain-of-thought (6.47%, P=0.029), and multiagent prompts (5.97%, P=0.043). The integration of descriptive text significantly enhanced diagnostic accuracy for Claude 3.5 Sonnet (46.3% to 66.2%, P<0.001), GPT-4o (43.5% to 57.5%, P<0.001), and Gemini-1.5-Pro-002 (39.8% to 60.4%, P<0.001). Model performance was significantly influenced by case rarity for GPT-4o (rare: 6.7% vs. nonrare: 53.9%, P=0.001) and by knowledge cutoff dates for Claude 3.5 Sonnet (post-cutoff: 23.5% vs. pre-cutoff: 64.0%, P=0.005).
Conclusion
Claude 3.5 Sonnet achieved the highest diagnostic accuracy in radiological quiz cases, followed by GPT-4o and Gemini-1.5-Pro-002. The use of AI-generated prompts and the integration of descriptive text inputs enhanced model performance.
5.Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN
Ultrasonography 2025;44(3):220-231
Purpose:
This study aimed to evaluate the diagnostic accuracy of three multimodal large language models (LLMs) in radiological image interpretation and to assess the impact of prompt engineering strategies and input conditions.
Methods:
This study analyzed 67 radiological quiz cases from the Korean Society of Ultrasound in Medicine. Three multimodal LLMs (Claude 3.5 Sonnet, GPT-4o, and Gemini-1.5-Pro-002) were evaluated using six types of prompts (basic [without system prompt], original [specific instructions], chain-of-thought, reflection, multiagent, and artificial intelligence [AI]–generated). Performance was assessed across various factors, including tumor versus non-tumor status, case rarity, difficulty, and knowledge cutoff dates. A subgroup analysis compared diagnostic accuracy between imaging-only inputs and combined imaging-descriptive text inputs.
Results:
With imaging-only inputs, Claude 3.5 Sonnet achieved the highest overall accuracy (46.3%, 186/402), followed by GPT-4o (43.5%, 175/402) and Gemini-1.5-Pro-002 (39.8%, 160/402). AI-generated prompts yielded superior combined accuracy across all three models, with significant improvements over the basic (7.96%, P=0.009), chain-of-thought (6.47%, P=0.029), and multiagent prompts (5.97%, P=0.043). The integration of descriptive text significantly enhanced diagnostic accuracy for Claude 3.5 Sonnet (46.3% to 66.2%, P<0.001), GPT-4o (43.5% to 57.5%, P<0.001), and Gemini-1.5-Pro-002 (39.8% to 60.4%, P<0.001). Model performance was significantly influenced by case rarity for GPT-4o (rare: 6.7% vs. nonrare: 53.9%, P=0.001) and by knowledge cutoff dates for Claude 3.5 Sonnet (post-cutoff: 23.5% vs. pre-cutoff: 64.0%, P=0.005).
Conclusion
Claude 3.5 Sonnet achieved the highest diagnostic accuracy in radiological quiz cases, followed by GPT-4o and Gemini-1.5-Pro-002. The use of AI-generated prompts and the integration of descriptive text inputs enhanced model performance.
6.Diagnostic performance of multimodal large language models in radiological quiz cases: the effects of prompt engineering and input conditions
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN
Ultrasonography 2025;44(3):220-231
Purpose:
This study aimed to evaluate the diagnostic accuracy of three multimodal large language models (LLMs) in radiological image interpretation and to assess the impact of prompt engineering strategies and input conditions.
Methods:
This study analyzed 67 radiological quiz cases from the Korean Society of Ultrasound in Medicine. Three multimodal LLMs (Claude 3.5 Sonnet, GPT-4o, and Gemini-1.5-Pro-002) were evaluated using six types of prompts (basic [without system prompt], original [specific instructions], chain-of-thought, reflection, multiagent, and artificial intelligence [AI]–generated). Performance was assessed across various factors, including tumor versus non-tumor status, case rarity, difficulty, and knowledge cutoff dates. A subgroup analysis compared diagnostic accuracy between imaging-only inputs and combined imaging-descriptive text inputs.
Results:
With imaging-only inputs, Claude 3.5 Sonnet achieved the highest overall accuracy (46.3%, 186/402), followed by GPT-4o (43.5%, 175/402) and Gemini-1.5-Pro-002 (39.8%, 160/402). AI-generated prompts yielded superior combined accuracy across all three models, with significant improvements over the basic (7.96%, P=0.009), chain-of-thought (6.47%, P=0.029), and multiagent prompts (5.97%, P=0.043). The integration of descriptive text significantly enhanced diagnostic accuracy for Claude 3.5 Sonnet (46.3% to 66.2%, P<0.001), GPT-4o (43.5% to 57.5%, P<0.001), and Gemini-1.5-Pro-002 (39.8% to 60.4%, P<0.001). Model performance was significantly influenced by case rarity for GPT-4o (rare: 6.7% vs. nonrare: 53.9%, P=0.001) and by knowledge cutoff dates for Claude 3.5 Sonnet (post-cutoff: 23.5% vs. pre-cutoff: 64.0%, P=0.005).
Conclusion
Claude 3.5 Sonnet achieved the highest diagnostic accuracy in radiological quiz cases, followed by GPT-4o and Gemini-1.5-Pro-002. The use of AI-generated prompts and the integration of descriptive text inputs enhanced model performance.
8.Comparison of micro-flow imaging and contrast-enhanced ultrasonography in assessing segmental congestion after right living donor liver transplantation
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN ; Dong Ik CHA ; Kyowon GU ; Jinsoo RHU ; Jong Man KIM ; Gyu-Seong CHOI
Ultrasonography 2024;43(6):469-477
Purpose:
This study aimed to determine whether micro-flow imaging (MFI) offers diagnostic performance comparable to that of contrast-enhanced ultrasonography (CEUS) in detecting segmental congestion among patients undergoing living donor liver transplantation (LDLT).
Methods:
Data from 63 patients who underwent LDLT between May and December 2022 were retrospectively analyzed. MFI and CEUS data collected on the first postoperative day were quantified. Segmental congestion was assessed based on imaging findings and laboratory data, including liver enzymes and total bilirubin levels. The reference standard was a postoperative contrast-enhanced computed tomography scan performed within 2 weeks of surgery. Additionally, a subgroup analysis examined patients who underwent reconstruction of the middle hepatic vein territory.
Results:
The sensitivity and specificity of MFI were 73.9% and 67.5%, respectively. In comparison, CEUS demonstrated a sensitivity of 78.3% and a specificity of 75.0%. These findings suggest comparable diagnostic performance, with no significant differences in sensitivity (P=0.655) or specificity (P=0.257) between the two modalities. Additionally, early postoperative laboratory values did not show significant differences between patients with and without congestion. The subgroup analysis also indicated similar diagnostic performance between MFI and CEUS.
Conclusion
MFI without contrast enhancement yielded results comparable to those of CEUS in detecting segmental congestion after LDLT. Therefore, MFI may be considered a viable alternative to CEUS.
9.Comparison of micro-flow imaging and contrast-enhanced ultrasonography in assessing segmental congestion after right living donor liver transplantation
Taewon HAN ; Woo Kyoung JEONG ; Jaeseung SHIN ; Dong Ik CHA ; Kyowon GU ; Jinsoo RHU ; Jong Man KIM ; Gyu-Seong CHOI
Ultrasonography 2024;43(6):469-477
Purpose:
This study aimed to determine whether micro-flow imaging (MFI) offers diagnostic performance comparable to that of contrast-enhanced ultrasonography (CEUS) in detecting segmental congestion among patients undergoing living donor liver transplantation (LDLT).
Methods:
Data from 63 patients who underwent LDLT between May and December 2022 were retrospectively analyzed. MFI and CEUS data collected on the first postoperative day were quantified. Segmental congestion was assessed based on imaging findings and laboratory data, including liver enzymes and total bilirubin levels. The reference standard was a postoperative contrast-enhanced computed tomography scan performed within 2 weeks of surgery. Additionally, a subgroup analysis examined patients who underwent reconstruction of the middle hepatic vein territory.
Results:
The sensitivity and specificity of MFI were 73.9% and 67.5%, respectively. In comparison, CEUS demonstrated a sensitivity of 78.3% and a specificity of 75.0%. These findings suggest comparable diagnostic performance, with no significant differences in sensitivity (P=0.655) or specificity (P=0.257) between the two modalities. Additionally, early postoperative laboratory values did not show significant differences between patients with and without congestion. The subgroup analysis also indicated similar diagnostic performance between MFI and CEUS.
Conclusion
MFI without contrast enhancement yielded results comparable to those of CEUS in detecting segmental congestion after LDLT. Therefore, MFI may be considered a viable alternative to CEUS.
10.Image Quality and Focal Lesion Detectability Analysis of Multiband Variable-Rate Selective Excitation Diffusion-Weighted Imaging of the Liver Using 3.0-T MRI
Ja Kyung YOON ; Yong Eun CHUNG ; Jaeseung SHIN ; Eunju KIM ; Nieun SEO ; Jin-Young CHOI ; Mi-Suk PARK ; Myeong-Jin KIM
Investigative Magnetic Resonance Imaging 2024;28(1):8-17
Purpose:
Acquisition time reduction in diffusion-weighted imaging (DWI) can be achieved by the combining multiband and variable-rate selective excitation (MB-VERSE). This study attempted to evaluate and compare the image quality (IQ) and focal lesion detectability of the respiratory-triggered MB-VERSE DWI with conventional DWI for liver magnetic resonance imaging.
Materials and Methods:
The acquisition time, IQ, and focal lesion detectability of MBVERSE DWI and conventional DWI were compared in 144 patients. Qualitative (overall IQ, IQ at the liver dome, sharpness of the liver margin, and degree of artifacts) and quantitative (signal-to-noise ratio [SNR], contrast-to-noise ratio [CNR], and apparent diffusion co efficient) IQ parameters were compared with the Wilcoxon signed-rank test. The diagnostic accuracy for focal lesion detectability was estimated with the mean figure of merit (FOM) from the area under the jackknife alternative free-response receiver operating characteristic curve.
Results:
The MB-VERSE DWI exhibited significantly shorter scan time (153.1 ± 34.5 s vs.225.1 ± 33.0 s, p < 0.001), poorer qualitative IQ (3.4 vs. 3.9, p < 0.001), lower SNR (34.4 vs. 50.0, p < 0.001), but comparable CNR (57.5 ± 49.0 vs. 78.9 ± 75.6, p = 0.070) compared to those of the conventional DWI. The MB-VERSE DWI exhibited similar per-lesion sensitivities (85.1%–88.1% vs. 88.1%–92.5%) and specificities (99.7%–99.8% vs. 99.5%–99.8%) of focal lesion detectability (p > 0.050) and similar diagnostic accuracy (FOM, 0.958 vs.0.957, p = 0.583) compared to those of the conventional DWI.
Conclusion
MB-VERSE DWI exhibited a significantly shorter acquisition time than conventional DWI, with compromised overall IQ and lower SNR but preserved CNR and focal liver lesion detectability. MB-VERSE DWI may be a useful alternative for patients requiring a short acquisition time.

Result Analysis
Print
Save
E-mail