Comparative Performance of Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in Simplifying Turkish Abdominal Ultrasound ReportsHasan Eryeşil1, Halil Ibrahim Altunbulak1, Yusuf Eryeşil21Clinic Of Radiology, Tatvan State Hospital, Bitlis, Türkiye 2Department Of Computer Engineering, Faculty Of Technology, Selcuk University, Konya, Türkiye
AIM: To evaluate and compare the performance of three advanced Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in simplifying real-world Turkish abdominal ultrasound reports to enhance patient comprehension. MATRERIAL and METHOD: This retrospective, cross-sectional study included 100 adult abdominal ultrasound reports categorized by finding severity (normal, benign, and suspected malignancy/emergencies). Original reports were simplified to a 5th-grade reading level using standardized prompts for each model. Readability was measured quantitatively using the Ateşman Readability Formula. Two independent radiologists qualitatively assessed the outputs for medical accuracy, inclusiveness, readability, and hallucination rates RESULTS: All AI models significantly improved readability from a baseline Ateşman score of 69.53 ± 5.51 to a range of 73.24–82.19 (p < 0.001). Claude 3.5 Sonnet achieved the highest quantitative readability score (82.19 ± 5.51). However, qualitative expert evaluation revealed ChatGPT and Gemini were superior to Claude in inclusiveness and conversational readability (p < 0.001). ChatGPT (4.82 ± 0.45) specifically excelled in integrating normal anatomical findings without losing clinical context. Hallucination rates were minimal across all models (1.3% - 2.0%, p=0.565). CONCLUSION: LLMs demonstrate considerable potential in simplifying morphologically complex Turkish radiology reports while main taining clinical accuracy. While Claude 3.5 Sonnet shows strong performance in formula-based metrics, ChatGPT 5.0 and Gemini 3.0 tend to offer more favorable clinical fluency and inclusiveness. Despite high performance, rare hallucination risks such as “normality faking” necessitate a human-in-the-loop approach for patient safety
Keywords: Large Language Models, Natural Language Processing, Patient-Centered Care, Ultrasonography, Health Literacy
Hasan Eryeşil, Halil Ibrahim Altunbulak, Yusuf Eryeşil. Comparative Performance of Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in Simplifying Turkish Abdominal Ultrasound Reports. Med Records. 2026; 1(1): 1-7
Sorumlu Yazar: Hasan Eryeşil, Türkiye |
|