Medical Records

[Med Records]
Med Records. 2026; 1(1): 1-7 | DOI: 10.5505/medrec.2026.60352  

Comparative Performance of Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in Simplifying Turkish Abdominal Ultrasound Reports

Hasan Eryeşil1, Halil Ibrahim Altunbulak1, Yusuf Eryeşil2
1Clinic Of Radiology, Tatvan State Hospital, Bitlis, Türkiye
2Department Of Computer Engineering, Faculty Of Technology, Selcuk University, Konya, Türkiye

AIM: To evaluate and compare the performance of three advanced Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in simplifying real-world Turkish abdominal ultrasound reports to enhance patient comprehension.
MATRERIAL and METHOD: This retrospective, cross-sectional study included 100 adult abdominal ultrasound reports categorized by finding severity (normal, benign, and suspected malignancy/emergencies). Original reports were simplified to a 5th-grade reading level using standardized prompts for each model. Readability was measured quantitatively using the Ateşman Readability Formula. Two independent radiologists qualitatively assessed the outputs for medical accuracy, inclusiveness, readability, and hallucination rates
RESULTS: All AI models significantly improved readability from a baseline Ateşman score of 69.53 ± 5.51 to a range of 73.24–82.19 (p < 0.001). Claude 3.5 Sonnet achieved the highest quantitative readability score (82.19 ± 5.51). However, qualitative expert evaluation revealed ChatGPT and Gemini were superior to Claude in inclusiveness and conversational readability (p < 0.001). ChatGPT (4.82 ± 0.45) specifically excelled in integrating normal anatomical findings without losing clinical context. Hallucination rates were minimal across all models (1.3% - 2.0%, p=0.565).
CONCLUSION: LLMs demonstrate considerable potential in simplifying morphologically complex Turkish radiology reports while main taining clinical accuracy. While Claude 3.5 Sonnet shows strong performance in formula-based metrics, ChatGPT 5.0 and Gemini 3.0 tend to offer more favorable clinical fluency and inclusiveness. Despite high performance, rare hallucination risks such as “normality faking” necessitate a human-in-the-loop approach for patient safety

Keywords: Large Language Models, Natural Language Processing, Patient-Centered Care, Ultrasonography, Health Literacy


Hasan Eryeşil, Halil Ibrahim Altunbulak, Yusuf Eryeşil. Comparative Performance of Large Language Models (ChatGPT 5.0, Gemini 3.0, Claude 3.5 Sonnet) in Simplifying Turkish Abdominal Ultrasound Reports. Med Records. 2026; 1(1): 1-7

Sorumlu Yazar: Hasan Eryeşil, Türkiye


ARAÇLAR
Tam Metin PDF
Yazdır
Alıntıyı İndir
RIS
EndNote
BibTex
Medlars
Procite
Reference Manager
E-Postala
Paylaş
Yazara e-posta gönder

Benzer makaleler
PubMed
Google Scholar