Skip to content

Annals of Clinical and Analytical Medicine

E-ISSN: 2667-663X · Monthly · English

Use of large language models in radiological reports: a study on simplifying turkish MRI findings

Use of large language models on simplifying turkish MRI reports

Abstract

AimAdvanced Large Language Models (LLMs), like ChatGPT, are known for their human-like expression and reasoning abilities. They are used in many fields, including radiology. This study is pioneering in evaluating and comparing the effectiveness of LLMs in simplifying Magnetic Resonance Imaging (MRI) findings in Turkish.MethodsIn our study, we simplified 50 fictional MRI findings in Turkish language using different LLMs, including ChatGPT-4, Gemini Pro 1.5, Claude 3 Opus and Perplexity. We compared the responses based on Ateşman’s readability index and word count. Additionally, three radiologists assessed the medical accuracy, consistency of suggestions, and comprehensibility of the answers, scoring each model on a scale of 1 to 5.ResultsThere was no statistically significant difference between the scores of Gemini 1.5 Pro (average: 4.9; median: 5.0), Opus (average: 4.8; median: 5.0), and ChatGPT-4 (average: 4.8; median: 5.0) (P > .05). However, there was a significant difference between the scores of Gemini 1.5 Pro and Perplexity (average: 3.7; median: 4.0) (P < .001). According to the readability index, Gemini 1.5 Pro had the highest average score of 59.3, which was significantly higher than the other LLMs (P < .005). In terms of word count, ChatGPT-4 used the most words (151.5), while Perplexity used the fewest (88.4).ConclusionThis study is the first to evaluate the ability of LLMs to simplify MRI findings in Turkish. The results suggest that radiologists find these models effective in making radiology reports more understandable. However, additional research is necessary to confirm these findings.

Keywords

large language modelradiology reportsreadabilityhealth communication

Introduction

Natural language processing (NLP) tools, especially large language models (LLMs), have gained the ability to generate highly accurate and human-like text through extensive training on large datasets.1 ChatGPT, a state-of-the-art NLP model released by OpenAI in November 2022, has gained worldwide attention for its human-like expression and reasoning abilities. It has been applied in various fields, including writing, summarization, medical knowledge, and medical education.2-3
In the medical field, LLMs have been the focus of many studies and have sparked interest in radiology.4-5 Their success in radiological assessments, understanding of radiological guidelines, and contributions to differential diagnosis and decision-making have recently generated significant excitement among radiologists.6-7
Radiology reports summarize radiologists’ findings and opinions based on imaging studies. They are critical in daily practice. However, medical terminology and key insights in these reports can be hard to understand for patients and physicians from other specialities. LLMs now enable the adaptation, simplification, and translation of professional radiology reports into different languages. This makes the reports comprehensible to individuals without medical knowledge and highlights key points.8-10
Successful simplification of radiology reports by LLMs can greatly enhance health communication and clarity for patients and their relatives. To the best of our knowledge, there are no studies on this subject for MRI findings in Turkish yet. Given the lack of studies on this topic in Turkish language, we aimed to evaluate and compare the effectiveness of LLMs, particularly in simplifying Magnetic Resonance Imaging (MRI) findings in Turkish.

Materials and Methods

In our study, we evaluated and compared the abilities of large language models (LLMs) such as ChatGPT-4, Gemini 1.5 Pro, Claude 3 Opus, and Perplexity to simplify Turkish MRI findings. Our study only included fictional MRI findings and did not use actual radiology reports, so it did not require ethical board approval. The study design followed the Standards for Reporting of Diagnostic Accuracy Studies (STARD) and the principles outlined in the Declaration of Helsinki.11
For our study, the authors jointly created 50 fictional Turkish MRI findings used in radiology reports. Care was taken to ensure these findings were common in daily practice and portrayed realistically. The findings included 20 related to neuroradiology, 15 to musculoskeletal radiology, and 15 to abdominal radiology. Table 1 showcases 20 of these findings as examples.
We utilized LLMs named ChatGPT-4, Gemini Pro 1.5, Perplexity, and Claude 3 Opus. We chose these models because of their timeliness, powerful capabilities, and the fact that they come from different companies.4-5,12 The designed findings were entered into each LLM via their respective websites following the prompt, “I will write the findings from the MRI report below. Please explain them in a way that someone without a medical background can understand.” in Turkish. Each finding was processed in a new window with default settings used for each model. The study was conducted between April 25 and April 28, 2024.
We analyzed the responses from the LLMs using the Ateşman Readability Index, a well-established Turkish readability measure, to determine readability levels.13 This index includes the number of syllables of words and the number of words of sentences in its formula.[198,825 - (40,175 x number of syllables/number of words) - (2,610 x number of words/number of sentences)] We used the publicly available and free website “www.okunabilirlikindeksi.com” for this process.
Three authors jointly rated the responses of the LLMs using a Likert scale from 1 to 5, based on medical accuracy, consistency of recommendations, and comprehensibility.
We also measured and compared the word count of the LLMs’ responses to assess their word efficiency, which is believed to directly impact readers’ reading time. Figure 1 summarizes the workflow of the study.
We used IBM SPSS Version 26 for statistical analyses. We checked data distribution with the Kolmogorov-Smirnov and Shapiro-Wilk tests. The Levene test assessed data variance. Descriptive statistics included minimum, maximum, average, median, standard deviation, interquartile range, and percentages. To find significant relationships between quantitative data in dependent groups, we used the Friedman and Wilcoxon tests. We used Spearman correlation analysis to examine the linearity of correlations between quantitative data.

Results

There was no significant difference between the scores of Gemini 1.5 Pro (average: 4.9; median: 5.0), Opus (mean: 4.8; median: 5.0), and ChatGPT 4 (mean: 4.8; median: 5.0).[P > .05] However, Gemini 1.5 Pro had significantly superior scores compared to Perplexity (mean: 3.7; median: 4.0).[P < .001] Table 3 summarizes the characteristics of each model in our study.
Gemini Pro 1.5 received a score of 5 in 48 out of 50 questions and a score of 4 in 2 questions. Opus scored 5 in 44 questions, 4 in 2 questions, and 3 in 4 questions. ChatGPT 4 scored 5 in 43 questions and 4 in the remaining 7 questions. The lowest performing LLM in our study was Perplexity, which scored 5 in 9 questions, 4 in 25 questions, 3 in 11 questions, 2 in 4 questions, and 1 in 1 question. Figure 2 shows the box-plot graphs of the scores obtained by the language models.
According to Ateşman’s readability index, Gemini 1.5 Pro had the highest average score of 59.3, which was significantly higher than the other LLMs.[P < .005] Table 3 There was no significant difference between the average scores of ChatGPT 4 (53.85) and Opus (52.86).[P = .543] Perplexity’s average readability index (47.01) was significantly lower than all other LLMs.[P < .001]
ChatGPT 4 had the highest average number of words used (151.5), which was significantly higher than all other language models.[P < .001] Table 3 Gemini 1.5 Pro, with the second-highest average (121.1), also used significantly more words compared to the other models.[P < .001] Opus (103.1) used significantly more words than Perplexity (88.4).[P = .008]
There was a linear correlation between the number of words in the designed findings and the number of words produced by Gemini 1.5 Pro (correlation coefficient = 0.706).[P < .001] and ChatGPT 4 (correlation coefficient = 0.585).[P < .001] However, no linear correlation was found for Opus.[P = .224] and Perplexity.[P = .420]
A linear correlation existed between the readability index of the designed findings and the readability indices of the responses from Opus (correlation coefficient = 0.545).[P < .001] Perplexity (correlation coefficient = 0.387).[P = .005] and Gemini 1.5 Pro (correlation coefficient = 0.294).[P = .038] However, no correlation was found between the readability index of the findings and the readability index of ChatGPT 4.[P = .402]

Discussion

Our study showed that even fictional, commonly used Turkish findings in MRI reports can be simplified by large language models (LLMs) for people without medical backgrounds. Gemini 1.5 Pro, Claude 3 Opus, and ChatGPT 4 received near-perfect scores from three radiologists for accuracy, consistency, and comprehensibility. Perplexity also scored above average (mean: 3.7/5; median: 4/5). Gemini 1.5 Pro had the highest Ateşman readability index score (59.3), significantly higher than all other models. ChatGPT 4 (53.85) and Opus (52.86) had mid-range scores. Perplexity had the lowest score (47.01), significantly lower than the other models. We used the Ateşman readability index to measure the readability of the simplified MRI reports generated by the LLMs in Turkish. Ateşman developed this formula in 1997.13 It evaluates the readability of Turkish texts based on the average number of syllables per word and words per sentence. Scores range from 1 to 100, with higher scores indicating easier reading.13 Ateşman highlighted that a text’s success depends on both readability and comprehensibility. Readability is evaluated with quantitative data, while comprehensibility is assessed qualitatively using the content of the text.13 We evaluated the readability of the responses with the Ateşman index and their comprehensibility using a Likert scale. We acknowledge that ratings by individuals without medical backgrounds would provide more valuable insights into comprehensibility. ChatGPT 4 used significantly more words (151.5) than the other models, indicating that its responses would take more time to read. Gemini 1.5 Pro followed with 121.1 words, then Opus with 103.1 words, and finally Perplexity with 88.4 words. We evaluated the word counts of the responses, considering the relationship between the number of words and the time required for users to read them. In a similar study, Jeblick et al. used the prompt, “Explain this medical report to a child using simple language,” with ChatGPT 3.5 for three fictional radiology reports.8 Fifteen radiologists rated the responses based on accuracy, comprehensiveness, and potential harm to the patient. Almost all responses were rated as accurate, comprehensive, and unlikely to cause harm.8 Our study was conducted entirely in Turkish and examined various LLMs in comparison to ChatGPT 3.5. This approach allowed us to evaluate the performance of other LLMs as well. In a similar study, Schmidt et al. used ChatGPT 3.5 to simplify knee MRI findings of varying complexity (simple, moderate, and complex) with five different prompts.9 Four doctors (two orthopedists and two radiologists) and 20 patients evaluated the simplified reports. The doctors rated the reports as “neutral” for informativeness but agreed they were “good” in accuracy and comprehensibility, posing no harm to patients. Patients felt better after understanding the simplified reports.9 Lyu et al. examined 62 thorax CT and 76 brain MRI reports.10 Each report had three simplified versions based on different prompts: making the report easier to understand, providing patient advice, and offering healthcare professional recommendations.10 Two radiologists rated these simplified reports on the overall score, comprehensiveness, and accuracy. ChatGPT 3.5, using more comprehensible language, scored 4.27/5. ChatGPT 4 produced even better-quality reports. They also explored how different prompts could create varied reports for patients with different education levels and found no significant differences.10 We did not use such specific prompts. Instead, we compared the baseline responses of the models using the prompt “in a way that someone without a medical background can understand.” Studies on how prompt engineering can improve readability levels in Turkish reports are needed to see how they would affect the readability levels of the responses. Li et al. studied 100 X-ray, ultrasound, CT, and MRI reports.14 They examined their lengths, Flesch reading ease scores, and Flesch-Kincaid reading levels.14 They used the prompt, “Explain this radiology report to a patient in layman’s terms: .” The simplified reports were statistically shorter, easier to read, and at a lower reading level than the originals.14 We did not use specific commands like making the report shorter, longer, or easier to read. Studies exploring the impact of such commands in Turkish reports through prompt engineering would be beneficial for demonstrating how much readability and word count can be improved.

Limitations

Our study is the first to examine the simplification of Turkish MRI findings by LLMs for individuals without medical backgrounds. However, there are limitations. The main limitation is the lack of patient inclusion, so we do not have their opinions on these simplified reports. Future studies should compare patient understanding of standard MRI reports with those simplified by LLMs. This would provide valuable feedback on the comprehensibility and usefulness of the simplified reports from the end-users perspective. Another limitation is that we used only fictional findings related to a single condition, not actual radiology reports. More complex reports covering all relevant findings might yield different results. Lastly, we used only one prompt. Different prompts might produce better or worse results depending on the models’ capabilities.

Conclusion

Our study suggests that large language models might effectively simplify Turkish MRI findings, potentially enabling patients to read and understand their MRI reports. This understanding could lead to better patient comprehension of their diagnoses and treatments, possibly resulting in enhanced compliance. Nevertheless, further research is required to address the limitations identified in our study and to validate these preliminary findings.

Declarations

Animal and Human Rights Statement

All procedures performed in this study were in accordance with the ethical standards of the institutional and/or national research committee and with the 1964 Helsinki Declaration and its later amendments or comparable ethical standards.

Data Availability

The datasets used and/or analyzed during the current study are not publicly available due to patient privacy reasons but are available from the corresponding author on reasonable request.

Conflict of Interest

The authors declare that there is no conflict of interest.

Funding

None.

References

  1. Thirunavukarasu AJ, Ting DSJ, Elangovan K, et al. Large language models in medicine. Nat Med. 2023;29(8):1930-1940. doi:10.1038/s41591-023-02448-8
  2. Grewal H, Dhillon G, Monga V, et al. Radiology gets chatty: the ChatGPT saga unfolds. Cureus. 2023;15(6). doi:10.7759/cureus.40135
  3. Kung TH, Cheatham M, Medenilla A, et al. Performance of ChatGPT on USMLE: potential for AI-assisted medical education using large language models. PLOS Digit Health. 2023;2(2). doi:10.1371/journal.pdig.0000198
  4. Li D, Gupta K, Bhaduri M, Sathiadoss P, Bhatnagar S, Chong J. Comparing GPT-3.5 and GPT-4 accuracy and drift in radiology Diagnosis Please cases. Radiology. 2024;310(1). doi:10.1148/radiol.232411
  5. Ueda D, Mitsuyama Y, Takita H, et al. ChatGPT’s diagnostic performance from patient history and imaging findings on the Diagnosis Please quizzes. Radiology. 2023;308(1). doi:10.1148/radiol.231040
  6. Yilmaz EC, Belue MJ, Turkbey B, Reinhold C, Choyke PL. A brief review of artificial intelligence in genitourinary oncological imaging. Can Assoc Radiol J. 2023;74(3):534-547. doi:10.1177/08465371221135782
  7. Akinci D’Antonoli T, Stanzione A, Bluethgen C, et al. Large language models in radiology: fundamentals, applications, ethical considerations, risks, and future directions. Diagn Interv Radiol. 2024;30(2):80-90. doi:10.4274/dir.2023.232417
  8. Jeblick K, Schachtner B, Dexl J, et al. ChatGPT makes medicine easy to swallow: an exploratory case study on simplified radiology reports. Eur Radiol. 2024;34(5):2817-2825.
  9. Schmidt S, Zimmerer A, Cucos T, Feucht M, Navas L. Simplifying radiologic reports with natural language processing: a novel approach using ChatGPT in enhancing patient understanding of MRI results. Arch Orthop Trauma Surg. 2024;144(2):611-618.
  10. Lyu Q, Tan J, Zapadka ME, et al. Translating radiology reports into plain language using ChatGPT and GPT-4 with prompt learning: results, limitations, and potential. Vis Comput Ind Biomed Art. 2023;6(1):9. doi:10.1186/s42492-023-00136-5
  11. Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. Radiology. 2015;277(3):826-832. doi:10.1148/radiol.2015151516
  12. Horiuchi D, Tatekawa H, Shimono T, et al. Accuracy of ChatGPT-generated diagnosis from patient’s medical history and imaging findings in neuroradiology cases. Neuroradiology. 2024;66(1):73-79. doi:10.1007/s00234-023-03252-4
  13. Ateşman E. Türkçede okunabilirliğin ölçülmesi [Measuring readability in Turkish]. Dil Derg. 1997;58(2):71-74.
  14. Li H, Moon JT, Iyer D, et al. Decoding radiology reports: potential application of OpenAI ChatGPT to enhance patient understanding of diagnostic reports. Clin Imaging. 2023;101:137-141. doi:10.1016/j.clinimag.2023.06.008

Tables

Table 1. The English translations of 20 of the fictional Turkish MRI findings designed for the study are shown

Table 2. The Ateşman readability Index and its corresponding readability level are shown

Table 3. Descriptive findings of the responses of the large language models are shown

*Likert scores: In our study, authors rated the accuracy of the explanations, consistency, and comprehensibility of the suggestions made by the large language models on a scale of 1 to 5. SD: Standard deviation, IQR: Interquartile range

Additional Information

Publisher’s Note
Bayrakol MP remains neutral with regard to jurisdictional and institutional claims.

Rights and Permissions

Creative Commons License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0). To view a copy of the license, visit https://creativecommons.org/licenses/by-nc/4.0/

About This Article

How to Cite This Article

Turay Cesur, Eren Çamur, Yasin Celal Güneş. Use of large language models in radiological reports: a study on simplifying turkish MRI findings. Ann Clin Anal Med 2024;15(8):586-590. doi:10.4328/ACAM.22266

Publication History

Received:
17.05.2024
Accepted:
02.07.2024
Published Online:
10.07.2024
Printed:
01.08.2024