Medical ChatGPT – A systematic Meta-Review
Institute for Artificial Intelligence in Medicine (IKIM), Essen University Hospital (AöR), Girardetstraße 2, 45131 Essen, Germany
Center for Virtual and Extended Reality in Medicine (ZvRM), Essen University Hospital (AöR), Hufelandstraße 55, 45147 Essen, Germany
Cancer Research Center Cologne Essen (CCCE), University Medicine Essen (AöR), Hufelandstraße 55, 45147 Essen, Germany
Institute of Computer Graphics and Vision, Graz University of Technology, Inffeldgasse 16, 8010 Graz, Austria
Department of Pathology, Microbiology, and Forensic Medicine, School of Medicine, University of Jordan, Amman 11942, Jordan
Department of Clinical Laboratories and Forensic Medicine, Jordan University Hospital, Amman 11942, Jordan
Department of Physics, TU Dortmund University, Otto-Hahn-Straße 4, 44227 Dortmund, Germany
Department of Oral and Maxillofacial Surgery, University Hospital RWTH Aachen, Pauwelsstraße 30, 52074 Aachen, Germany
Institute of Medical Informatics, University Hospital RWTH Aachen, Pauwelsstraße 30, 52074 Aachen, Germany
*Corresponding author: jan.egger@uk-essen.de (J.E.)Abstract
Since its release at the end of 2022, ChatGPT has seen a tremendous rise in attention, not only from the general public, but also from medical researchers and healthcare professionals. ChatGPT definitely changed the way we can communicate now with computers. We still remember the limitations of (voice) assistants, like Alexa or Siri, that were “overwhelmed” by a follow-up question after asking about the weather, not to mention even more complex questions, which they could not handle at all. ChatGPT and other Large Language Models (LLMs) turned that in the meantime upside down. They allow fluent and continuous conversations on a human-like level with very complex sentences and diffused in the meantime into all kinds of applications and areas. One area that was not spared from this development, is the medical domain. An indicator for this is the medical search engine PubMed, which comprises currently more than 36 million citations for biomedical literature from MEDLINE, life science journals, and online books. As of March 2024, the search term “ChatGPT” already returns over 2,700 results. In general, it takes some time, until reviews, and especially systematic reviews appear for a “new” topic or discovery. However, not for ChatGPT, and the additional search restriction to “systematic review” for article type under PubMed, returns still 31 contributions, as of March 19 2024. After filtering out non-systematic reviews from the returned results, 19 publications are included. In this meta-review, we want to take a closer look at these contributions on a higher level and explore the current evidence of ChatGPT in the medical domain, because systematic reviews belong to the highest form of knowledge in science.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This study did not receive any funding
Introduction
Comprehensive systematic reviews that synthesize the findings of multiple studies can provide a more robust and reliable understanding of ChatGPT’s performance, its limitations, and the challenges that still need to be addressed, which are particularly useful in fields where research is abundant and rapidly evolving, such as the adoption of ChatGPT in healthcare. They allow researchers to quickly grasp the key aspects of a topic without having to sift through numerous individual studies. The increasing popularity of ChatGPT among healthcare professionals and researchers alike leads to a surge in research focusing on the evaluation of ChatGPT’s performance in different application scenarios, such as medical consultation [1], research [2], education [3], or different medical specialties, such as neurology [4, 5], pediatric [6, 7], cosmetic surgery [8, 9] and dermatology [10, 11]. Each of these fields presents unique challenges and opportunities for ChatGPT. Understanding how ChatGPT performs in these different contexts is crucial for its continued development and improvement.
Due to the large amount of publications produced in a relatively short time, it is demanding for researchers to stay up-to-date with the latest development of ChatGPT in their specific domain and understand its challenges and limitations. Systematic reviews provide a quick and informative overview on the use of ChatGPT in a particular scenario or specialty [12, 13, 14, 15, 16, 17, 18, 19], which keep researchers updated with the field.
A meta-review on ChatGPT in healthcare synthesizes the findings of multiple systematic reviews, and therefore takes a broader view of the field, going beyond the scope of systematic reviews that focus only on specific scenarios or specialities. It aims to provide a concise and high-level summary of the current state of ChatGPT in the general healthcare sector [20], and help researchers, healthcare professionals, and policymakers understand the big picture in order to make sound decisions and policies. Future research directions can also be suggested.
We position our work as a meta-review of systematic reviews of ChatGPT in healthcare, which look at systematic reviews across various application scenarios and medical specialties. The aim of the meta-review is to provide a concise yet comprehensive summary of the status quo of ChatGPT in the healthcare sector, assess the overall performance of ChatGPT, identify common trends and patterns, highlight key challenges and limitations, and point out areas where further research is needed. However, it is important to notice that, at this stage, the number of comprehensive systematic reviews is still small, and the majority of related publications remain to be high-level commentaries or small-scale evaluations [21], which highlights a gap in the literature and an opportunity for future research.
Conclusion
Based on the systematic reviews about ChatGPT in the medical field in this contribution and the above discussions, we project that the number of systematic reviews will continue to increase, with more and more research on the evaluation of ChatGPT in a specific domain being published in the near future. However, it is important to note that, despite these reviews keeping track of the latest development of ChatGPT in healthcare, the quality and performance of the AI tool still depends on the underlying large language model, which needs to be continuously improved by researchers and practitioners in natural language processing (NLP). As healthcare professionals and researchers, our primary responsibility is to rigorously test the tool in our respective medical specialities in order to provide feedback on the actual capabilities of ChatGPT and expose its limitations and the ethical and legal concerns that arise. These in turn help NLP researchers to improve the language model and provide legislators the basis to formulate proper regulations. Summarized, we extracted the following three main challenges for ChatGPT in healthcare, but also medical NLPs in general, from our meta-review:
- (1.) Dependence on the Underlying Language Model: The quality and performance of ChatGPT are reliant on the underlying large language model, which is trained on vast amounts of text data and responsible for generating appropriate responses. Nevertheless, like any machine learning model, ChatGPT is prone to mistakes or failure given complex or ambiguous inputs. Therefore, continuous improvements by NLP researchers and practitioners are essential for fine-tuning the language model for a specific application. Furthermore, developing an artificial general intelligence (AGI) model with a wide spectrum of medical knowledge is a challenging but rewarding task.
- (2.) Role of Healthcare Professionals and Researchers: As users of ChatGPT in the medical field, healthcare professionals and researchers play an important role, who continuously provide feedback to NLP developers. Rigorous testing of the tool in their respective medical specialties is essential in providing valuable feedback on its actual capabilities in order to improve ChatGPT’s performance. The expert-in-the-loop process involves not only identifying applications that ChatGPT is good at, e.g., providing medical information or assisting with patient communication, but also helps to expose the limitations of ChatGPT, such as hallucination or inability to provide accurate information for complex medical questions. The feedback from healthcare professionals not only help NLP researchers to improve the language model but also provide legislators with a basis to formulate proper regulations.
- (3.) Ethical and Legal Concerns: Using AI tools like ChatGPT in healthcare can also inevitably raise ethical and legal issues, such as patient privacy and data security. A clear guideline on when and how ChatGPT should be used in a healthcare setting is required. Healthcare professionals, policymakers and regulators should work jointly to address these concerns.
To conclude, the development and use of AI tools like ChatGPT in healthcare is a collaborative effort that should involve NLP researchers, healthcare professionals and legislators. Each group has a crucial role to play in ensuring that these tools are (in this order) safe, beneficial and effective for patient care.
Data Availability
All data produced in the present work are contained in the manuscript
Acknowledgements
This work was supported by the REACT-EU project KITE (Plattform für KI-Translation Essen, https://kite.ikim.nrw/, EFRE-0801977) and FWF enFaced 2.0 (KLI 1044, https://enfaced2.ikim.nrw/). Behrus Puladi was funded by the Medical Faculty of the RWTH Aachen University in Germany as part of the Clinician Scientist Program. Furthermore, we acknowledge the Center for Virtual and Extended Reality in Medicine (ZvRM, https://zvrm.ume.de/) of the University Hospital in Essen, Germany.
Disclaimer
For some parts of the paper, Microsoft’s Bing Chat was used to generate hints using customized prompts.