Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications
1Life Supporting Technologies Research Group, Universidad Politécnica de Madrid, Avda Complutense 30, 28040, Madrid, Spain
*Corresponding author: mrujas@lst.tfo.upm.esAbstract
The development of Artificial Intelligence (AI) in the healthcare sector is generating a great impact. However, one of the primary challenges for the implementation of this technology is the access to high-quality data due to issues in data collection and regulatory constraints, for which synthetic data is an emerging alternative. This Scoping review analyses reviews from the past 10 years from three different databases (i.e., PubMed, Scopus, and Web of Science) to identify the healthcare domains where synthetic data are currently generated, the motivations behind their creation, their future uses, limitations, and types of data. A total of 13 main domains were identified, with Oncology, Neurology, and Cardiology being the most frequently mentioned. Five types of motivations and three principal future uses were also identified. Furthermore, it was found that the predominant type of data generated is unstructured, particularly images. Finally, several future work directions were suggested, including exploring new domains and less commonly used data types (e.g., video and text), and developing an evaluation benchmark and standard generative models for specific domains.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This study did not receive any funding
1.Introduction
The development and application of Artificial Intelligence (AI) in contemporary society is growing rapidly [1]. This technology has proven its capacity to revolutionize and influence the progress of various industries, including agriculture, transportation, and education [2]. Nevertheless, one of the most impactful areas for AI is the healthcare sector. In recent years, AI has demonstrated the potential to assist healthcare professionals in enhancing the diagnosis, treatment, and monitoring of different diseases [3]. The expected impact of AI extends beyond clinical improvements to significant economic benefits, with projected savings between 200 and 300 billion dollars only in the United States [4].
The implementation of AI in healthcare faces numerous constraints and barriers, including ethical, technological, regulatory, liability, personnel, patient safety and social issues [5]. A crucial factor in this context is the availability and quality of data, which can accelerate the implementation process by addressing some of these barriers and encouraging open scientific research. This emphasis on data aligns with a broader shift within the AI paradigm from a model-centric to a data-centric approach, where improving data access and quality is paramount to developing better AI systems [6].
In healthcare, access to high-quality data is particularly challenging due to the difficulties involved in data collection, such as the low prevalence of rare diseases, the critical conditions of some patients and the added burden for medical professionals [7]. In addition, privacy issues pose major obstacles due to the sensitive nature of health data and their potential misuse. Various techniques, such as federated learning and advanced encryption methods, are used to address these privacy concerns. However, an increasingly popular alternative is the generation of synthetic data, which can mitigate some privacy, access, intellectual property and regulatory issues while still providing valuable information [8].
Synthetic data can be defined as an artificial re-expression of real data through statistical processes, designed to mitigate privacy concerns and promote broad dissemination and open science [9]. The main goal of synthetic data is to provide a data resource that can be used for a variety of applications, such as testing and training machine learning models, while avoiding some of the risks and limitations associated with real data. Another advantage of synthetic data is its potential to improve fairness in artificial intelligence models. Synthetic datasets can be manipulated to better represent populations rather than simply reflecting the current state of the world. This means they can be designed to avoid racial or gender discrimination, helping to mitigate biases that might otherwise be present in real-world data.
Despite their advantages, synthetic data also present some challenges, particularly in terms of monitoring the results. Ensuring the accuracy and consistency of results from synthetic data can be complex, especially with complex datasets. The quality of synthetic data also depends heavily on the quality of the original data and the data generation model used. If the original data contain biases, these biases may be reflected in the synthetic data, potentially compromising its unbiasedness and usefulness. In addition, efforts to manipulate datasets to create fair synthetic data may inadvertently lead to inaccuracies, as overly sanitised data may not accurately reflect real-world conditions. These challenges indicate the importance of scrutiny and rigorous validation of synthetic data quality, also assuring the explainability for AI applications. In addition, all these considerations must be in line with compliance with emerging regulations such as the EU AI Act, first-ever legal framework on AI, which includes provisions for the use of synthetic data under Art.10 (Data and data governance), especially for training, validation and testing data sets in a high-risk sector such as healthcare.
As the field continues to evolve, there is a need to explore and understand the specific health domains in which synthetic data are generated and use cases addressing underrepresented data types like device, image, and genomic data. This comprehension may help to identify best practices, address potential bottlenecks and maximise the benefits of synthetic data in advancing health innovations.
2.Materials and Methods
The methodology employed for this Scoping review adheres to the guidelines established in the “Preferred Reporting Items for Systematic reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR)” [11]. The strategy followed included defining the research question, identifying and selecting relevant studies, and charting and reporting the findings.
2.1.Search strategy
The literature search was performed on April 25, 2024, across PubMed, Web of Science, and Scopus, utilizing key terms associated with synthetic data and healthcare (e.g., “synthe*”, “record” or “health*”). The queries were further refined by applying various filters available within each search engine. The complete search strategy is detailed in Table 1.
2.2.Inclusion and exclusion criteria
The inclusion and exclusion criteria were established following the JBI methods for a Scoping Review [12]. Articles were included in the review if they met the following conditions: (1) the study involved human subjects; (2) the publication was a review or systematic review because they already analyse the limitations and gaps in the primary literature so that by including only reviews, less researched areas and opportunities for future research can be easily identified; and (3) the study was published between 2014 and 2024. Articles were excluded based on the following criteria: (1) the study was not published in English; (2) the focus of the article was not related to health; (3) the study did not fall within the scope of the research questions; (4) the study involved animal subjects; (5) the article was later than 10 years; and (6) the publication type was a book, paper, clinical trial, meta-analysis, or randomized controlled trial.
2.3.Search and screening process
Following the article search and extraction process, duplicate articles were removed. Initially, two independent authors (M.R. and B.M.) reviewed the titles and abstracts to identify articles suitable for inclusion in the study. Any disagreements regarding the inclusion or exclusion of an article were resolved by a third reviewer (R.M.G.). Subsequently, the full-text articles were retrieved for a second review conducted by two reviewers (M.R. and R.M.G.). Data were then extracted from the articles that met the review criteria, including: authors, main healthcare domain and subdomain, motivations for creating synthetic data, future uses of the synthetic data, type of data generated, limitations identified by the authors related to the generation of synthetic data, and type of review. If there were any uncertainties during the data extraction process, the referenced article within the review study was thoroughly examined to extract the necessary information.
2.4.Data charting, appraisal and synthesis of results
The various extracted characteristics were documented using an Excel spreadsheet. For the fields of motivations, future uses, type of data, and limitations, the data were represented as enumerated lists, as multiple items from each category might be present within a single study. Additionally, within the data type category, a two-level classification was performed: initially determining whether the data was structured or unstructured, and subsequently specifying the exact type of data (e.g., images or text). After extraction, the data was analysed field by field using counters and groupings to better understand and report the different trends and patterns in this topic.
3.Results
The study results are organized as follows: first, the numerical outcomes of the search and screening process will be displayed using a PRISMA chart, providing clarity on the process and aiding future similar reviews. Next, a table with the information extracted from the articles will be presented. Finally, the specific results for each field of extracted information will be detailed.
3.1.Search and screening results
The PRISMA chart illustrating the search and screening process is shown in Figure 1.
Initially, 346 articles were retrieved from three search engines: 271 from Scopus, 39 from Web of Science, and 36 from PubMed. After removing duplicates, the titles and abstracts of 294 articles were reviewed. Of these, 142 articles passed the initial screening; however, the full text of 7 articles could not be retrieved. Following the full-text review, 42 articles were included for data extraction, while 93 were excluded for the following reasons: they did not address a specific healthcare domain (e.g., cardiology or oncology) but rather a broader topic (e.g., medical imaging or precision medicine) (n = 32); they discussed the potential for synthetic data generation rather than its current application (n = 26); they did not include anything related to synthetic data generation (n = 20); they were original research articles (n = 10); they were Method reviews that compared different synthetic data generation methods without providing real evidence of synthetic data generation in the domain beyond the study itself (n = 3); or they included non-human subjects (n = 2);
3.2.Data extraction results
The information extracted from the various articles is summarized in Table 2.
3.3.Specific results
The main findings for each specific topic are detailed in the subsequent subsections.
3.3.1.Main Healthcare Domain and Subdomain
The analysed articles encompass 13 main healthcare domains. These domains, listed from most to least frequently mentioned, are: Oncology (n = 10), Neurology (n = 7), Cardiology (n = 5), Endocrinology (n = 4), Epidemiology (n = 3), Gastroenterology (n = 3), Ophthalmology (n = 3), Immunology (n = 2), Dermatology (n = 1), Gynecology (n = 1), Nephrology (n = 1), Otorhinolaryngology (n = 1), and Pneumology (n = 1). Regarding subdomains, most reviews addressed different subdomains. The only subdomains appearing in more than one review were Diabetes (n = 4) and COVID-19 (n = 2).
3.3.2.Motivations
The reviewed articles identify 45 different motivations for generating synthetic data, which can be consolidated into five main categories: data privacy and security, data scarcity, data quality, AI development, and direct medical and clinical applications. Figure 2 illustrates the distribution of these motivations within each category and their relative frequencies. A detailed table summarising the frequency of these 45 motivations across the articles is provided in Annex 1 of the supplementary material.
3.3.3.Future Use
The generated data have been applied in 14 specific use cases, which can be broadly categorized into three main areas: AI development, including training and validation of models and enhancing the generalizability and interpretability of existing models; enabling secondary use of data, such as data sharing and conducting analyses and studies; and enhancing clinical knowledge, serving as educational material or support during diagnosis and therapy. Figure 3 depicts the distribution of these use cases within each category and their relative frequencies. A comprehensive table summarising the occurrences of these use cases in the articles is available in Annex 2 of the supplementary material.
3.3.4.Data Types
Regarding the types of data generated, 30 out of 42 articles (71.43%) refer to the generation of unstructured data, while 18 out of 42 articles (42.86%) refer to the generation of structured data. Among the unstructured data, images are the most commonly generated type (n = 28), followed by text (n = 2) and video (n = 1). For structured data, both time series (n = 11) and cross-sectional data (data collected at a single point in time, providing a snapshot) (n = 12) are similarly represented.
3.3.5.Limitations
Only 14 of the 42 reviews mentioned limitations. The most frequently cited limitations were related to GAN technical issues (n = 7), the absence of a standard evaluation benchmark (n = 5), and the transmission of biases from training data to synthetic data (n = 5). Additional limitations included the challenges posed by small and heterogeneous datasets, issues with the generalization of augmented data, and the loss of complexity inherent in real-world scenarios.
3.3.6.Review Type
Among the 42 reviews analysed, 22 do not specify their type. The remaining 20 reviews are categorized as follows: 10 Systematic reviews, 3 Systematic reviews and Meta-analyses, 2 Scoping reviews, 2 Narrative reviews, 2 Surveys, and 1 State of the Art review.
4.Discussion
In this section, we emphasize the contributions of this scoping review of reviews. This study has systematically analysed all the reviews from the last decade that discuss the generation of synthetic data in different domains of healthcare. Specifically, we examined the main domains and subdomains where synthetic data are being generated, the motivations behind this generation, the potential future applications of these data, the types of data generated, the limitations identified by the authors regarding synthetic data generation, and the types of reviews analysed. Understanding the current state and potential growth areas for synthetic data in healthcare is important to identify how these data can help in addressing key challenges and contribute to advancements in healthcare domain. General reflections from the full-text screening process suggest that while synthetic data generation is a promising field with significant potential to address many current challenges in a multidisciplinary manner, it still needs to be further integrated and applied across the different areas of healthcare. This highlights the need for focused research on the current generation and applications of synthetic data to fully realize its potential and overcome existing barriers to its widespread adoption.
Our findings towards healthcare domains indicate that synthetic data generation is most prevalent in fields such as Oncology, Neurology, and Cardiology, which reflects a high demand for data in these areas due to challenges like data scarcity and privacy concerns. Less frequently mentioned domains, including Dermatology, Gynecology, and Pneumology, suggest emerging interest and potential for further exploration. Regarding subdomains, there is greater variety, as most articles do not share common subdomains. The exceptions are the endocrinology articles focusing on diabetes [22,32,40,46] and articles addressing COVID-19 [28,54], reflecting the impact of these issues on society. This wide range of healthcare domains and subdomains where synthetic data is currently generated illustrates the versatility of this technology in the healthcare sector.
Arising from our analyses, we have found that the motivations for generating synthetic data, while diverse, raise several critical concerns that are worthy of further consideration. These motivations can be broadly classified into five categories: data privacy and security, data scarcity, data quality, AI development, and direct medical and clinical applications. The emphasis on data privacy and security, while valid, often oversimplifies the complexities involved in ensuring truly anonymised synthetic datasets. The assumption that synthetic data can completely mitigate privacy risks overlooks potential vulnerabilities in the data generation process, such as re-identification risks if synthetic data are not sufficiently differentiated from real data especially in the health field. The issue of data scarcity, especially in cases of rare diseases in patients, highlights an important gap in health research that allows synthetic data to provide an answer. However, the reliance on synthetic data to fill these gaps can lead to a false sense of data adequacy and we must consider principles mentioned above to ensure a minimum of quality. This limitation is often exacerbated by the quality of the available data, which is often unbalanced or incomplete.
The development of AI as a motivation for synthetic data generation is compelling, given the need for large datasets to train sophisticated models. However, the quality of the AI models produced is intrinsically linked to the quality of the synthetic data. The risk of perpetuating existing biases or introducing new ones is a major concern that needs to be rigorously addressed. Potential pitfalls, such as privacy risks, data quality issues and biases, make it clear that the recent AI Act regulation adopted by the European Commission, and which entered into force on 1 August 2024 will need to consider address these issues.
Furthermore, synthetic data are used to promote open science and secondary data use through data sharing for analyses and studies, as well as to improve clinical knowledge by assisting in various tasks such as diagnostics and personalized therapies or serving as educational material. This demonstrates that synthetic data are valuable not only for technological advancements in the healthcare sector, such as the development of decision support systems, but also in the scientific and academic fields, facilitating open research and the sharing of higher quality information.
The reviewed articles predominantly discuss the generation of unstructured data, particularly images, reflecting the critical role of medical imaging in healthcare. However, there is a notable gap in the generation of other types of unstructured data, such as video and text, which are increasingly relevant with the advent of more complex generative models. Structured data, including time series and cross-sectional data, also play a significant role in capturing comprehensive patient information and warrant further investigation to enhance their utility in healthcare applications.
Only one-third of the articles identified limitations related to synthetic data generation. The most frequently mentioned limitation was technical issues with GAN models, such as instability during training and mode collapse. Additionally, there is a strong emphasis on the need for standard evaluation benchmarks, as highlighted in other studies like the one conducted by Murtaza et al. [10], and the transmission of biases present in the original data to the synthetic data, which is a necessary consideration. These three limitations are not exclusive to the health domain but are relevant to any type of synthetic data. This implies that as research on synthetic data generation advances, its application across all the different domains will continue to evolve.
Lastly, it is worth noting that several studies have identified synthetic data generators for specific domains that are widely accepted and commonly used within the community. This practice should become more widespread in the coming years, fostering scientific progress across various health sectors and promoting research in synthetic data. The synthetic data generators referenced in several studies include the UVA/PADOVA Type 1 Diabetes Simulator [55], which is cited by 3 out of the 4 articles concerning diabetes; and the simulator for single-cell RNA sequencing data (SPLATTER) [56], which is cited by the article focused on cancer biology. These two synthetic data generators serve as an ideal reference for the scientific community, facilitating greater access to quality health data without compromising individual privacy.
4.1.Limitations and Strengths of the study
This scoping review demonstrates several strengths. Firstly, the exclusive analysis of domain-specific reviews ensures that the extracted content is reliable and relevant within the specified domain. Additionally, a systematic approach was employed, adhering to established guidelines and standards, enhancing the rigor of the review. To minimize bias, the extraction and analysis of results were conducted by three different researchers. These combined strengths ensure that the information provided in this review is aligned with the current knowledge in the literature.
However, a notable limitation of this study is the rapid growth of AI and its scientific output, which means that this review provides only a baseline snapshot that will continue to evolve in the coming years. Furthermore, the exclusion of articles discussing the future potential of synthetic data might have limited the scope of insights into emerging trends and anticipated developments in this field. Despite these limitations, this Scoping review complements the more technologically focused reviews by providing insights into the clinical and practical applications of synthetic data generation within healthcare, offering a robust foundation for future research in this field.
5.Conclusions and Future Work
Synthetic data generation is a promising technology with high potential to enhance healthcare and healthcare research. This study has reviewed 42 articles to provide a comprehensive overview of the primary healthcare domains and subdomains where these techniques are applied, their motivations, the types of data generated, their future uses, and the limitations encountered. While the analysis indicates that synthetic data with various characteristics and typologies are currently being generated across many main healthcare domains, the technology’s versatility and relatively early stage suggest considerable potential for future applications.
Firstly, further investigation and application are needed in domains where synthetic data is just beginning to be used, such as Immunology, Dermatology, and Gynecology. It is also essential to extend their application to new domains and subdomains that have not adopted these techniques yet. Additionally, the generation of videos and text, which remains underexplored in the healthcare field, has great possibilities, especially with the recent advancements in generative models and Large Language Models. Finally, efforts are needed to define a standard evaluation benchmark, as previously highlighted in the literature, and to develop reference models for specific domains such as the UVA/PADOVA Type 1 Diabetes Simulator to foster open research and investigation in the generation of synthetic health data.
Supporting information
Data Availability
All data produced in the present work are contained in the manuscript
Declaration of generative AI and AI-assisted technologies in the writing process
During the preparation of this work, the authors used ChatGPT 4 in order to proof of read some of the text. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the published article.