Social media discourse and internet search queries on cannabis as a medicine: A systematic review
1Department of General Practice, Melbourne Medical School, Faculty of Medicine, Dentistry & Health Sciences, The University of Melbourne, Victoria, Australia
2Centre for Digital Transformation of Health, Victorian Comprehensive Cancer Centre, The University of Melbourne, Victoria, Australia
3St Vincent’s Health - Department of Addiction Medicine, Melbourne, Victoria, Australia
*Corresponding authors Email: Sedigh.Khademi@unimelb.edu.au; hallinan@unimelb.edu.auAbstract
Objective
To systematically review the current published literature that uses social media data and internet search engine queries to identify emerging themes on the use of cannabis as a medicine. In this study, the term cannabis as a medicine refers to the use of cannabis with therapeutic intent and includes prescribed cannabis, over-the-counter cannabis, and non-prescribed cannabis used specifically for self-medication (not for recreational purposes).
Materials and Methods
For this systematic review, Medline, Scopus, Web of Science and Embase databases were searched for peer reviewed studies published in English between January 2010 and March 2022. All study types that examined social media data and cannabis as a medicine were included in the review.
Results
Forty studies were included in this review. 21 studies used manually labeled data, 4 studies used existing meta-data, 2 studies used data from social media analytics companies and 13 used computational methods for annotating data. More than half of the studies 22/40 (55%), were published in the last three years. The incidental use of cannabis as a medicine was found in eleven studies that were focused on more general health-related issues, rather than on medicinal cannabis only.
Conclusion
Our systematic review has revealed the growing interest in analyzing user-generated content for studying cannabis as a medicine. This review has also highlighted the need for studies into cannabis use for specific health conditions and for automatic processing of larger datasets using computational methods (machine learning technologies).
Article notes
Competing Interest Statement
Yvonne Bonomo is a Principal Investigator on an open label study to evaluate the safety, tolerability, and pharmacokinetics of a medicinal cannabinoid oil formulation in chronic non-cancer pain patients for Zelira Therapeutics ACTRN12619001013156.
Funding Statement
This work was undertaken as part of the MDRP program as part of Melbourne Medical School, Department of General Practice at the University of Melbourne, and the Australian Centre for Cannabinoid Clinical and Research Excellence (ACRE) is established through the National Health and Medical Research Council (NHMRC) Centre of Research Excellence scheme. It draws together over twenty Australian research leaders and clinicians from major national universities and research institutions to establish a research evidence base to inform safe clinical use of medicinal cannabinoids and to guide policy as cannabinoids are introduced into therapeutic practice in Australia.
Introduction
The plant cannabis sativa was first cultivated for use as a medicine in the Stone Age (1). From the 1800’s individuals accessed cannabis as a medicine by either prescription or as an over the counter therapeutic (2). Yet by the mid-20th century, cannabis use was prohibited in many parts of the developed world with the passing of legislation in the USA, the UK and various European countries that proscribed its use (3–6) More recently, new evidence regarding the clinical efficacy of cannabis for some conditions (7) has stimulated public interest in cannabis and cannabis-derived products (8, 9), resulting in a global trend towards public acceptance, and subsequent legalization of cannabis for both medicinal and recreational use.
There is emerging evidence of cannabis efficacy for childhood epilepsy, spasticity and neuropathic pain in multiple sclerosis, acquired immunodeficiency syndrome (AIDS) wasting syndrome, and cancer chemotherapy-induced nausea and vomiting (10–12). Although researchers are investigating cannabis for treating cancer, psychiatric disorders (13), sleep disorders (14), chronic pain (15) and inflammatory conditions such as rheumatoid arthritis (16), there is currently insufficient evidence to support its clinical use. Scientific studies on emerging therapeutics typically exclude vulnerable populations such as pregnant women, young people, the elderly, and those with comorbidities, and those who depend on multiple medications, thus limiting the availability of evidence for cannabis effectiveness in these population groups (17).
Cannabis as medicine is associated with a rapidly expanding industry (18). Patient demand is increasing, as reflected in an increasing number of approvals for prescriptions over time (19), with one study showing that 61% of Australian GPs surveyed reported one or more patient enquiries regarding medical cannabis (20). With this increasing demand, is sophisticated marketing by medicinal cannabis companies that leverages evidence from a small number of studies to promote their products (21, 22). In light of these development, concerns regarding patient safety are warranted especially when marketing for cannabinoid products is associated with inadequate labelling and/or inappropriate dosage recommendations in some cannabis products (23), the provision of over-the-counter cannabis products that do not require a prescription (24) and the illicit drug market (25). Given this dynamic interplay between marketing, product innovation, regulation, and consumer demand new innovative surveillance methods are required to augment existing established monitoring approaches.
Although there is considerable engagement from many stakeholders to improve the scientific evidence regarding the efficacy and safety of cannabis through randomized control trials, many gaps remain in the literature (26). However, data from real-world and patient reported data sources could provide opportunities to address this evidence deficit (27). This real-world data can be captured from a variety of sources such as found in routinely collected health care and health services records that include but is not limited to patient generated data from medical, administrative and claims data as well as patient reported data from surveys, wearable trackers, patient registries, and social media (28–30).
People readily consult the internet when looking for and sharing health information (31, 32). According to 2017 survey of Health Information National Trends, almost 78% of US adults used online searches first to inquire about health or medical information (31). Data resulting from these online activities is labelled ‘user generated’ and is increasingly becoming a component of surveillance systems in the health data domain (33). Monitoring user-generated data on the web can be a timely and inexpensive way to generated population-level insights (34). The collective experiences and opinions shared online are an easily accessible wide-ranging data source for tracking emerging trends – which might be unavailable or less noticeable by other surveillance systems.
The objective of this systematic review is to understand the utility of online user generated text in providing insight into the use of cannabis as a medicine. In this review, we aim to systematically review existing work that utilizes user-generated content to explore cannabis as a medicine. The objective of this systematic review is to synthesise quality primary research that uses social media discourse and internet search engine queries to answer the following questions:
- Does online user-generated text provide a useful data source for studying cannabis as a medicine?
- What are the research questions motivating the studies of online user-generated text that discuss medicinal use of cannabis?
- How can future research leverage user-generated content to study the use of cannabis as a medicine?
Materials and methods
For this systematic review we used a framework for systematic reviews and meta-analyses, the PRISMA guidelines to inform our methods (35). A manual search of Medline, Embase, Web of Science, and Scopus databases was independently conducted by SKH in May 2021 and again by CMH in March 2022. The search was limited to English-language studies that were published between January 1974 and March 2022. Given that much of the computational research in this area is published in peer-reviewed computer science venues that are not necessarily indexed by Medline we also used complementary literature databases (Embase, Web of Science, and Scopus) to generate our initial list of publications for screening.
Literature database queries were developed for four categories of studies. The first three categories used social media text as a data source, the fourth relied on internet search engine query data. For the first category, the database queries combined words used to describe social media forums, and cannabis-related keywords and general medical-related keywords (Table 1 Category 1). The second category also included the social media and cannabis-related keywords, but used keywords specific to psychiatric disorders, for which the use of medical cannabis has been described. Our search terms for this second category were informed by a systematic review of medicinal cannabis for psychiatric disorders (13) (Table 1 Category 2). The third category included social media and cannabis related keywords but focused on non-psychiatric medical conditions for which cannabis is sometimes used (Table 1 Category 3). The fourth category included studies using Internet search engine queries as a data source, there were no medical conditions included in these searches (Table 1 Category 4).
The inclusion criteria for this systematic review were: (i) primary research studies, (ii) studies which used online user-generated text as a data source, and (iii) research that was either directly focused on cannabis and cannabis products that have an impact on health, or were health-related studies that found medicinal use of cannabis.
Exclusion criteria comprised published: (i) editorials, letters, commentaries, and book chapters; (ii) studies that used social media for recruiting participants; (iii) studies where the full text of the publication was not available; (iv) studies primarily focused on electronic nicotine delivery systems adapted to deliver cannabinoids and (v) studies that used bots or autonomous systems as the main data source and (vi) studies that focused exclusively on synthetic cannabis.
Included studies were reviewed by CMH, YB, MC, and SKH. Where initial disagreement existed between reviewers regarding the inclusion eligibility of a study, team members met to discuss the disputed article’s status until consensus was achieved. For all the articles that met the inclusion criteria, pairs of reviewers independently critiqued the articles using a checklist developed for this study. The purpose of the checklist was to provide an overall assessment of quality rather than generate a specific score, a summary version of this checklist is presented (S1). Assessments of quality were based on evidence of weakness in aims or objectives, main findings, data collection method, analytic methods, data source, and evaluation and interpretations of the study.
Studies were excluded if they had: (i) poorly defined objectives; (ii) poorly described methodology; (iii) inappropriate analytic approach; (iv) small volume of data and/or (v) results that are not clearly reported).
Of the 1,271 titles identified in the electronic database searches, 781 were duplicates and 441 were excluded based on lack of relevance. This screening provided 49 potentially relevant articles for inclusion. Of these, five were excluded due to full text inaccessibility, and four were excluded based on reasons listed earlier. This provided 40 papers for inclusion in the review. The PRISMA flow diagram is presented in Fig 1. The PRISAM checklist is reported in S2 Checklist.
Results
Of the 40 papers synthesized for this review most were journal articles 38/40 (95%), and two were conference-based publications 2/40 (5%). Table 2 provides a summary of each paper that includes author names, publication year, data source, duration of the study, number of collected posts, number of analyzed posts, and the coding/labelling approach used.
Data collection and annotation
Table 2 shows that the largest dataset labelled by the researchers belongs to one of the earliest studies in this domain, where around 47,000 tweets are manually labelled (36). This paper was one of 21/40 (53%) of the papers which have either collected a limited number of data points or have sampled their collected data and manually coded the data to gain an in-depth understanding of the domain. 4/40 (10%) have used existing meta-data including geo-location data. 2/40 (5%) have used social media analytics companies’ data (37, 38). 13/40 (32%) have used an automated method for labelling data - these include machine learning, lexicon or rule-based algorithms. The earliest study that employed an automated labelling method was performed in 2017 (39), which used machine learning, thereafter the papers show a trend for using automation. The ability of such automated data-centric (40) approaches to handle more data means they can be utilized to broaden the available analyses, as more data enables the use of more sophisticated techniques such as machine learning models. Insights gained from the smaller qualitative studies can be used to inform the design of automated approaches, and even the manually labelled data can be used for training these models, so in this way automated methods can be built upon and leverage initial manual methods.
Data analysis
Most of the manually labelled studies focused on manual or statistical analysis (36, 41, 42, 46–50, 58, 63–65, 67, 70–74), while three analyzed online videos (44, 45, 54). Studies that utilized a large volume of data allowed for the use of computational methods, including sentiment analysis, topic modeling, and rule-based text mining. Sentiment analysis aims to analyse people’s sentiments, opinions, and attitudes (75, 77). Topic modelling is a machine learning method for automatically discovering common themes in a collection of documents (75, 78). Rule-based text mining involves classification of posts into pre-existing health-related categories (76). Studies also explored social network analysis, which examines how social media users connect through friendship, sending links and information, or tagging and following (79).
Research Themes
In this review, we categorized 40 research articles related to user-generated online text (social media discourse and internet search engine queries) into six broad themes. The themes are centered around the research questions motivating the studies.
Cannabis mode of use
Seven studies reported on the use of cannabis as a medicine in relation to its mode of use (Table 4). These studies collected data using keywords such as ‘vape’, ‘vaping’, ‘dabbing’ and ‘edibles’. Conversations around modes of use revealed a theme about lacking, seeking, or sharing knowledge about possible health consequences of the modes of use. Another theme was around the perceived health benefits of various modes of use including sleep improvement and relaxation resulting from dabbing oils (46) or consuming ‘edibles’ (47). The findings suggest that for emerging modes of use such as dabbing, where the availability of evidence-based information is limited, people seek information from others’ experiences.
Cannabis as a medicine for a specific health issue
Six studies were included in this theme (Table 5). These studies investigated conversations around the use of cannabis or cannabidiol for a specific health issue. The health conditions included glaucoma (63), PTSD (39), cancer (60, 70), Attention Deficit Hyperactivity Disorder (ADHD) (48) and pregnancy (71). These studies mostly discovered that conversations claimed benefits of cannabis as an alternative treatment for these health conditions, although mentions of harm, and both harm and therapeutic effects, were also present (48).
Cannabis as a medicine as part of discourse on illness and disease
Eleven studies were included in this theme. In this category, the research focus was on social media topics relating to management and treatment options for a range of health conditions rather than on medicinal cannabis per se (Table 6). There were eleven studies, and the health conditions included inflammatory and irritable bowel disease(59), opioid use disorder (37, 55), pain (38), ophthalmic disease (41), cluster headache and migraine(49), asthma (44), cancer (52, 67), autism disorder (66) and brachial plexus injury (73).
Cannabidiol (CBD)
There were six studies in the cannabidiol (CBD) category (57, 64, 65, 68, 74, 75). These studies concentrated on conversations related to the benefits of CBD products, product sentiment (positive, negative, or neutral), the factors that impact on a person’s decision to use CBD products, and the trends in therapeutic use of CBD.
Adverse drug reactions and adverse effects
One study focused explicitly on adverse reaction detection (Table 8) This study explored the prevalence of internet search engine queries relating to the topic of adverse reactions and cannabis use. Mentions of adverse effects of cannabis consumption were found in eight studies, these included respiratory effects and anxiety (56, 58, 76), addiction and withdrawal syndromes (50), neurological harm (76), harm in pregnancy (71, 76) and general harm (45, 46, 48, 76).
Discussion
Currently, there exist systematic reviews of cannabis and cannabinoids for medical use based on clinical efficacy outcomes from randomized clinical trials (15) and reviews on the use of social media for illicit drug surveillance (81). However, to our knowledge, this paper constitutes the first systematic review examining studies that used user-generated online text to understand the use of cannabis as a medicine in the global community.
Our systematic review found that the use of social media and internet search queries to investigate cannabis as a medicine is a rapidly emerging area of research. Over half of the studies included in this review were published within the last three years, this reflects not only increase community interest in the therapeutic potential of cannabinoids, but also world-wide trends towards cannabis legalization (4, 82–86). Regarding social media platforms, Twitter was the data source in seventeen (42.5%) of the 40 studies, almost three and a half times the number of studies using Reddit (5, 12.5%) and just under three times the number of studies using data from
Online forums (6, 15%). Three (7.5%) GoFundMe studies and three (7.5%) Google Trends studies were also included in the review. Hence, much of the data in this systematic review comprised posts from the Twitter platform. Several factors may explain this finding, firstly Twitter is real-time in nature, it has a high volume of messages, and it is publicly accessible. These factors makes it a useful data source for public health surveillance (87).
Regarding the subjects of the studies, eleven (29%) focused on general user-generated content regarding the treatment of health conditions (glaucoma, autism, asthma, cancer, bowel disease, brachial plexus injury, cluster headaches, opioid disorder). These studies were not explicitly designed to investigate cannabis as a medicine, yet they generated results that incidentally found cannabis mentioned as an alternative or complimentary treatment, either formally prescribed or via self-medication.
Qualitative studies feature prominently in the research, but while their contribution is valuable, especially in the context of hypothesis generation, they tend to be limited by their smaller datasets, which frequently comprised manually annotated samples. The recent emergence of powerful machine learning-based natural language processing (NLP) models suggests that it should be possible to automate the continuous processing of far larger datasets using NLP technologies, built upon the insights gained from initial qualitative studies, and even leveraging their annotated data for training purposes. Recent trends in the social science data landscape have shown a convergence between social science and computer science expertise, where the ability to use computational methods has greatly assisted the collection and validation of robust datasets that can form the basis of deeper social science research (88).
We found much heterogeneity in approaches applied to analyse user-related content, and inconsistent quality in the methodologies adopted. While we endeavored to include as many studies as possible, some of the publications initially identified as suitable for inclusion were not suitable based on a minimum quality requirements checklist (S1). This checklist was designed to ensure that selection of data source, choice of platform, data acquisition and preparation, analysis and evaluation delivers data and conclusions that are appropriate for answering the research questions.
The utilization of user-generated content for health research is subject to several inherent limitations which include the lack of control that researchers have in relation to the credibility of information, the frequently unknown demographic characteristics and geographical location of individuals generating content, and the fact that social media users are not necessarily representative of the wider community (89). Furthermore, the uniqueness, volume, and salience of social media data has implications that need to be considered when used for health information analysis (90). Volume is usually inversely related to salience: a platform such as Twitter has a very high volume of information, but, apart from its frequency, much of which is not highly pertinent for analyzing an effect; whereas the volume of information contained in a blog will much less but likely more salient for analysis. Notwithstanding these limitations, user-generated content comprises large-scale data that provides access to the unprompted organic opinions and attitudes of cannabis users in their own words, and is an effective medium through which to gauge public sentiment. To date, insights regarding cannabis as medicine have gained primarily through surveys or focus groups which have their own limitations regarding the format of data collection and potential bias in participant recruitment.
Conclusion
This systematic review has shown that user-generated content as a data source for studying cannabis as a medicine is a growing area worthy of investigation given that it provides another means to understand how cannabis is being used and perceived by the community. As such, it is another potential ‘tool’ with which to engage in pharmacovigilance of, not only cannabis as a medicine, but also other novel therapeutics as they enter the market.
Data Availability
N/A
Acknowledgments
This systematic review was supported by the Australian Centre for Cannabinoid Clinical and Research Excellence (ACRE), funded by the National Health and Medical Research Council (NHMRC) through the Centre of Research Excellence scheme (NHMRC CRE APP1135054).
Supporting information
S1. Quality assessment checklist (Doc)
S2. PRISMA checklist (Doc)