Cannabis Use Documentation within the Electronic Health Record: A Use Case for Natural Language Processing Methods
aCenter for Pharmacy Innovation and Outcomes, Geisinger College of Health Sciences, Danville, PA, US
bDepartment of Pharmacy, Geisinger College of Health Sciences, Danville, PA, US
cDepartment of Health Promotion & Policy, University of Massachusetts, Amherst, MA, US
dDepartment of Population Health Sciences, Geisinger College of Health Sciences, Danville, PA, US
eBiomedical Research Informatics Center, Nemours Children’s Health, Wilmington, DE, US
fDepartment of Developmental Medicine, Geisinger College of Health Sciences, Danville, PA, US
gCenter for Substance Use Research and Education, Geisinger College of Health Sciences, Danville, PA, US
hSection Chief of Palliative Medicine, Department of Supportive Oncology, Levine Cancer Institute, Charlotte, NC, US
iDepartment of Emergency and Hospital Medicine, Jefferson Health-Lehigh Valley, Allentown, PA, US
jDepartment of Medical Education, Geisinger College of Health Sciences, Scranton, PA, US
kDepartment of Investigational Medicine, Geisinger College of Health Sciences, Danville, PA, US
#Corresponding Author: Apoorva M. Pradhan, E-mail address: apradhan1@geisinger.eduAbstract
Introduction
Recreational and medical cannabis use (CU) information is often available within the electronic health record (EHR) in a format that is impractical for health care provider use. Transformation of free-text EHR documentation in notes to discrete elements is possible using natural language processing (NLP) and has the potential to characterize CU efficiently. The objective of this study was to develop an NLP algorithm to identify documentation of CU within EHR unstructured clinical notes.
Methods
We identified EHR notes with cannabis-related terminologies through a keyword search among all Geisinger patients with at least one encounter between 1/1/2013 and 6/30/2022. We trained four NLP models to classify notes into six categories based on time, context, and reliability of CU documentation identified through manual annotation. We compared the demographic characteristics of patients with positive classification for CU using the best-performing model to those of the overall population.
Results
Of the over 1.7 million eligible patients, 150,726 (8.6%) were flagged as cannabis users. The Bio-ClinicalBERT, a transformer-based NLP model, achieved close to human performance in classifying CU (weighted Precision=91.4, Recall=93.3, F-score=92.4). Cannabis users had higher BMI and were at least nine-fold more likely to use tobacco, alcohol, and illicit substances.
Conclusion
Our study evaluated the prevalence of CU documentation across the entire corpus of EHR notes data without population segmentation. The NLP methodologies used achieved performance close to that of human annotation and laid the foundation for identifying and classifying CU within unstructured data sources, with future applications in research and patient care.
Plain Language Summary
Marijuana, also known as cannabis, may impact the health of patients, yet it is not routinely captured in medical records, and when documented, it is often found in unstructured formats (e.g., progress notes) rather than in discrete fields. Incomplete and unstructured capture limits many functional capabilities within the EHR that enhance patient care (e.g., drug interactions, notifications) and limit researchers from identifying patients routinely exposed to marijuana use. The transformation of free-text documentation of cannabis use (CU) into discrete elements can be performed using natural language processing (NLP). The objective of this study was to develop an NLP model to identify CU in unstructured clinical notes in the EHR. We examined the EHRs of Geisinger patients in Pennsylvania over a 10-year period. Among 1.7 million patients, 9% were identified as CU. One of the NLP models tested, Bio-ClinicalBERT, achieved the highest performance. Cannabis users had a higher BMI and were ten-fold more likely to be tobacco users, ten-fold more likely to use alcohol, and nine-fold more likely to use illicit substances. NLP can be used to better understand the risks and benefits of CU at a population level and may improve patient identification to assist clinical decision-making. Future CU epidemiological research should continue to explore other avenues to automate and improve CU documentation by leveraging rapidly evolving technologies, such as artificial intelligence-driven tools.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
The study is funded by Ascend Wellness Holdings, Clinical Registrant through the Geisinger Academic Clinical Research Center. The funder was not involved in the study or the decision to submit for publication.
Introduction
Cannabis products, both recreational and medical, are commonly used in the United States (US)[1–3]. The 2024 National Household Survey on Drug Use and Health estimated that 44.3 million people, age > 12, had past-month marijuana use[4]. The Commonwealth of Pennsylvania (PA) passed its Medical Marijuana (MMJ) Act in 2016[5]. As of November 2025, more than one out of every twenty-five PA (PA) adults (3.37%) were certified for medical marijuana[6] with anxiety disorders, chronic pain, and post-traumatic stress disorder accounting for the preponderance (85.6%) of certifying conditions (Supplemental Figure 1)[7,8].
The medical marijuana certification process in PA, like most other states, is typically completed by a physician with whom the patient does not have a long-standing relationship, as they have with their primary care provider, often by telemedicine [9,10]. Three-quarters of states, including PA, that have a medical marijuana program do not report any information to their Prescription Drug Monitoring Program [11]. With the executive order from December 2025 facilitating the potential reclassification of cannabis from a Schedule I to a Schedule III substance, there are a multitude of reasons why physicians and pharmacists need readily available CU information. These include recognition of cannabis-drug interactions [12,13], alignment with evidence driven utilization for medical conditions for which CU is indicated [14–16], awareness of CU as a substitute for prescription medications [17], aiding diagnosis of CU-related conditions (e.g. cannabis use disorder, hyperemesis syndrome, etc.) [18], and awareness of CU can assist clinicians in discussing safe use practices (e.g. storage in homes with children) [19]. The examples listed above highlight some of the scenarios that explain the necessity behind improving the understanding of an individual’s cannabis use and its importance in making informed health decisions.
Despite the need, the amount of actionable information documented in electronic health records (EHR) remains inadequate (e.g., only 4.8% of medical CU is documented in the EHR compared to 35.1% that is implicitly reported)[20,21]. The most common source of CU documentation is unstructured clinical notes [22–31] or patient portal messages [32]. A prior report from members of this team in PA discovered that secure patient portal messages between patients and health care providers discussing cannabis from 2012-2022 peaked in 2019 after dispensaries opened in 2018 [32]. Natural language processing (NLP) models have been leveraged to identify CU-related documentation from unstructured text [26–31]. Most prior studies applying NLP for CU identification have focused on specific sub-populations by disease states, such as psychosis patients, primary care patients, older adults undergoing surgery [26–28,30,31,33] or limited age ranges like pediatric and young adults, [26,29,31], thereby limiting their applicability for understanding CU across the entire population captured within the EHR. The purpose of this investigation was to extend prior studies [25,32] and to use validated NLP methodologies to classify CU from unstructured EHR notes, providing a proof of concept that free-text EHR notes can be extracted for CU using an automated mechanism from a broad, unstratified population [26].
Methods
We adhered to the STROBE reporting guidelines for reporting purposes (Supplement F, [34]). Figure 1 presents the pipeline and phases involved in this study.
Study Setting
This investigation was conducted at Geisinger, an integrated health system in the US, spanning 24 counties in central and northeast PA [Supplemental Figure 2A]. Geisinger comprises ten acute care hospitals, 133 primary and specialty clinics, a research institute, and an insurance company, Geisinger Health Plan. Geisinger began implementing its EHR, Epic® Corporation (Verona, WI), in 1996 at all ambulatory sites and integrated it across inpatient care sites in 2003. This infrastructure provides a comprehensive view, including clinical information, imaging, and administrative claims for rural and urban residents, offering a real-world examination of dynamic changes in this population.
Patient Population
The study included patients with at least one encounter (e.g., inpatient, outpatient, emergency department) at a Geisinger facility between January 1, 2013, and June 30, 2022. This population will henceforth be referred to as the overall population. For patients included in the overall population, we only pulled EHR notes that included one of the cannabis-related terms from the lexicon. Although population declines occurred in 20 of 24 (83.3%) counties (Supplemental Figure 2B), nativity levels remained high in the Geisinger footprint (Supplemental Figure 2C), supporting the use of this population for the extended period employed in this investigation.
Lexicon Development and Dataset Construction
This report used snippets of encounter notes curated from the EHR, including nursing notes, appointment notes, problem list-based free-text documentation, visit reason documentation, and follow-up notes. (Supplement A – Data sources)
We developed a CU lexicon to narrow the scope of data retrieval. The CU-related terms in the lexicon were finalized based on prior literature [26,27] and input from the study team’s subject matter experts (MPD, CKK, BJP). The lexicon included the following terms: ‘marijuana,’ ‘cannabis,’ ‘MJ,’ ‘THC,’ ‘CBD,’ ‘weed,’ ‘MMJ,’ ‘indica,’ ‘sativa,’ ‘cannabinoid,’ ‘spice,’ ‘tetrahydrocannabinol,’ ‘pot,’ ‘cannabidiol,’ ‘ganja,’ ‘grass,’ ‘hash,’ ‘hashish,’ ‘bong,’ ‘Mary Jane,’ ‘edibles,’, and ‘joint.’ Among the 1.76 million patients seen at Geisinger during the study period, we identified 1.71 million patients (135,858,323 EHR notes) with exact-word matches in their EHR notes using this lexicon. We then trimmed these EHR notes to retain only 300 characters before and after the CU-related term. We will refer to these trimmed snippets of notes simply as EHR notes in the context of this study moving forward. We observed that specific terms were used in contexts unrelated to cannabis; for instance, the term ‘pot’ was used in the context of ‘neti pot,’ ‘netti pot’, or as a part of the word ‘hypotension,’ ‘potential,’ or ‘potassium,’ while the term ‘CBD’ was used in an anatomical context for the ‘Common Bile Duct’ such as “CBD measure” and “CBD stricture”, these phrases and words were removed while data cleaning. ‘CBD’ or ‘pot’, only when mentioned in the context of marijuana use, was retained in the final analytical sample for classification. The term ‘joint’ was also primarily and extensively used in an anatomical context, such as “joint pain,” “joint space,” “joint deformity,” and “joint disease”. To minimize such misclassifications, we removed EHR notes that contained these specific phrases or words (Supplement A – Primary Data Cleaning), resulting in a working sample of 17,789,648 EHR notes. The data cleaning process also involved the removal of notes when the term of interest was negated (e.g., “no cannabidiol use” or “tetrahydrocannabinol (THC): no.”) This led to a more concise final list of terms and a final analytical sample of 2,790,896 EHR notes with at least one mention of a key term in patient charts. We used SAS Enterprise Guide 8.3 for dataset construction and cleaning.
Manual Annotation Phase
We developed an annotation schema and decision rules for labeling CU in EHR notes. VS and AP manually reviewed and double-coded a random sample of 100 EHR notes using labeling categories informed by previous literature [26]. Labeling categories were refined following a discussion of the assigned codes, with further input from team members to resolve disagreements. (Supplement B – Labeling Guidelines) This review phase eventually led to the creation of six categories: Three annotators were trained using the decision rules constructed for each labeling category and by reviewing the 100 initially sampled EHR notes. (Supplement C – Decision Rules) Each annotator subsequently independently reviewed a new random sample of 50 EHR notes (300 potential labels) (Supplement D – Examples), and we calculated an inter-rater reliability kappa for each category between each reviewer. (Supplement E – Inter-rater reliability kappa) Annotators then labeled EHR notes independently, generating 3,650 unique training examples. Throughout this phase, annotator questions regarding uncertain labeling decisions were discussed and resolved among the study team’s experts to refine category definitions and decision rules further. The annotated sample was drawn based on cannabis-related terms in the lexicon and was not adjusted for any patient characteristics. All annotation work was conducted using Microsoft Excel™.
- True mention – The identified key term in the EHR notes refers to cannabis
- Patient use – The true mention applies to the patient’s use, not another person
- Indication of use – The EHR note explicitly mentions or indicates that the patient has used cannabis at some point
- Medical use – The EHR note explicitly mentions that the patient used it for a medical purpose
- Current use – The EHR note explicitly mentions that the patient currently uses
- History of use – The EHR note explicitly mentions that the patient has a history of use
Model Development and Training
We trained two traditional models and two transformer-based models to identify CU in EHR notes: (1) logistic regression (LR), (2) support vector machine (SVM), (3) Bidirectional Encoder Representations from Transformers (BERT) [35], and (4) Bio-ClinicalBERT [36]. Prediction of each CU category was treated as a separate binary classification task for each model. Model development and training were conducted using Python 3.10.13.
For the LR and SVM models, we used a bag-of-words approach using unigrams, bigrams, and trigrams to represent each EHR note. We removed stop words and built bag-of-words representations using the sci-kit learn library [37]. We used elastic net regularization for dimensionality reduction in the LR models, with an L1 ratio of 0.5, which applies both L1 and L2 penalties equally. For SVM, a linear kernel was applied. We tested regularization weights of 0.01, 0.1, 1, 10, and 100 for LR and SVM. We used a regularization weight of 0.01 for LR models and 10 for SVMs, as model performance was highest with these weights (based on f-score across all CU categories). We conducted 5-fold cross-validation to train and evaluate models, averaging performance metrics across each test fold. Precision, recall, and F-score were used to assess model performance.
For BERT and Bio-ClinicalBERT, we used each model’s pre-trained transformer layers and tokenizers available through Hugging Face’s transformer library [36,38]. Models were fine-tuned on the corpus of labeled EHR notes. The training ran for ten epochs with a batch size of 8, and the AdamW algorithm was used for optimization. We performed 5-fold cross-validation for training and evaluation, and used precision, recall, and F scores averaged across each test fold as performance metrics. Precision, or more commonly known as positive predictive value, is the proportion of positive results that are true positives in the predicted sample [39]. It answers the question, “if the patient is classified as a CU, what is the probability they actually engage in CU.” [39] Recall, also known as sensitivity or true positive rate, quantifies how well a test identifies true positives (i.e., correctly identifies subjects who truly have the condition of interest)[39]. An F-score combines precision and recall into a single score and is used to evaluate the performance of the machine learning model. All BERT and ClinicalBERT models were implemented using Pytorch 2.1.0. Our model training and testing scripts are available at [40].
CU prevalence and patient characteristics
After evaluating each trained model, we used the highest performing model (in terms of weighted F score) to assign ‘True Mention’ and ‘Indication of Use’ classifications to all unlabeled EHR notes. Patients who had an EHR note positively classified for both the ‘True Mention’ and ‘Indication of Use’ categories were flagged as cannabis users. Any patient with at least one EHR note being flagged positive for CU was subsequently identified as a cannabis user for the remainder of the study period for descriptive analysis purposes.
We assessed the prevalence of CU documentation in the EHR using the CU flag. We computed descriptive statistics in patients flagged for CU at the time when the patient was first identified as a cannabis user, and in the overall population at the time of their first encounter during the study period. We calculated means or medians with SDs or interquartile ranges (IQRs) for continuous variables and frequencies and percentages for categorical variables, when appropriate, to evaluate the distribution of patients’ demographic and clinical characteristics including current substance use (smoking, alcohol, and illicit drug use), and BMI in the annotated sample, overall and CU populations. We further assessed the magnitude of difference in the demographic and clinical characteristics between the annotated sample, CU, and the overall population using t-tests and chi-square tests. All analyses were performed using SAS statistical software (SAS Enterprise Guide 8.3, Cary, NC).
Results
We retrieved 2,790,896 EHR note snippets from 370,885 patients; 98.7% were captured within nursing notes. ‘Marijuana’, along with its misspelling (‘Marjiuana’), was the most used terminology (56.7%), followed by ‘CBD’ (13.3%), ‘Cannabis’ (7.9%), ‘THC’ (7.5%), and ‘Pot’ (5.1%) (Table 1). From the 3,650 EHR notes in the annotation set, 1,945 (53.3%) had a ‘True mention’ of CU. Among the ‘True mentions’, 1,911 (98.3%) were explicitly discussing the patient’s personal use at the time of the encounter. 1,476 (77.2%) had documentation about the ‘Indication of use,’ 359 (18.8%) indicated medically recommended use, 973 (50.9%) had documentation indicating ‘Current Use,’ and 551(28.8%) had a match indicating ‘Past Use.’ A high kappa (> .83) was observed between the three annotators for all the labeling categories in the training sample, except for “History of use” (Supplement C – labeling categories). Supplement Table D includes example snippets for each labeling category.
Classification of CU was close to human performance for the ‘true mention’ and ‘indication of use’ categories across all models, with an F-score range of 93.2 % to 95.7% and 84.2% to 88.0%, respectively (Table 2). ClinicalBERT achieved the highest performance for the ‘true mention’ and ‘indication of use’ categories with a weighted average F-score of 92.4%. Except for the ‘patient use’ category, the performance thresholds were comparatively lower for all models for the ‘medical use,’ ‘current use,’ and ‘history of use’ categories. For these remaining categories, SVM performed slightly better than ClinicalBERT (weighted average F-score 81.6% vs. F-score 80.8%). SVM also performed the best for identifying medical use (P-82.8%, R-58.7%, F-68.4%) and current use (P-78.1%, R-68.1%, F-72.7%), while ClinicalBERT performed the best for identifying patient use (P-95.4%, R-95.1%, F-95.3%) and past use (P-57.9%, R-74.1%, F-64.5%) (Table 3).
The trained ClinicalBERT models for ‘True Mention’ and ‘Indication of use’ subsequently classified over two-fifths (40.6%, 150,726) of the patients from the analytical sample as positive for CU. This represented 8.6% of the overall population. The distributions of all demographic variables were statistically different between the annotated sample, CU documented population, and the overall population (p < 0.0001). The average age of patients with CU documentation was similar to that of the overall population (CU - 37.6 (SD – 18.8) vs. overall - 38.1 (SD-24.6)). The proportion of patients over 65 (CU-9.3% vs. overall-17%) and under 18 (CU-11.1% vs. overall-24.6%) was higher in the overall population than in patients with CU documented in their EHR. Additionally, just over half of patients with CU were male (52.5%), compared with the overall population, where just over half were female (52.3%). Patients with CU were less likely to be Asian and more likely to be African American than the overall population. Nearly two-thirds of cannabis users (62.8%) had a Body Mass Index (BMI) greater than 25, and of these, 36.8% had a BMI > 30 (i.e., obese). The proportion of obese patients was lower in the overall population (23.9%). The average BMI of cannabis users was also greater than the overall population (CU mean= 28.5, SD=8.6 vs. overall mean= 26.9, SD=9.0, Table 4). The distribution of some patient characteristics also differed between the annotated sample and those for CU and the overall population, as the random sample used for annotation was based on CU-related terminology of interest. (Table 4). Upon comparing substance use behaviors between cannabis users and overall population, we observed that CU were unlike the overall population in their current or history of use of various drugs in that they were ten-fold more likely to report tobacco smoking (49.3 % vs. 5.1%), ten-fold more likely to report alcohol use (48.2% vs. 4.9%), and nine-fold more likely to report illicit drug use (4.7% vs. 0.5%, Fig. 2).
Discussion
This study used NLP algorithms to identify patients within an integrated health system over a ten-year period who had CU documented in their EHR notes. This novel NLP application builds on prior reports [26–32] and demonstrates that CU can be identified with high precision and recall from unstructured data sources using advanced NLP models. While individual model performance varied across cannabis-related categories, the models generally performed well. ClinicalBERT, a transformer-based model, and a more traditional machine learning model, the SVM, performed nearly equally well for CU identification. Despite variations in each model’s performance for the other classification tasks, such as medical, past, current, and patient use, ClinicalBERT and SVM consistently outperformed BERT and LR. Based on these models, the prevalence of CU-related documentation was 8.6% in the overall cohort curated using only the keyword search strategy.
EHR notes and other non-discrete fields are a highly sought option for documenting cannabis and non-allopathic substances and medicines, thereby making them a key source of information for understanding utilization patterns [27,41–43] [27,43–45]. The current analysis provides the first NLP-driven, overarching insight into unsegmented EHR-based unstructured notes data on CU documentation within a single health system. This study also includes a diverse cross-section of racial breakdown in rural (e.g., Sullivan County, population 5,845) and urban (e.g., Luzerne County, population 326,496) populations (Supplemental Figure 2D). The use of traditional and transformer-based NLP models to identify CU documentation within different types of EHR notes and other unstructured data sources is rapidly increasing [22,23,26,29,44,45] and models like ours could be applied and adopted across different health systems. This investigation’s rigorous manual annotation phase, with a high interrater reliability (kappa >.83), also ensured consistency in the classification of CU. Furthermore, because the models were trained on various EHR notes spanning different specialties and disease states, as well as documenting providers, the algorithms could be easily applied to further targeted clinical and epidemiological research across the system and across different disease states.
A key report from an integrated health system in Washington state determined that for every person with past-year marijuana use documented in their EHR, seven others used marijuana, but who would only disclose this on a survey [20]. CU identified in our study over 9.5 years (8.6%) was less than half that reported in PA in the 2021-2022 National Survey of Drug Use and Health among those aged > 12 for marijuana use in the past year (19.2%) [46] affirming similar incomplete documentation as noted in Washington State. Despite the recent push for cannabis being reclassified as a Schedule III substance [47] and potentially soon to be reported through state’s prescription drug monitoring program, self-reported CU information will continue to be recorded in the EHR. This is because medical use accounts for only a small portion of the total use among high schoolers [1,2] as well as adults [20], and only 24 states, three Territories, and Washington, D.C. have legalized adult use, thereby resulting in a continued reliance on self-reported CU information for assessing prevalence [5,48]. A hesitancy to disclose CU is perhaps understandable as patients might perceive that this documentation could negatively impact one’s ability to possess a firearm [49], employment [50], or even child custody [51]. The authors of an EHR report in California from 2013 to 2017 (i.e., before adult use passage) of primary care patients interpreted their social history CU documentation (0.38%) as “surprisingly low” [52] and at least ten-fold below what might be anticipated based on past-year use. Even within this context of persistent cannabis stigma [53] and under-reporting of tobacco [52] and alcohol use [Supplemental Figure 3, 4], the pronounced higher rates of tobacco, alcohol, and other illicit drug use among CU relative to the overall population are clear. The higher prevalence of CU tobacco, alcohol and other illicit drug users aligns with findings from literature[54,55], There is also some evidence that indicates that CU is more likely to be documented in the EHR for substance users compared to non-users, indicating potential biased approach towards not just who is asked or tested for CU, but also when it is documented.[29]. We could potentially hypothesize that CU is higher among those who report illicit drug use, in part, because individuals who already disclose illicit drug use may feel less need to hide or underreport their CU. However, more research is needed to validate this assumption.
Demographics and clinical characteristics differed between our annotated sample, the overall population, and CU patients. However, the clinical interpretation of these findings may not be clear due to potential confounding, given the large sample sizes, and focus on having an equitable distribution of CU-related terminologies in the annotated sample over patient characteristics [56]. For example, in our study, we found a higher proportion of obese patients in the CU cohort, which is in contrast with other studies, which found that compared to non-users, the prevalence of obesity and being overweight was lower among cannabis users in the US compared to non-users due to the increase in metabolic rates and consequent reduction in BMIs [57–60]. However, in our analysis, we did not stratify patients based on length and frequency of CU, nor did we adjust for the higher prevalence of alcohol intake in this group, or potential clinical comorbidities that may be associated with high obesity prevalence, all of which may have influenced this difference in prevalence observed [59,61]. Furthermore, we observed a higher prevalence of CU documentation within African Americans and a lower prevalence amongst Asians. While these findings align with prior research, there is existing evidence that indicates that preconceived biases, stigma, and racial profiling and discrimination surrounding who is asked about or tested for CU may be a reason behind the higher prevalence observed[29,62].
The EHR landscape is rapidly evolving, with the diffusion of agentic and generative AI-based tools implemented across multiple platforms to collect and transform data, predict potential health outcomes, and improve documentation. As the landscape continues to evolve, healthcare systems will need to develop internal capacity to tailor and fine-tune universal models to meet local and system-specific needs. In the context of CU, models developed as part of this study raise the potential for this technology to improve the current capture of CU within EHR systems and to fine-tune newer models in the future. However, it is also noteworthy that the NLP models in the current analysis performed poorly while classifying medical use of cannabis, with F1 scores ranging from 45% to 68%, indicating that the algorithms accurately classified medical use only about half the time. The distinction between medical use and non-medical use is legally clear but may have limited epidemiological utility. Among primary care patients in Washington state who indicated cannabis use, 40.1% categorized their use as medical, 31.8% as nonmedical, and 28.2% as both medical and non-medical [20]. Similarly, among medical patients in New England surveyed before the passage of laws for legalizing medical use, about half described their use as on a continuum that involved at least some recreational use [7, Supplemental Figure 4][17]. Among adults living in a state that condones medical use, the most common source of cannabis (57.5%) was from a non-dispensary source [63]. Together, these findings indicate that the distinction between medical and recreational CU is not clear.
Our investigation had several limitations. First, the results from this study are based on data collected from a single health system. CU documentation within a system could be influenced by the legalization status of cannabis at the time of data collection (i.e., currently condoning medical but not recreational in PA) [64], health system-level regulations, equity concerns about cannabis documentation attributable to racial diversity (Supplemental Figure 2D [29,65], and clinician-level characteristics, including but not limited to personal biases. Second, while using fivefold cross-validation improved the internal validity of these models, they were not externally validated, thereby limiting their generalizability for use at external institutions. Third, we identified the EHR notes sample using a keyword search for terminologies. There is still a possibility that some terminologies, variations, or misspellings could have been missed (e.g., “marijuana” can be misspelled as “marihuana” or “marjiuana” or “marajuana”), impacting the number of notes identified and restricting the sample. We curated literature and sought stakeholder input to assemble a comprehensive lexicon that would minimize the impact of this limitation. Additionally, based on our models, we could not conclude medical use or distinguish between current and past use with certainty due to poor model performance. This could be attributed to semantic variability in the documentation of this information and to low coder agreement during the manual annotation phase when classifying these factors. Prior research has encountered similar difficulties when leveraging NLP algorithms [26,27]. Consistency in documenting practices across the system may help address ambiguity in reported information and improve model performance. Lastly, data captured in the EHR remains susceptible to the general limitations of EHR data, such as inaccuracy, incompleteness, and insufficient granularity [66]. System-level initiatives that ensure regular data quality checks could help overcome this.
Conclusion
Our study successfully leveraged NLP methodologies to classify CU documentation from EHR notes with high precision. System-level initiatives supporting automated extraction of this information using NLP and its potential transformation into data that can be discretely queried, as well as leveraging the rapidly evolving artificial intelligence capabilities within the EHR, may lead to an improved understanding of patients’ health and facilitate future clinical and epidemiological research surrounding CU prevalence and patterns among different patient populations.
Supporting information
Acknowledgement
We would like to thank our Summer Research Immersion Program student, Ms. Alivia Roberts, and Ms. Faith Garasich, a Summer Undergraduate Research Program student, for their assistance in the manual annotation work on this study, and Wei Wei, PhD, and Joseph J. Dewalle, for fruitful discussions, as well as Megan Gunther and Iris Johnston for technical support. Vanessa Troiani is currently the director of the Geisinger’s Academic Clinical Research Center, which operates through the Center for Substance Use Research and Education.
Statement of Ethics
This study protocol was reviewed and approved by the Geisinger IRB (approval number 2022-0498).
Conflict of Interest Statement
The authors had no conflicts of interest to report during the study period.
Funding Sources
The study is funded by Ascend Wellness Holdings, Clinical Registrant through the Geisinger Academic Clinical Research Center. The funder was not involved in the study or the decision to submit for publication.
Data Availability Statement
Data and models may be shared upon reasonable request in the future, contingent on IRB approval. Specific requests for access to the data should be directed to the corresponding author, who will facilitate the process, contingent upon approval and adherence to the necessary ethical and institutional guidelines.