Artificial Intelligence Approximates Human Affect Ratings of Cannabis Images
aCenter for Technology and Behavioral Health, Dartmouth Geisel School of Medicine, Lebanon, NH, USA
bDepartment of Biomedical Data Science, Dartmouth Geisel School of Medicine, Hanover, NH, USA
cDepartment of Psychological Sciences, Florida State University, Tallahassee, FL, USA
dDepartment of Computer Science, Dartmouth College, Hanover, NH, USA
eDartmouth College, Hanover, NH, USA
fVirginia Polytechnic Institute and State University, Blacksburg, VA, USA
*Corresponding Author: Jacob T. Borodovsky, PhD., Email: Jacob.borodovsky@dartmouth.edu, 46 Centerra Pkwy, Lebanon, NH 03766, (603) 646-7000Abstract
Cannabis imagery is proliferating online and can elicit affective responses related to use. Scalable tools are needed to evaluate how this proliferation could influence population health. This pilot study tested whether multimodal generative artificial intelligence (MGAI) can reproduce subjective human affect ratings of cannabis images. Four MGAI agents (model: gpt-4o-2024-11-20) were created to parallel the four human participant subgroups from Macatee et al. 2021, defined by primary method of cannabis administration (bong, bowl, joint/blunt, vaporizer). Using Macatee et al.’s participant instructions and standardized image set, each agent rated images of its primary method of administration on valence, arousal, and urge constructs. For each image-construct pair, n=100 ratings were generated in separate conversational threads using zero-shot prompting. Image-level MGAI mean ratings were compared with human mean ratings using Two One-Sided Tests of equivalence and Spearman correlations. Although formal statistical equivalence was rare (4% valence, 11% arousal, 3% urge), MGAI ratings approximated human ratings closely (Mean difference of mean ratings = – 0.31, SD = 1.23) and correlations between MGAI and human mean ratings were moderate to high: rs(valence) = 0.55, rs(arousal) = 0.34, rs(urge) = 0.56. MGAI also reproduced the parabolic relation between rating means and standard deviations observed in human data. These preliminary results indicate that MGAI can approximate human cannabis cue-reactivity patterns closely enough to justify continued refinement. MGAI could potentially be developed into a Cannabis Regulatory Science tool to aid regulatory oversight of online cannabis marketing.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
Funding for this study was provided by grants from the National Institute on Drug Abuse (NIDA) R01DA050032, R21DA062816, P30DA037202
INTRODUCTION
The US legal cannabis industry continues to expand its e-commerce sales and online advertising. Consequently, images of cannabis-related products and paraphernalia (e.g., bud, vapes, edibles), brands, logos, and lifestyles have become pervasive online (Berg et al., 2023; Carlini et al., 2022; Cavazos-Rehg et al., 2019; Duan et al., 2023). When individuals use cannabis, the sensory stimuli present during use (e.g., sight of cannabis bud, lighter, pipe) become associated with the rewarding effects of intoxication and come to function as conditioned cues that independently elicit craving. When these cues are depicted in digital images, they can re-activate conditioned associations and elicit similar craving responses (Aston et al., 2015; Cousijn et al., 2013; Field et al., 2006; Field & Cox, 2008; Henry et al., 2014; Lundahl & Greenwald, 2015, 2016; McRae-Clark et al., 2011; Metrik et al., 2016; Norberg et al., 2016; Ruglass et al., 2019; Strickland et al., 2019; Vafaie & Kober, 2022; Wölfling et al., 2008). Marketers are aware of these dynamics and leverage them in advertising strategies to shape affective responses (Martin et al., 2013; Wang et al., 2023).
Compliance with cannabis advertising regulations remains low across many states with legal markets (Cao et al., 2020; Carlini et al., 2022; Marinello et al., 2024; Moreno et al., 2022; Sheikhan et al., 2021). The rapid expansion and decentralized structure of online cannabis marketing make oversight difficult and allow noncompliant content to persist (Caputi et al., 2018; Y. Cui et al., 2024; Duan et al., 2023). As a result, advertising often continues to feature salient cues that leverage use-related associations. Scalable tools that can navigate large digital ecosystems are therefore needed for effective monitoring of online cannabis imagery.
Multimodal generative artificial intelligence (MGAI) could be useful for regulatory oversight of aggressive online cannabis marketing. Recent studies show that MGAI can identify alcohol and nicotine products in digital images (Bonela et al., 2023; Kuntsche et al., 2023; Vassey et al., 2025). These studies represent important progress, but they have not focused specifically on cannabis imagery and have emphasized product detection rather than psychological impact. There is growing evidence that MGAI models can reasonably approximate various dimensions of human psychology (Argyle et al., 2023; Bail, 2024; Z. Cui et al., 2025). If MGAI can reproduce human affective responses to cannabis cues, it could potentially be used to help regulators triage review of online advertisements.
The purpose of the present study was to understand the extent to which, and ways in which, MGAI can approximate aggregate human affective ratings of cannabis cues. We instructed MGAI to provide ratings of valence, arousal, and urge for images in Macatee et al.’s (2021) standardized cannabis cue image set (Macatee et al., 2021). MGAI ratings were then compared with the corresponding human data reported in Macatee et al. We hypothesized that MGAI mean ratings would show statistical equivalence and strong rank order correlation with human mean ratings.
METHODS
Original Study Design
Macatee et al. (2021) developed and validated a standardized cannabis cue image set for cue-reactivity research. The image set contained 280 cannabis-related images and 80 neutral tooth-brushing images matched on visual and motor characteristics. Twenty cannabis images depicted cannabis flower alone, and 260 images depicted four methods of administration: bong, bowl, joint/blunt, and vaporizer. Each method-of-administration set consisted of 60 images distributed across three subtypes (20 per subtype): object-only images (e.g., bong shown alone), object-with-hand images, and object-with-face images (e.g., bong being used). The vaporizer category included 20 extra images to capture the diversity of vape products available. Participants were assigned to one of four subgroups based on their primary method of administration. Each participant rated the neutral and cannabis-flower images, and the images corresponding to their primary method (e.g., bong users rated bong images). Participants rated each image three times, once for each of the three constructs: valence, arousal, and urge. Ratings used a 1–9 Self-Assessment Manikin scale (Bradley & Lang, 1994) with labeled anchors at 1, 5, and 9. The present study used the object-only subset of method-of-administration images. MGAI agents provided ratings on the same constructs and scale, and MGAI mean ratings were compared with human mean ratings of the same object-only subset of images.
MGAI Agents
We created four MGAI agents to parallel the four Macatee participant subgroups (bong, bowl, joint/blunt, vaporizer). Each agent had fixed “system instructions” that reflected the aggregate cannabis use characteristics (primary method of use, age at first use, age at regular use, and years of regular use; Figure 1, “Agent Setup”) of the corresponding human subsample from the Macatee study. Note that system instructions are different from the “prompt instructions” given to the agent. Verbatim system instructions for all four agents are provided in the appendix. All agents used GPT-4o (model: gpt-4o-2024-11-20), default hyperparameter values, and were queried via the OpenAI® “Assistants” API.
Statistical Analysis
We first tested statistical equivalence of MGAI and human mean ratings at the image–construct level using Two One-Sided Tests (TOST). For each image–construct pair, we calculated the pooled standard deviation of MGAI trial-level ratings and human participant-level ratings. We multiplied the pooled standard deviation by 0.5 to define equivalence bounds (k = 0.5). We set α = 0.05. Equivalence required both one-sided p-values < 0.05 and a 90% confidence interval that fell entirely within the equivalence bounds. We planned 240 equivalence tests (20 images × 3 constructs × 4 method-of-administration categories). For example, the mean MGAI valence rating for bong image #1 was compared to the mean human valence rating for bong image #1. One test (valence for images V7) could not be conducted because MGAI trial-level standard deviations equaled zero. No alpha corrections were applied as we treated each image-construct comparison as an individual hypothesis (Rubin, 2021). We also compared MGAI and human ratings using Spearman correlation (rs) within each construct, and within each method-of-administration category. All analyses were performed in R using dplyr for data wrangling, purrr for iteration, and TOSTER for equivalence testing. This study was deemed exempt from human subjects review by the Dartmouth Committee for the Protection of Human Subjects.
RESULTS
TOST-based equivalence of Human and MGAI ratings was rare. Overall, 4%, 11%, and 3% of MGAI’s mean valence, arousal, and urge ratings, respectively, were statistically equivalent to human mean ratings. When stratified by method of administration, approximately 5%, 8%, 3%, and 7% of MGAI’s mean ratings for bowl, vape, blunt/joint, and bong, respectively, were statistically equivalent to human mean ratings. We also observed that approximately 98%, 90%, and 3% of the differences (MGAI mean minus human mean) in valence, arousal, and urge ratings, respectively, were negative, indicating that MGAI consistently underestimated valence and arousal ratings and overestimated urge ratings.
Despite the lack of TOST-based equivalence, differences between MGAI mean ratings and human mean ratings were small (mean difference of means = -0.308, SD = 1.230)(Figure 2a), and the Spearman correlations between MGAI mean ratings and human mean ratings were high: valence (rs = 0.552), arousal (rs = 0.335), and urge (rs = 0.562). When stratified by method of administration, Spearman correlations between MGAI mean ratings and human mean ratings were: bowl (rs = 0.203), vape (rs = 0.004), blunt/joint (rs = 0.287), and bong (rs = 0.441). Notably, MGAI rating standard deviations were markedly lower than human rating standard deviations. Average standard deviations were: valence (MGAI: 0.340 vs. human: 2.019), arousal (MGAI: 1.164 vs. human: 2.428), and urge (MGAI: 0.561 vs. human: 2.341). However, the relationship between the mean and standard deviation of ratings showed similar parabolic relationships for MGAI and humans (Figure 2b).
DISCUSSION
We tested whether multimodal generative artificial intelligence (MGAI) could reproduce human ratings of cannabis images across four methods of administration (bong, bowl, joint/blunt, vaporizer) on three psychological constructs (valence, arousal, urge). Statistical equivalence of mean MGAI ratings and mean human ratings was rare. However, MGAI–human differences in mean ratings were small, and the rank ordering of mean ratings of constructs observed in humans (valence < arousal < urge) was well replicated by MGAI both overall and within method-of-administration categories. Thus, the lack of statistical equivalence seems to be due to smaller standard deviations of MGAI ratings rather than large MGAI–human discrepancies in mean ratings. Although further work is needed to clarify the sources of MGAI–human discrepancies, these findings provide initial evidence that MGAI could be developed into a Cannabis Regulatory Science (Borodovsky et al., 2021; Borodovsky & Budney, 2018) tool for assisting regulatory review of online cannabis advertising.
Across method-of-administration categories, MGAI preserved the relative rank order pattern of mean ratings of constructs observed in the human data. However, MGAI systematically “amplified” that pattern by producing mean ratings farther from the midpoint of the scale (Figure 2a). Several recent studies have reported similar findings (Alrasheed et al., 2025; Romeo & Testolin, 2025; Shirahama et al., 2025). Despite this amplification bias, MGAI still reproduced the parabolic relation between rating means and standard deviations observed in our human sample and other human samples (Pollock, 2018). This correspondence is noteworthy because, in principle, MGAI could have produced mean ratings at the scale midpoint with narrow standard deviations (e.g., MGAI rates the image as a 5 in all 100 trials).
Several limitations qualify these findings. First, procedural differences between the MGAI and human tasks may have influenced results. MGAI agents were required to produce a five-sentence rationale before each rating (Chu et al., 2024) and used text-based anchors, whereas participants in Macatee et al. (2021) were not asked for a rationale and provided ratings using Self-Assessment Manikin pictorial scales. Second, all 100 MGAI trials for a given image– construct pairing originated from the same agent identity across separate threads. Using a larger pool of independent agents, with each agent providing a single rating per image, could have better approximated the distributions of independent human raters. Third, in the absence of a field standard, we relied on a conservative k = 0.5 bound for our TOST analysis, which strongly influences whether a result is labeled “equivalent”. Fourth, we evaluated only a single MGAI platform and model (ChatGPT, version GPT-4o-2024-11-20); performance may differ across other commercially available MGAI systems (e.g., Gemini, Claude) or future OpenAi ® models.
These preliminary findings indicate that MGAI can approximate average human cannabis cue-reactivity patterns closely enough to justify continued methodological testing and refinement. With further development under a Cannabis Regulatory Science framework, these tools could be used to aid Federal or State oversight of online cannabis-based product advertisements (Food and Drug Administration, 2025; Padon et al., 2025) that violate public health regulations.
Supporting information
Data Availability
All data produced in the present study are available upon reasonable request to the authors
Acknowledgments
Funding for this study was provided by grants from the National Institute on Drug Abuse (NIDA) R01DA050032, R21DA062816, P30DA037202. The funding organizations had no role in the design and conduct of the study; collection, management, analysis, and interpretation of the data; preparation, review, or approval of the manuscript; and decision to submit the manuscript for publication. Author JTB employed ChatGPT and Grammarly to help refine grammar, sentence structure, and word choice. The author(s) reviewed and retain full responsibility for the manuscript’s content.