Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study
Department of Radiology & Imaging Sciences, Indiana University School of Medicine, Indianapolis, IN, United States of America
Department of BioHealth Informatics, Indiana University Luddy School of Informatics, Computing, and Engineering, Indianapolis, IN, United States of America
Department of Radiology and Imaging Sciences, Emory University School of Medicine, Atlanta, Georgia, United States of America
Department of Radiology and Biomedical Imaging, University of California San Francisco, San Francisco, California, United States of America
* Corresponding author Email: jolburns@iu.edu (A1)Abstract
Radiology specific clinical decision support systems (CDSS) and artificial intelligence are poorly integrated into the radiologist workflow. Current research and development efforts of radiology CDSS focus on 5 main interventions, based around exam centric time points– at time of radiology exam ordering, after image acquisition, intra-report support, post-report analysis, and radiology workflow adjacent. We review the literature surrounding CDSS tools in these time points, requirements for CDSS workflow augmentation, and technologies that support clinician to computer workflow augmentation.
We develop a theory of radiologist-decision tool interaction using a sequential explanatory study design. The study consists of 2 phases, the first a quantitative survey and the second a qualitative interview study. The phase 1 survey identifies differences between average users and radiologist users in software interventions using the User Acceptance of Information Technology: Toward a Unified View (UTAUT) framework. Phase 2 semi-structured interviews provide narratives on why these differences are found. To build this theory, we propose a novel solution called Radibot - a conversational agent capable of engaging clinicians with CDSS as an assistant using existing instant messaging systems supporting hospital communications. This work contributes an understanding of how radiologist-users differ from the average user and can be utilized by software developers to increase satisfaction of CDSS tools within radiology.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
John Burns - No disclosures or competing interests. Marc Kohli - Travel support from SIIM and RSNA during the study period. Co-Founder and Shareholder in Alara Imaging, which was not involved in the study and sells products related to quality measures and edge-to-cloud gateways. Josette Jones - No disclosures or competing interests. Saptarshi Purkayastha and Judy W. Gichoya are funded by US National Science Foundation (grant number 1928481) from the Division of Electrical, Communication & Cyber Systems. Judy W. Gichoya is also funded by the National Institute of Biomedical Imaging and Bioengineering (NIBIB) MIDRC grant of the National Institutes of Health (75N92020C00008 and 75N92020C00021). Saptarshi Purkayastha has no competing interests with regards to this work. Judy W. Gichoya has no competing interests with regards to this work.
1Introduction
Radiology domain-specific clinical decision support systems (CDSS) applications are poorly integrated into the radiologist workflow (1). In 2017, Dreyer and Geis described a transition in radiology moving towards integrating Artificial Intelligence (AI) into the radiologist workflow. "In the past, radiology was reinvented as a fully digital domain when new tools, PACS and digital modalities, were combined with new workflows and environments that took advantage of the tools. Similarly, a new cognitive radiology domain will appear when AI tools combine with new human-plus-computer workflows and environments." They describe the concept of a "Centaur Radiologist" as a physician utilizing AI-augmented CDSS workflows to increase efficiency (2). We expand this term as “future radiologist,” inclusive of non-AI techniques in CDSS.
However, the future radiologist will not happen if the tools are poorly integrated, with cumbersome human-computer interfaces (3). Deliberate and sustained effort by using inter-disciplinary knowledge from human-centered computing, psychology, cognitive sciences, and medicine is required to build CDSS for the future radiologist (4). In this work we create a basis of knowledge in the theory of radiologist-decision tool interaction using a sequential explanatory study design. The study consists of 2 phases, the first a quantitative survey and the second a qualitative interview study. The phase 1 survey identifies differences between average users and radiologist users in software interventions using the User Acceptance of Information Technology: Toward a Unified View (UTAUT) framework (5). Phase 2 semi-structured interviews provide narratives on why these differences are found. To build this theory, we propose a novel solution called Radibot - a conversational agent (CA) capable of engaging clinicians with CDSS as an assistant using existing instant messaging (IM) systems supporting hospital communications. This work contributes an understanding of how radiologist-users differ from the average user and can be utilized by software developers to increase satisfaction of CDSS tools within radiology.
1.1Background
We expect that the future radiologist will routinely interact with CDSS at each stage of their workflow. We designed Radibot for diagnostic radiologists, with interventions at each of the following time-points: after image acquisition, during report creation, after report creation, and between studies. A brief overview of existing interventions in each time point follows. Diagnostic radiologist’s clinical work is fully completed using systems, including PACS, RIS, and VR, with every interaction being digitally augmented (40). Given the fully digital clinical workflows, radiology specific CDSS implementations are uniquely positioned to provide support and affect change. Radiology specific guidelines for "advisor systems" were laid out by Teather et al. in 1985, while Khorasani in 2006 provides features for the development of clinical decision support systems (41, 42). Outside of radiology, CDSS are built following the Ten Commandments for Effective Clinical Decision Support: Making the Practice of Evidence-Based Medicine a Reality (43). Commandments 2, 3, 7, 10, and 1 – anticipate needs, fit into user workflow, simple interventions, knowledge system maintenance, and speed – appear with a higher frequency when aligned with radiology specific guidance. An alignment of the general CDSS and radiology specific CDSS guidelines are found in table 1. Differences in CDSS priorities underscore the need for more research in this area and are mapped to UTAUT concepts and the hypotheses for phase 1.
- After Image Acquisition - radiologists combine a variety of data to make interpretations of images. Interventions include Computer-Aided Detection (CAD), Computer-Aided Diagnosis (CADx), and patient history/metadata presentation. These interventions generally function within the Picture Archiving and Communication System (PACS), though some will interface with the Radiology Information System (RIS), Voice Recognition system (VR), or in an external client (6-16).
- During Report Creation – these interventions surround embedding evidence-based guideline processes during dictation and are found within VR. Guidelines are navigated using drill-through commands or natural language processing (NLP) of the dictation to generate report text (17, 18).
- After Report Creation - In most RIS, reports are stored as unstructured text. Interventions in post-report analysis include extracting categorical data, automating radiologist-clinician communication, and quality improvement systems. By generating summative report metadata, these interventions enable context-switching and reduce fatigue when a radiologist is asked to return to a finished report (19-31).
- Between Studies – existing adjacent to radiologist workflow, these interventions influence decision making at an individual or business level and consist of workflow-prioritization, management, and feedback tools. These tools utilize metadata found in Health Level 7 (HL7) or Digital Imaging and Communications in Medicine (DICOM) messages. Users interface with them outside of clinical systems, IE. web dashboards, or they are integrated into PACS/RIS/VR presentation layers (32-39).
1.2Instant Messaging and Conversational Agents in Healthcare
IM is found throughout the healthcare enterprise, including in disease management, patient-clinician interactions, medical education, among patient populations and workforce members for extra-clinical activities. IM can be inclusive of voice, video calling, and file sharing (44). Extra-clinically, IM tools facilitate socialization, catharsis, and professional connectiveness functionalities when applied in clinical settings (45, 46). IM is asynchronous and short-form, leading to advantages over other communication methods, particularly in the area of articulation work - answering medical questions, coordinating logistics, addressing social information for patients, and querying staff/equipment locations or status (47). IM is integrated into many PACS, RIS, and VR, serving many purposes within radiology including care discussions and facilitating remote tele-radiology communications (29, 48-58).
CA, or chatbots, are natural language human-machine interfaces. CA can apply 4 methods for negotiating user interactions: immediate, negotiated, mediated, and scheduled (59). Consumer health care CA are currently scheduling appointments, providing basic symptom identification and recommendation, and assisting with long term care such as sensor monitoring/alerting and medication reminders (60). Most healthcare CA are built for patients (interview, data collection, or telemonitoring), while clinician focused CA are designed around data collection (61). Other efforts in clinician focused CA include interpreting spoken language into clinical facts and drug interaction/alternative drug recommendation systems (62-64). IM impact on task completion is not fully understood, especially in the context of automated IM interventions. There is evidence that non-relevant messages can increase or reduce task completion times depending on the message initiator; at a cost of quality of the task output (65). Disruptiveness of IM specific interventions is reduced when IM are relevant to the task being completed or if delivered at time-points that fit the users workflow (66). IM interactions among a professional workforce are found to support task completion, accuracy, and quality of outcomes (65).
2Methods
2.1Population
Our study population consists of 2 sets of radiologists – attendings and trainees at a large academic health system. The attendings set is a subset of the approximately 112 radiologist faculty at a teaching hospital system. The trainee’s set is a subset of the 62 residents/fellows within the same system. Our population is acquired through convenience sampling. Of 174 possible participants, 98 responded affirmative that they would complete the survey and 3 that they did not want to participate. 39 participants responded that they would complete an interview and 11 responded that they would participate in the survey but not the interview. In total, 94 surveys were submitted, and 23 interviews were conducted.
2.2Survey
An electronic survey was created using Qualtrics (67), we collect population composition and quantitative data surrounding intervention feasibility, usability, and acceptance. Within the UTAUT framework we focused on behavioral intention to use the system (BI), attitude toward using the technology (AF), effort expectancy (EOU), performance expectancy (PE) and anxiety (ANX). We chose to not utilize questions in social influence, facilitating conditions, and self-efficacy due to applicability to a prospective study of a tool not yet implemented in practice. A full listing of UTAUT questions by construct and factor are found on EDUTECH’s Wiki (68). Due to respondent time constraints we chose to utilize 12 of 19 questions in the chosen constructs, with each construct having at least 2 questions asked. Questions were eliminated if they were not relevant to a system that does not yet exist (Example: Working with the system is fun).
Other frameworks exist for testing usability and user experience for software design. However, UTAUT is unique in the number of constructs it can capture quickly. Measures like the System Usability Scale or Technology Acceptance Model can capture intent to use, but do not create the linkages to moderating factors of interest. Contrasting the UTAUT concepts with the CDSS commandments, we create the following links:
- Performance Expectancy
- 1. Speed is everything
- 2. Anticipate Needs and Deliver in Real Time
- 5. Recognize that physicians will strongly resist stopping
- 7. Simple interventions work best
- Effort Expectancy
- 3. Fit into the user’s workflow
- 4. Little things can make a big difference
- 6. Changing Direction is Easier than Stopping
- 8. Ask for Additional Information Only When You Really Need It
- Anxiety
- 5. Recognize that physicians will strongly resist stopping
The survey in full is included in the S1 Appendix A.1. Fig 1 highlights the intervention and proposed capabilities.
2.3Interview
Using the research statements developed with the survey (S1 Appendix A.5), we generated hypothesis and began developing the semi-structured interviews. As we did not have a working system, we prototyped 5 interventions and created video examples of each to use during the interview. Fig 2 highlights what these videos looked like during a demo. The videos highlighted interventions during each workflow time point in the following ways:
- After Image Acquisition
- Video 1 – Radibot identifies potential for 3d reconstruction, asks radiologist permission to process, and then suggests the correct VR template.
- During Report Creation
- Video 2 – Radiologist engages Radibot to query the Electronic Medical Record (EMR) for cardiac risk factors. Radibot performs this query as the radiologist returns to reviewing images, then returns all risk factors that meet these criteria.
- Video 3 – Radibot identifies VR dictation of left adrenal nodule then engages radiologist in stepping through adrenal nodule flow chart. Completion of the flowchart inserts guideline recommend text and citation into report.
- After Report Creation
- Video 4 – Based on text of report, Radibot engages radiologist for follow up communication.
- Between Studies
- Video 5 – Radibot presents possible studies for radiologists to engage with, removing the need to navigate the worklist. Includes suggestions of cross coverage of busier worklists and high priority studies.
An interview guide was created (S1 Appendix B.1) following the UTAUT framework. The guide begins with video 1, loops through each video asking the same questions, then has a set of questions after all videos have completed. A portable interview setup was created consisting of one laptop, a 4k portrait monitor mimicking a diagnostic monitor, and a microphone for collecting audio. Interviews occurred in offices/conference rooms located near interview candidates normal work locations. Subjects were presented with consent and informed that no names would be utilized during the interview for confidentiality. Zoom was utilized to record the screen and interview narrative to the laptop (73).
39 survey participants responded that they would complete an interview. 23 interviews occurred before the research team agreed that response saturation was achieved. Interviews were transcribed using Otter.AI, then a research assistant and study team member reviewed each video separately and corrected any transcription errors (74). Transcriptions were downloaded in docx format, then loaded into ATLAS.ti 9.0.19.0 for qualitative analysis. The study team created labels for text analysis (S1 Appendix B.2) and linked these by semantic domain (UTAUT construct). 2 research assistants were hired and trained by the study team to annotate interview text using ATLAS.ti. The research assistants separately annotated interview 1, then the study team reviewed and provided additional guidance. They then separately annotated the remaining interview narratives, and the annotated narratives were merged, and inter-rater agreement is measured. Because semantic domains are established and we did not segment quotes in advance, Krippendorff’s CU Alpha is utilized to measure semantic domain agreement by quote. An overall agreement level of α ≥ .8 is set for all documents (75).
3Results
3.1Survey Data Analysis
Resulting data was downloaded from Qualtrics in Comma Separated Values (CSV) format and analyzed using Excel. Irrelevant metadata fields were removed. A total of 88 survey responses were used for analysis, representing 50.6 percent of the total sample population. After removing 4 outliers that took over an hour to complete the survey, average completion time was found to be 6 minutes and 45 seconds. Raw survey data is available in S2 Survey Data.
Qualitative questions were bucketed into numbers ranging from 0-5 (IE 0 to 5 years = 1; 5 to 10 years = 2; etc.). A full set of questions, response bucketing, and UTAUT constructs are included the S1 Appendix A.2. Summary data surrounding survey responses used in the analysis are listed in table 2.
Partial Least Squares (PLS) Structured Equation Modeling (SEM) was utilized to investigate the relationship between constructs. PLS-SEM calculations were performed using SmartPLS V. 3.2.9. Complete data analysis steps are included in the Supplemental Data Analysis (S1 Appendix A.3). SEM began with connecting all possible paths, then eliminating construct relationships that were insignificant. The final SEM is presented in Fig 3 and details in table 3. T statistics for each path are greater than 1.95 and p values are below 0.05, indicating that each relationship is statistically significant.
Table 4 Cronbach’s Alpha report shows that the t statistic is greater than double the standard deviation, and this indicates the model fits 95% of the data. Table 5 Average Variance Extracted additionally shows strong model fit. Fig 4 Partial Least Squares model was created to determine path coefficients – table 6, and construct validity – table 7.
These final sets of reports explain the model and variance encountered in the model. The weakest relationships surround ANX. Based on this analysis, we know that Clinical Tools strongly influences ANX, however, Clinical Tools has the lowest Cronbach’s Alpha and Adjusted Rho of all reviewed items. ANX also has a less than ideal Cronbach’s Alpha, but other indicators show that it is likely a reliable concept.
3.2Interview Data Analysis
The average interview time was 39.93 minutes. Krippendorff’s CU Alpha was generated at an individual narrative (S1 Appendix B.3) and overall level. Interviews were eliminated until the overall level reached α ≥ .8, resulting in α = 0.82. Code co-occurrence was measured by hypothesis and Sankey diagrams generated (S1 Appendix B.4).
3.3Survey Results
Table 8 includes the outcomes of each hypothesis for the survey. Results are expanded upon in S1 Appendix A.4.
Hypotheses were tested at a 95% significance level.
3.4Interview Results
Table 9 includes the outcomes of each hypothesis for the interview. Interpretation, code co-occurrence tables, and Sankey diagrams supporting results are found in S1 Appendix B.4.
5Discussion
Radiologists have a high intent to use and positive attitudes towards IM based CDSS and the presented interventions overall. We determined that years of experience, and Consumer Tools (IM and CA) were not moderating variables in our model. In any given path, the t statistic was too low and p value too high to consider this in our analysis. These questions are not a part of the UTAUT model, and we found them not to be factors relevant to our efforts. The following UTAUT expected paths were additionally removed, and speculation as to why is included:
5.1Age and Intent to Use
This is a deviation from UTAUT. Potentially, radiologists are technologically saturated users; they perform their work functions using a wide variety of complex technological solutions. Among clinicians, radiologists chose this specialty because of their interest in technology solutions within healthcare. We were unable to measure this result during interviews.
5.2Expected Efforts Influence on Attitude
The survey and interview studies have opposing results for expected efforts influence on attitude. The survey deviated from UTAUT in not finding an association with attitude. However, the interview study showed that decreasing effort is linked to positive attitude and positive intent to use. Common themes on effort/attitude interactions-
- Reducing time to acquire and apply clinical knowledge.
- “…however many seconds it takes for everyone…and it’s different for everyone to figure out how they want to go about finding this information. We all kind of I think most people know where to look IE the ACR guidelines, but…having this thought process kind of forcing us to focus on this dialog box kind of streamlines that whole process. So I think overall, it should enhance the workflow”
- “…these are the things when we’re not given enough information…Some people perseverate on the lack of information more than others. And some people are really dutiful and want to go into the EMR and look, but that could be one to two minutes, and then compound that over an entire shift. Do that a couple times. That could be an hour that you’ve saved if you had this information in a ready format, or in a readily available format, so I think this definitely makes you, from this particular type of interactions definitely makes you more efficient, I believe.”
- Increasing multitasking
- “I think it is a great idea. I think it helps you do multiple things, maybe not just in this cardiac workup. But like for lung nodules when your kind of trying to decide what the appropriate workup is. We always have a caveat that takes seconds to say but it’s still seconds that you have to say it every time. You know, if the patient has high risk for pulmonary malignancy recommend whatever. We know that they’re already high risk for whatever, then I feel like we don’t need to say that. Or if that even auto populates the patient has these risk factors that we would recommend discussing these risk factors.”
- Trusting CDSS as safety nets
- “…we touted on AI is not to replace your diagnostic skills, but eight other things, whether it’s making you more efficient or providing kind of a little safety net, right. Maybe you forgot to mention a follow up or something that really should be a critical result.”
5.3Anxieties influence on Attitude
The survey shows a small negative relationship with attitude. The interview study asked many questions to understand anxiety surrounding this intervention, however, we were unable to strongly correlate with attitude. Overall, anxiety is the least grounded concept throughout the interview.
5.4Expected Performance as the Major Influencer of Attitude and Intent to Use
Overall, expected performance is a major influence on attitude and intent to use. Within the survey results it has significantly more influence than any other factor. However, the interview results show a stronger correlation of expected effort with attitude and intent to use. There is a strong negative relationship between performance and effort present in both phases of the study, another deviation from the UTAUT model. There is potential that radiologists’ system use is derived from performance, maybe measured in clinical outcomes. However, we cannot assume these performance metrics overcome effort needs. Common themes from factors influencing attitude/intent to use-
- Radiologists expect to be interrupted or context switch quickly.
- “Interrupt My normal workflow? Well, I guess it depends on what is normal. This would not interrupt my normal workflow. We’re constantly getting interrupted. It would just be another interruption among a series of normal interruptions.”
- “…this is kind of the thought process that I, this is I go through this checklist. Basically, every time I close a study, we look at the work list again. I’m thinking to myself, looking over my shoulder at the residents and looking at their, their work list, and thinking about [county hospital] over here, looking at how deep my work list is how far so I basically run this checklist mentally, in between each exam.”
- Reducing effort is highly embraced.
- “I love the idea that I’m not having to call someone and that automatically reminds me and I can just either do one click and go one click would be nice to just be like, Yes.”
- “the status quo is quicker or definitely is quicker than then that interface I just saw on the video.”
- “…I’m responsible for all those things that I don’t see on a regular basis…Let me go back to that algorithm figure out what I need to say. This would be a really great tool for me in those cases, because I don’t have to worry as much I think I’m missing a recommendation or something like that. I don’t have to hedge as much. I don’t have to hurry try to get to get to my list. Luckily, we don’t aren’t too inundated so it’s not an issue but I do feel like this will help me to put the appropriate things in with the appropriate recommended”
- “Well, it would shortcut having to call a technologist and initiate a conversation about what the patient was, what the study was, the post processing that you needed, done. So if the radbot could predict that you might need it. And could figure out what you needed quicker and would negate having that phone call that would be positive.”
- Radiologists will trade effort for performance.
- “So it slows you down slightly, but in the long run of collections and all that stuff. Yeah, I think it would [improve performance], because you’re making sure you get reimbursed and get the correct RVU amounts for the right study.”
- “… the amount of time it takes to look those things up, and it’s not super frequent, but not super infrequent…either it’s taking you more time to figure it out or you also end up with more a variation amongst different radiologists for the recommendation. So you’re not only might save time, but you might also decrease the heterogeneity of the recommendations. And probably more, you’d be probably more likely to actually be following the guidelines since you’d be prompted to to adhere to them.”
- “…I think maybe it slows you slightly on the front end, but on the back end, it helps you, and it helps clinicians too”
- “I think it would decrease my productivity very minimally. But for a good cause.”
5.5Conclusion
“No, I think again, this is it’s all these videos have demonstrated processes which are mostly done mentally by all radiologists. And again, it’s not always easy to kind of put these things on a screen or you know, because your kind of you’re, you’re juggling a couple different priorities at the same time. So, I think this is kind of taking an existing thing and it’s making a more organized and streamlined fashion.”
Radiologist’s interactions with decision support tools, or at least this intervention, differs from the standard user software interaction model. The positive relationship from performance to effort is the most major deviation, allowing increasing effort if the outcomes are desirable enough. This relationship is supported by both the survey and interview studies. Further, because performance and effort make up most of attitude and intent to use, there are a lot of opportunities for CDSS to provide novel workflow changes that increase patient outcomes. CDSS should be designed to streamline activities, and we see particular interest in tools to enable clinical knowledge gathering and context switching.
“Yeah, assuming that you would have gone to the EMR it was important enough to go there. And if you went there to take more time, okay. But at the same point, you know, it might change the threshold at which you would ask a question, right? It’s like, it’d be nice if we knew this and it’s easy just to query it. But you otherwise might not go to the EMR.”
Anxiety is another large deviation from the standard user model. In both parts of the study anxiety had the weakest relationships and was often secondary to the excitement of new clinical solutions. The most common source of anxiety surrounds the maintenance of CDSS “I guess the part that causes me to pause is who’s going to be mining for new updates? And how can we be sure that we’re staying current on recommendations? You know, how is that? Who’s going to handle that part of it?”
Radiologists deviate from the standard clinician with regards to the 10 commandments of CDSS. Commandments 2, 3, 7, 10, and 1 – anticipate needs, fit into user workflow, simple interventions, knowledge system maintenance, and speed – are all highlighted within radiology specific guidance, and we do find these present for radiologists in our study. However, the relationship between performance and effort highlights that radiologist CDSS doesn’t need to always hit every commandment. Radiologists expect workflow modification, they routinely use complex interventions, and they are not overwhelmed by CDSS information gathering. As we design for the future radiologist, we can trade effort in these commandments for increasing positive outcomes.
Data Availability
All relevant survey data are within the manuscript and its Supporting Information files. Interview data is withheld due to IRB restrictions.
Supporting Information Captions
S1 Appendix. Detailed information on the study including the full survey, expanded hypothesis results, Semi-structured interview guide, code co-occurrence tables, full data analysis/findings, and additional diagrams.
S2 Survey Data. Raw quantitative survey data in Excel format.