Characterizing Trainee Workload Using Paired Daily Surveys and EHR Use Data: A Mixed-Methods Pilot Study
1second-year Fellow, Division of Pulmonary, Critical Care and Sleep Medicine, Department of Internal Medicine, College of Medicine; The Center for the Advancement of Team Science, Analytics, and Systems Thinking in Health Services and Implementation Science Research, The Ohio State University Wexner Medical Center
2Senior Instructor, University of Colorado, Department of Medicine, Division of Hospital Medicine
3Senior instructor, Veterans Health Administration, Eastern Colorado Health Care System, Section of Hospital Medicine
4third-year Internal Medicine Resident, University of Colorado, Department of Medicine, Internal Medicine Residency Program
5Data Analytics Program Manager, University of Colorado, Department of Medicine, Division of Hospital Medicine
6Associate Professor and Program Director of the Hospitalist Training Track, University of Colorado, Department of Medicine, Division of Hospital Medicine, Internal Medicine Residency Program
7Associate Professor, University of Colorado, Department of Medicine, Division of Hospital Medicine
8Full Professor and Division of Hospital Medicine Head, University of Colorado, Department of Medicine, Division of Hospital Medicine
*Correspondence should be addressed to: Karan Rai MD, MHA, 241 W 11th Ave, Suite 5082, Columbus, OH 43201, Karan.Rai@osumc.edu, Phone: (614) 338-9045Abstract
Purpose
High clinical workload is associated with worse patient and hospital outcomes and is a well-established driver of clinician burnout. Trainees may be particularly exposed, shouldering both clinical and educational responsibilities. Evidence-based work design offers a data-driven approach to healthcare work but relies on robust workload measurements. Trainee workload remains poorly characterized, as commonly used metrics (e.g., duty hours, patient census) overlook cognitive and contextual dimensions. This pilot evaluated the feasibility of combining survey-based and electronic health record (EHR) data to characterize internal medicine (IM) trainees’ workload.
Methods
A pilot study was conducted including IM and Medicine-Pediatrics residents (postgraduate years 1-4) between March 31 and June 22, 2025. Participants completed daily surveys during a seven-day inpatient schedule assessing workload and work experience domains, including environment, professional fulfillment, psychological safety, autonomy, and rounding experience, using validated instruments where available. Concurrently, EHR data captured chart review, documentation, orders, and secure messaging activity. Associations between survey and EHR data were assessed.
Results
Among 37 eligible residents, 28 (76%) participated in the pilot capturing 166 shifts. Trainees spent 4.4 ± 1.6 (mean ± SD) minutes completing daily surveys and 8.6 ± 2.3 minutes completing the final survey. Trainees reported working 11.6 ± 1.0 hours/day and a median census of 9.0 (IQR 6.0-11.0). NASA-TLX score was 50.8 ± 12.6. Positive shift ratings were associated with lower NASA-TLX scores and perceived rounding length. First-to-last EHR login duration was 15 ± 2 hours/day, and EHR data showed 204 ± 46 active minutes/day. Login duration correlated with self-reported hours (r=0.43, p<0.0001), and notes signed correlated with self-reported team (r=0.19, p=0.013) and personal census (r=0.34, p<0.0001).
Conclusions
Integrating survey-based and EHR-derived workload measures provides multidimensional insight into trainee work. This novel approach supports scalable measurement and evidence-based work design interventions to improve trainee well-being, education, and clinical efficiency.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This study was supported by internal organizational funds for participant incentives.
Introduction
High clinical workload is associated with worse clinical and hospital outcomes1,2 and contributes to clinician burnout.3 Trainees, who work up to an average of 80 hours per week, are at particular risk, with nearly half of physicians-in-training reporting burnout,4 underscoring longstanding concern about the intensity and structure of trainee work. Clinician burnout is linked to personal morbidity, such as poor health and increased suicide risk,5 as well as downstream effects on patient care such as higher rates of medical errors, lower patient satisfaction, and billions of dollars in societal cost through lost productivity.3,6 In recent years, trainees have increasingly turned to organizing efforts (i.e., unionization)7 to advocate for improved work conditions and workplace protections.8,9 While these efforts have demonstrated promising early results in securing benefits such as vacation and stipends, collective action alone does not ensure improvement in the design of day-to-day clinical work, educational quality, or lived work experience.10,11
These limitations highlight the need for complementary, system-level strategies that directly address how clinical work is structured. Evidence-based work design has emerged as a promising framework that systematically applies data and human factors principles to optimize staffing, task allocation, and workflows to support patient safety, operational performance, and clinician wellbeing.12–14 Rather than relying solely on advocacy or regulation, this framework seeks to redesign work itself and offers a potential path to improve both operational efficiency and trainee experience.15 Central to this approach is the availability of valid, actionable measures of workload and work environments. Traditional metrics, such as duty hours and patient census, remain widely used in graduate medical education, but they are increasingly believed to be insufficient alone, missing critical elements of patient acuity, cognitive and emotional burden, and context that shape how workload is experienced.12,13,16,17 While duty hour restrictions have historically been the major focus of several overhauls of trainee clinical work regulations18, their effectiveness in improving patient and resident outcomes has been mixed and heterogenous across training programs.19 Moreover, the clinical environment has shifted considerably since the early 2000s. Contemporary trainee work is increasingly dominated by electronic health record (EHR) activity20 and digital communication.21,22 These tools, while essential, can increase cognitive load and fragmentation of work23,24 in ways not reflected in duty hours or simple census count. Concurrently, qualitative and contextual features, such as psychological safety, autonomy, task-switching, rounding structure, clinical acuity, and workspace design, modulate workload but are rarely captured in operational data.25 Together, these changes highlight the need for a multidimensional understanding of optimal trainee workload and educational experience, extending beyond single, coarse indicators.
Taking advantage of this dynamic clinical environment, new approaches to understanding the trainee work environment are increasingly available to organizational leaders – including EHR use measures (i.e., digital clickstream data collected during regular work)20,26–29 and advanced predictive analytics.12,30,31 These emerging tools when paired with brief intermittent surveys offer a pragmatic strategy to link how work is experienced with how it is performed, enabling more precise evaluation of workload, workflow, and educational structures. This approach offers opportunities to provide early signals of strain at the individual or team level, supporting timely intervention.12 As such, we conducted a pilot feasibility study to evaluate whether pairing surveys with EHR use data could generate multidimensional, trainee-relevant workload measures that are pragmatic to collect and analytically informative in a real-world graduate medical education setting. This work aimed to lay the groundwork for shaping evidence-based work redesigns that optimize the training experience by maximizing educational value while reducing unnecessary cognitive and administrative burden.
Methods
Pilot Design
We performed a mixed-methods pilot to characterize and quantify daily workload and work environment among internal medicine (IM) trainees. The pilot combined daily and weekly survey data and EHR use measures. The pilot was deemed exempt by the Colorado Multiple Institutional Review Board (COMIRB #24-0303). We adhered to the STROBE guidelines for study reporting.32
Setting, participants, and study size
The pilot was conducted between March 31, 2025 and June 22, 2025, at a quaternary care academic hospital with over 700 beds and included IM and medicine-pediatrics residents. Residents were recruited across post-graduate years (PGY) 1-4 and invited to participate by a residency-wide email from a project team member in a non-supervisory role. Eligible residents were those working inpatient services on day shifts that spanned from 6 a.m. to 5 p.m. for seven-day stretches with one day off. Preliminary PGY-1 residents were not eligible for participation. Verbal consent to complete surveys and collect linked anonymized EHR use data was obtained prior to participation. Residents were able to end their participation at any time. Participants received $25 gift cards as compensation for participation in the study after completion of the final survey.
Survey development
We used the Chokshi and Mann process model33 for user-centered digital development to develop and pilot a modified version of a previously described workload survey application27 that includes patient census, task load, and health systems culture and validated measurement tools including the National Aeronautics and Space Administration Task Load Index (NASA-TLX)34 and the teamwork composite of the Agency of Healthcare Research and Quality (AHRQ) Hospital Survey on Patient Safety Culture (Figure 1).35 In phase one of the discover and define stage, we identified unique measures of workload that captured a multidimensional view of trainee work using focus groups and a modified Delphi panel.25 In phase two, using the previous survey application27 as a prototype for this work, we used a mixed methods approach to determine which measures to include for trainees. Specific measures felt to be less relevant to trainees were excluded, and additional measures were incorporated to address previously identified trainee-specific domains of workload, including autonomy, psychological safety, and rounding.25 For these domains, validated measurement tools were used when available, including the National Institute for Occupational Safety and Health (NIOSH) Worker Well-Being Questionnaire (WellBQ) on Physical Work Environment Satisfaction,36 Professional Well-being Academic Consortium (PWAC) Stanford Professional Fulfillment Index (PFI),37 Edmondson Psychological Safety scale,38 Work-life climate,39 and Moral Distress Thermometer.40 Where validated instruments were not available, survey questions were developed and, in phase one of the develop and deliver stage, user-tested with trainees for clarity and consistency. We additionally included a free text question to allow participants to share any other shift satisfiers, dissatisfiers, and notable events. The survey was user-tested among six trainees between January 6, 2025 and January 22, 2025 using structured interviews before finalization (Supplementary Material 1).
In phase two of the develop and deliver stage, we piloted the survey (Supplementary Material 2). Surveys were distributed via REDCap41 by email or text message, based on participant preference, at the end of each workday for a total of five daily surveys and one final end-of-week survey. For daily surveys, one reminder was sent after 1.5 hours if no response was received, and two reminders were sent for the end-of-week survey, at 1.5 hours and 24 hours if no response was received.
EHR data
We extracted aggregated vendor-derived EHR use data (i.e., data collected from regular EHR use during clinical work), including total time in the EHR, secure messaging volume (i.e., number of messages sent and received), and time spent performing various EHR activities, including chart review, documentation, in-basket activities, orders, and problem list management during the pilot week for each participant using Epic Signal.28,29,42,43 This data is aggregated at the week-level by participant. Additional more granular EHR use data, including notes signed and first and last login times, was also extracted using pre-existing datasets from the EHR vendor.
Analysis
Quantitative analyses were performed using SAS Enterprise Guide 9.4 (SAS Institute Inc.). Given the pilot nature of the study and the modest sample size, analyses were limited to unadjusted comparisons and correlations to characterize feasibility and directional associations rather than causal inference. Associations were organized into three domains: (1) subjective shift rating and discrete workload features; (2) correlations between objective workload (e.g., census, NASA-TLX) and subjective domains (environment, professional fulfillment, psychological safety, autonomy, and rounding experience); and (3) concordance between subjective workload and EHR-derived activity metrics (i.e., EHR-derived and granular EHR use data). Associations between the binary shift-rating measure and discrete variables, including patients on treatment team, patients personally responsible for, number of hospital units covered, NASA-TLX scores, and rounding scores, were evaluated using Students’ t-tests.
We then examined correlations among shift-level workload measures. Pearson correlation coefficients were used to assess associations between NASA-TLX scores and, (1) patients on treatment team, (2) patients personally responsible for, and (3) rounding scores.
We further evaluated correlations between each survey-based domain, including WellBQ score, professional fulfillment, work exhaustion, disengagement, psychological safety, work life climate, and moral distress, and shift-level: (1) patients on treatment team, (2) patients personally responsible for, and (3) NASA-TLX.
Finally, we assessed concordance between subjective workload reports and EHR-derived metrics by correlating, (1) self-reported shift hours with total time in the EHR, and (2) reported team and personal patient counts with EHR-derived patient counts. Given the number of comparisons, statistical significance was defined as p<0.002.
To capture subjective components of workload experience not explicitly queried through the survey questionnaire, we completed a hybrid inductive-deductive content analysis of additional free-text comments provided by participants. Four study team members (CF, JC, KR, NB) jointly reviewed comments to develop a codebook (Supplemental Table S2) with working definitions of overarching domains, then individually and separately coded each comment with discrepancies adjudicated as a group to reach consensus.
Results
Participants
A total of 37 trainees were eligible and invited for participation in the project, and 28 (76%) trainees participated in the pilot. Trainees included nine (32%) PGY-1, nine (32%) PGY-2, eight (29%) PGY-3, and one (4%) PGY-4. The participants included 14 (50%) trainees in the categorical track, six (21%) in the hospitalist track, five (18%) in the primary care track, and one (4%) in the physician scientist track. Post-graduate year data was missing for one participant, and training track data were missing for two participants. Participant demographics can be found in Table 1.
Survey and EHR use data results
Survey responses were completed for all six shifts in the pilot weeks for 26 (93%) participants, with one (4%) participant submitting a survey for four out of the six shifts, and one (4%) participant who completed surveys for five shifts and a partial survey for one shift. A total of 166 shifts were captured. Participants spent 4.4 ± 1.6 (mean ± standard deviation [SD]) minutes completing daily surveys and 8.6 ± 2.3 minutes completing the weekly survey (Supplemental Table S1).
Full survey results can be found in Table 2. Trainees worked a mean of 11.6 ± 1.0 hours per shift and cared for a median of 9 patients (interquartile range [IQR] 6–11) across a median of 3.2 ± 1.1 hospital units. Participants commonly supervised other learners (mean 1.7 ± 1.0 per shift) and most frequently worked with attending physicians (93%), medical students (67%), interns (35%), and senior residents (31%). Hallway rounds were the predominant rounding structure (70%), with bedside rounds reported in 12% of shifts. Mean ratings of rounds on a 1-10 scale reflected moderate educational value (6.1 ± 1.7), length (6.3 ± 1.5), and efficiency (5.7 ± 1.9). Weekly survey results showed low professional fulfillment on average (mean 2.3 ± 0.6), with 11% of participants meeting criteria for fulfillment (score ≥3.0). In contrast, 75% met criteria for work exhaustion and 57% for overall burnout (scores ≥1.33). Psychological safety scores on a 1-7 scale were high (mean 5.4 ± 0.8), and perceived autonomy during shifts closely matched desired autonomy levels. Qualitative content analysis of free-response questions indicated team structure, care coordination, and disruptions/task switching as common contributors to workload (Supplemental Table S2). The remaining free responses can be found in Supplemental Material 3.
We found that the mean NASA-TLX score was lower among shifts rated with “happy face” emojis compared to those with “sad face” emojis (46 ± 11 vs. 62 ± 9; p<0.0001) (Supplemental Table S3, Supplemental Figure S1). Similarly, mean rounding length scores were lower among shifts rated with “happy face” emojis compared to those with “sad face” emojis (6 ± 1 vs. 7 ± 2, p<0.0001) (Supplemental Table S4, Supplemental Figure S2). Mean rounding length scores were correlated with mean NASA-TLX scores (0.35 [95% CI: 0.21, 0.48], p<0.0001) (Supplemental Figure S3). No significant associations were found when comparing number of patients on the treatment team/patients personally responsible with shift rating, NASA-TLX, WellBQ score, professional fulfillment, work exhaustion, disengagement, psychological safety, work life climate, or moral distress scores. Similarly, NASA-TLX was not associated with any of these survey-based domains.
In regard to EHR use data, time from initial login to last login among trainees was 15 ± 2 hours per day (Table 3). Trainees spent an average of 204 ± 46 minutes per day in the EHR, primarily spent on documentation, chart review, and order management. Trainees signed 6 ± 3 notes per day. We found no associations between reported shift hours, and total time spent in the EHR but did find a correlation between reported shift hours and duration of shift measured as time from initial to last login into the EHR (r=0.43 [95% CI: 0.30, 0.55], p<0.0001). We also found a correlation between mean number of notes signed per day and mean reported team census (r=0.19 [95% CI: 0.04, 0.33], p=0.0130) and personal census (r=0.34 [95% CI: 0.20, 0.47], p<0.0001). Similarly, we found no associations between EHR-based identification of number of patients cared for, which was defined as average count of patients per day where the provider has co-signed, attested, or contributed characters to a note, and survey reports of number of patients on treatment team or personally responsible for.
Discussion
In this pilot feasibility study, we sought to determine whether regular, brief trainee surveys could be paired with EHR use data to capture multidimensional aspects of trainee workload in a pragmatic manner. We found high engagement with the survey instrument, minimal time burden, and positive feedback regarding continued participation, supporting feasibility in routine clinical practice. Subjective measures, including shift ratings and NASA-TLX scores, demonstrated consistent relationships with rounding duration, suggesting sensitivity to variations in cognitive load. EHR use data provided complementary, objective measures of task-level work that contextualized these insights. Together, these findings indicate that paired subjective and objective workload measurement is feasible in graduate medical education and yields information not captured by traditional metrics alone. These complementary data streams have the potential to lead to timely and meaningful change, in contrast to the prolonged runway to await policy or regulatory changes that have traditionally informed the structure of trainee work.
Although this pilot was not powered to test causal associations, several descriptive patterns aligned with existing literature and warrant consideration. Rounding emerged as a prominent contributor to perceived workload, with subjectively longer rounding duration associated with poorer shift ratings and higher cognitive load. Trainees frequently described prolonged and inefficient rounding practices with limited educational value, consistent with prior studies demonstrating resident rounding has substantially lengthened in recent years.44–46 Notably, in comparison to recent evidence among hospitalists,27 cognitive load among trainees was generally higher. These findings highlight rounding structure and cognitive load as potentially high-yield targets for future work design interventions.
EHR use data further demonstrated substantial time devoted to documentation, orders, chart review, and secure messaging volume, activities that consume cognitive bandwidth but are largely invisible to traditional conventional workload metrics. Both secure messaging and disruptions/task switching, which are known to be associated, emerged as workload domains in qualitative content analysis of free-text survey responses, and align with prior work suggesting that secure messaging volume is associated with EHR wrong-patient ordering errors47 and nonurgent and non-actionable secure messages are intervenable targets for improvement.22 Moreover, residents spent significantly more daily time in the EHR compared to hospitalists27, underscoring both the administrative burden borne by trainees and the importance of objective measures that capture digital work. Total EHR time in our study was significantly less than survey-based reports of total time spent at work, similar to previous findings.27 EHR-derived measures capture observable EHR interactions but may under-represent cognitive work performed outside of the system and can vary by vendor algorithms for “time-outs” due to pauses in mouse/keystroke data.26 This may explain the identified correlation between reported hours and time in EHR measured between first and last log-in, which avoids “time-out” issues, but may overestimate continuous work if sessions are interrupted (e.g. off-site logins). Together, these findings suggest that not all EHR-derived metrics reflect lived workload equally. Measures that capture the span of work (i.e., first-to-last login) aligned more closely with trainee-reported hours than aggregated active EHR time, highlighting the importance of selecting workload measures that map to how work is experienced.
The operational implications of this feasibility study suggest that integrated dashboards that combine survey-based signals with EHR-derived indicators could support more responsive approaches to monitoring trainee workload. When paired with brief, intermittent surveys that capture contextual features of work, these complementary data streams may enable earlier identification of emerging strain for teams and for specific rotations, as well as the empirical derivation of optimal work and educational design, well before downstream outcomes such as burnout or attrition occur. Beyond surveillance, these data can serve as baseline metrics to evaluate targeted interventions, such as restructuring rounds, optimizing messaging workflows, or streamlining documentation processes. Importantly, implementation of monitoring systems should be coupled with clear, non-punitive escalation pathways to ensure that identified workload concerns prompt timely support rather than unintended consequences.
Limitations
This study has several limitations to consider. First, our study population was limited to a single residency program with a small sample size and may reflect institution-specific findings rather than generalizable trends. Our study was conducted among IM and medicine-pediatrics residents assigned to inpatient ward primary services during the daytime and may not represent non-IM contexts nor be generalizable to outpatient settings or nocturnal coverage. Although survey data may suffer from response bias, our study sought to additionally evaluate EHR data as a measure of workload. Finally, given our study was conducted by an opt-in survey, it may suffer from participation bias. Our survey findings should be interpreted cautiously and considered hypothesis-generating and merit more rigorous testing through larger, multicenter studies and implementation trials in varying contexts.
Conclusion
Combining brief surveys with granular EHR audit-log data provides a feasible, practical, and scalable approach to capturing multidimensional aspects of trainee workload in real-world clinical settings. This paired measurement strategy extends beyond traditional indicators by integrating subjective experience with objective task-level work, offering actionable insight into how workload is structured and experienced. In this pilot, these data highlighted potentially modifiable contributors to workload, including cognitive load, rounding structure, and EHR-related burden, and demonstrated acceptability and analytic utility in graduate medical education. Embedding such measurement into routine practice may provide a foundation for evidence-based work design efforts aimed at improving the clinical learning environment while reducing unnecessary cognitive and administrative strain. Future studies should consider leveraging these tools to power trials of targeted work redesign and to evaluate downstream effects on education, clinician well-being, and patient outcomes.
Supporting information
Acknowledgements and Conflicts of Interest
Drs. Burden and Keniston report funding from the Agency for Healthcare Research and Quality, the National Institute for Occupational Health and Safety, University of Colorado Innovations digiSPARK award, and the American Medical Association not related to this work.
Drs. Burden and Keniston contributed to the development of GrittyWork, a digital workforce application, and a registered trademark of the University of Colorado. Dr. Burden also reports funding from Med-IQ and CDC NIOSH Center for Health, Work, and Environment (CHWE), a NIOSH Center of Excellence for Total Worker Health not related to this work. Dr. Burden reports funding from CU Thrive Precision Innovation Award.
The authors utilized the ChatGPT language model developed by OpenAI and Microsoft Copilot for editing original author content to improve readability. All information and materials in the manuscript are original.
Dr. Karan Rai had access to all the data in the project and takes responsibility for the integrity of the data and the accuracy of the data analysis.
The views expressed in this article are those of the authors and do not necessarily represent the views, policies, or positions of any affiliated institutions or employers. Similarly, concepts presented in the introduction and discussion are intended to illustrate general principles and do not describe or reflect any specific health system or organization.
Ethical approval
This study was deemed exempt by the Colorado Multiple Institutional Review Board (COMIRB#24-0303)
Previous presentations
none
Data availability
The data underlying this article will be shared on reasonable request to the corresponding author.