Regression-based modeling of pairwise genomic linkage data identifies risk factors for healthcare-associated infection transmission: Application to carbapenem-resistant Klebsiella pneumoniae transmission in a long-term care facility
1University of Michigan School of Public Health, Department of Epidemiology
2University of Michigan School of Public Health, Center for Social Epidemiology and Population Health
3University of Michigan, Department of Microbiology and Immunology
4Rush University Medical Center, Division of Infectious Diseases
5University of Michigan, Department of Medicine, Division of Infectious Diseases
#Corresponding Author Hannah Steinberg hsteinb@umich.eduAbstract
Background
Pathogen whole genome sequencing (WGS) has significant potential for improving healthcare-associated infection (HAI) outcomes. However, methods for integrating WGS with epidemiologic data to quantify risks for pathogen spread remain underdeveloped.
Methods
To identify analytic strategies for conducting WGS-based HAI surveillance in high-burden settings, we modeled patient- and facility-level transmission risks of carbapenem-resistant Klebsiella pneumoniae (CRKP) in a long-term acute care hospital (LTACH). Using rectal surveillance data collected over one year, we fit three pairwise regression models with three different metrics of genomic relatedness for pairs of case isolates, a proxy for transmission linkage: 1) single-nucleotide variant genomic distance, 2) closest genomic donor, 3) common genomic cluster. To assess the performance of these approaches under real-world conditions defined by passive surveillance, we conducted a sensitivity study including only cases detected by admission surveillance or clinical symptoms.
Results
Genomic relatedness between pairs of isolates was associated with room sharing in two of the three models and overlapping stays on a high-acuity unit in all models, echoing previous findings from LTACH settings. In our sensitivity analysis, qualitative findings were robust to the exclusion of cases that would not have been identified with a passive surveillance strategy, however uncertainty in all estimates also increased markedly.
Conclusions
Taken together, our results demonstrate that pairwise regression models combining relevant genomic and epidemiologic data are useful tools for identifying HAI transmission risks.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This research was supported by the National Institutes of Health [R01 AI148259]. Authors have no competing interests to report.
Introduction
Despite intensive research and scrutiny, healthcare-associated infections (HAIs) remain among the most frequent adverse events occurring in health facilities throughout the United States and the world. Improvements in broadly-effective infection prevention interventions such as hand hygiene and environmental cleaning, and targeted interventions such as pathogen decolonization, have been attributed with recent reductions in HAIs.1 Still, it is estimated that on any given day 1 in 31 hospital patients in the US has at least one HAI.1 HAIs are among the top 10 causes of death in the US and are associated with billions of dollars in excess healthcare costs.2 Antibiotic resistance is common in healthcare pathogens, and can make these infections harder to treat.
Colonization typically precedes infection, and many HAI prevention strategies work by interrupting transmission of colonizing pathogens. However, a better understanding of drivers of transmission and of which patients are more likely to transmit or acquire colonization could help in developing more effective interventions to reduce HAIs.
With the increased availability and falling cost of whole genome sequencing (WGS), there has been increased interest in the use of WGS to understand transmission pathways in healthcare settings.3 However, many studies of transmission in healthcare settings are descriptive in nature, e.g. identifying shared exposures among individuals with genomic linkage, but not quantitatively evaluating whether exposures are shared more than would be expected by chance. Thus, understanding risk factors for transmission and identifying putative targets for improved infection prevention will require more rigorous methods that integrate genomic and epidemiologic data to quantitatively identify transmission risk factors. While this has been done to some degree in outbreak settings,4,5 there is still a need for methods applicable in high-prevalence endemic settings, where the constant importation of resistant organisms makes delineating transmission links challenging, even with genomic data.
In addition to the lack of standard frameworks for integrating genomic with epidemiologic data, additional barriers to current methods (e.g. SNV-based regression models,6 machine learning algorithms,7 and probabilistic transmission models8) include the requirement of single nucleotide variant (SNV) cutoffs to infer transmission,6,7 needing data on uninfected controls,7 and models that are complex7 and/or require many assumptions about the transmission system hindering their generalizability.8 Additionally, models that do not account for the disproportionate effect of super-spreaders on model outcomes may overestimate confidence in risk-factor estimates.9 Recent work has shown that pairwise models that utilize individual, pairwise, and contextual data to describe the genetic relatedness of pathogen isolates in an endemic setting9 can identify drivers of transmission with fewer assumptions and computational needs than some previous studies and do not require data on non-cases or SNV cutoffs. This method involves a regression model in which the outcome is a measure of genetic similarity between a pair of isolates, and covariates are assessed for their influence on genetic similarity, which can be considered in many cases a proxy for transmission.
In this analysis, we evaluate the use of pairwise models to describe how carbapenem-resistant Klebsiella pneumoniae (CRKP), an important healthcare-associated pathogen, is transmitted in a long-term acute care hospital (LTACH). CRKP’s high prevalence in LTACHs make delineating transmission pathways complex with traditional epidemiologic methods. Although we have limited epidemiologic information in this dataset, we are able to identify known transmission risk factors with our models and hope this study can serve as an example of how to conduct this type of analysis in settings with richer epidemiologic data.
In addition to evaluating different approaches for incorporating genomic relatedness into pairwise statistical models, we also assess the sensitivity of these models to case capture. Sampling strategy may be important in understanding CRKP transmission as asymptomatic colonization with CRKP is a common precursor to invasive infection10 and potentially important in intra-facility spread,11 yet most facilities do not screen for asymptomatic colonization. To understand the impact of the sampling scheme on identification of transmission risk factors we evaluated models with a more passive surveillance strategy for detection of carriers.
Methods
Study population
CRKP surveillance samples were collected via rectal swab on admission and every two weeks from June 2012-June 2013 for all patients (n = 937 unique patients) in an LTACH in Chicago, Illinois (USA).12 Average daily patient census was 98 (SD: 7.4), and the median length of stay was 27 days (IQR: 17, 44). The mean age of patients was 60.5 years (SD: 15.8), and 43.1% of patient-days were for ventilated patients. This study was approved by the institutional review boards at Rush University Medical Center (Chicago, IL, USA) and the University of Michigan (Ann Arbor, MI, USA). Informed consent was waived.
Surveillance samples were cultured and unique colony morphologies were identified to species. Ertapenem disks were used to screen isolates for CRKP and a confirmatory PCR was conducted to detect blaKPC, the sole carbapenemase gene associated with CRE in the region during the study period. To capture contact patterns, each patient’s daily room and floor locations were recorded. Antibiotic usage over time was also recorded for each patient. Whole genome sequences were obtained for all positive isolates and recombination-filtered core genome alignments were produced for each sequence type.13 For this analysis, only sequence type 258 (ST258) isolates, the most common sequence type in the LTACH (70% of cases), were used.
Evaluation of the effect of serial sampling on risk-factor estimates
As most facilities do not have robust serial sampling strategies like our study facility implemented, we re-ran all models including only patients who tested positive on admission or had a positive CRKP test as part of clinical evaluation outside of the colonization study to examine the influence of serial sampling on the ability to make inferences on transmission dynamics.
Results
Prevalence of colonization and infection with CRKP in a single LTACH
In total, 255 individuals were colonized with at least one strain of CRKP during the study period (with an average prevalence of 32% throughout the year),13 180 of whom (70% of those colonized) were colonized with strain ST258. Of the 180 patients colonized with CRKP ST258, 87 (48%) were positive on admission, 72 (40%) had CRKP detected via clinical testing, and 54 (30%) were detected after admission during serial sampling and never had a clinical CRKP isolate. There were 37 genomic clusters of ST258 (2-16 isolates per cluster) previously identified in this study population with a threshold-free cluster detection approach that clustered each CRKP isolate acquired at the LTACH to the importation isolate with which it shared the greatest number of variants.13
Pairwise models suggest room sharing, residing on the floor that housed the high acuity unit, and shorter time lags between colonization detection are associated with genetic relatedness
In all pairwise models shared time on either Floor A or Floor D was associated with increased pairwise genomic relatedness, with Floor D (which contained the facility’s high acuity unit) having the larger effect on pairwise genomic relatedness, suggesting there may be more intra-floor risk on floors where patients require more intensive care (Table 2). In Model 1 (pairwise distance), both patients residing on Floor D was associated with SNV distances 44% (95% CI: 38% - 49%) closer. In Model 2 (closest donor), sharing time on floor D was associated with 7 (95% CI: 4-13) times greater odds of being the closest potential donor. In Model 3 (same transmission cluster), sharing time on floor D was associated with 10 (95% CI: 6-18) times greater odds of being in the same cluster. When collapsing the effect of floor sharing into a single covariate, sharing a floor during their exposure period was a significant predictor of genetic relatedness of CRKP isolate pairs in all models (Supplemental Table 1).
Other factors, such as sharing a room during a pairs’ exposure period, a shorter lag between positive cultures, and the exposure period occurring in the fourth quarter of the study period were positively associated with isolate similarity in each model, although certainty varied. Model 3 identified all three of these factors as significantly related to cluster comembership, Model 1 captured two of these factors (shared room and time period) as significantly associated with SNV distances, and Model 2 failed to show a statistically significant effect of any of these factors on the likelihood of being the closest potential donor (Table 2). Antibiotic exposure of the donor during the pairwise exposure period was not a meaningful predictor of CRKP relatedness in any of the models.
Case and admission-positive only models underestimate key risk factors
Although our study utilized serial surveillance for asymptomatic carriage, most facilities have only clinical culture isolates available, with some also testing for CRKP colonization on admission. Of the 180 patients colonized with ST258 CRKP in our study, 30% would never have been identified if only clinical and admission screening cultures were conducted, and 61% of closest potential transmission pairs (defined as in Model 2) would have been missed. When we excluded these patients from our analyses, culture date difference remained a significant predictor of genomic similarity, but room and floor sharing was not significant in any of the three models (although the qualitative direction of coefficients remained unchanged) (Supplemental Table 2).
Discussion
Using regression models of pairwise genomic relatedness, we were able to identify risk factors for CRKP transmission in an endemic LTACH setting. Although certainty varied, regardless of the metric of genomic relatedness employed, sharing a room or having an overlapping stay on the same floor, especially the floor that included the high acuity unit, predicted shorter genomic distances and a higher likelihood of membership in the same cluster. This is consistent with studies showing increased risk of CRKP infection among those with more intense care needs (e.g. fecal incontinence, mechanical ventilation) and those exposed to infected roommates.16–18 This work suggests that decolonization and other infection prevention efforts should be focused on close within-facility contacts of CRKP patients, with particular attention to high-acuity patients who have higher illness severity and are likely to have more medical interventions and direct hands-on contact with staff.
The availability of WGS data from colonization isolates gave us the ability to evaluate how individual, dyadic, and contextual factors predicted the genetic similarity of CRKP isolates, a proxy for transmission risk. Using these WGS data, we quantified the relatedness of CRKP isolate pairs in three ways: SNV distance, closest potential donor, and cluster co-membership. Each of these measures of relatedness yielded risk-factor estimates which were consistent with each other as well as existing literature on CRKP transmission. This suggests that the measure of relatedness used in pairwise models may be flexible. It may be important, however, to consider the sampling strategy and transmission dynamics of the pathogen of interest when selecting a pairwise metric. For instance, if serial sampling was not conducted and it is unlikely that direct transmission pairs have been identified, a closest donor approach may not be sensible. Or, if the pathogen of interest has a well-established SNV cutoff to determine cluster co-membership, using the criteria of a pair meeting that cutoff may be used as the model outcome. In our study population, it appears that a threshold-free cluster comembership model identifies transmission risk factors with the most certainty compared to a closest donor or SNV distance models.
Sensitivity analyses revealed that excluding data from serial surveillance isolates reduced our ability to identify the risk factors highlighted using the full dataset. This likely reflected the decreased number of cases overall in the reduced dataset as well as missed direct transmission links, highlighting the importance of serial culture surveillance as a tool for identifying transmission risk factors. When patients who did not have a clinical CRKP isolate during our study period and were negative on admission were excluded from analyses, 61% of probable transmission pairs were missed and our ability to detect an effect of room and floor sharing on genetic relatedness was weakened.
In addition to room and floor sharing, our models revealed that pairs are less likely to be closely genetically related if the time between collection of positive samples is longer. This suggests that a susceptible patient is more likely to get CRKP from someone who has more recently acquired CRKP than someone who has been colonized for a longer time. It could also indicate intra-host evolution between the time of acquisition and transmission. Two of our three models also suggested that individuals colonized during the last quarter of our study period were more closely related to their potential donors than those infected in the first period. This could be an artifact of model setup (as the study period progressed, the number of potential donors increased), the result of the introduction of a new strain into the facility with different transmission patterns, or more intra-facility transmission in this period. However, incidence of CRKP within the facility appeared to decrease throughout the study period,12 and thus this result may indicate the onward transmission of fewer strains within the facility resulting in those infected appearing to be more closely related to each other than if many strains were circulating due to a bottleneck effect. Lastly, although antibiotic exposure has been associated with CRKP acquisition risk in healthcare facilities,16,17 our results do not provide evidence that antibiotic exposure is associated with a change in the number of transmissions generated by a colonized or infected LTACH resident. However, the very high prevalence of antibiotic use in our patient population may have hindered detection of their impact on transmission.
Due to data availability and methodological constraints, our study can be considered to have the following limitations. First, we did not have access to information on patient-level procedures, devices, and healthcare worker exposures, which all may play a role in transmission and could help determine specific mechanisms increasing transmission risks, specifically on Floor D. However, we were able to identify known risk factors of CRKP transmission in an LTACH and hope this study will serve as a template for facilities that may have more detailed data available. Additionally, given that only 58% of closest donor pairs were in the same genomic cluster, it is likely we are missing some direct transmission links of CRKP in this LTACH. Thus, even the closest donor model may not be completely representative of direct transmission between two patients, and this may be why some transmission risks were not identified in the closest donor model. However, given our strategy of serially sampling every patient in the facility, there is a high probability of direct transmission links being represented in our model outcomes, so risk factors for direct transmission should be picked up even if not all transmission pairs are present in the data. This is supported by our results corresponding with known CRKP risk factors. Additionally, we are not only interested in direct transmission links but also transmission patterns of certain clusters of isolates, both of which could be identified in our models and helpful in infection prevention interventions. Finally, we chose to only include patient-pairs who had overlapping stays in the facility, and thus were unable to identify if sequential room occupation18 was a risk factor in the LTACH. We chose to limit our pairs to those who had overlapping stays to limit the dataset to more likely direct and staff-mediated transmission scenarios, as we did not find sequential room sharing to be a common transmission source in previous work with this facility.13
A caveat of our study design is that it includes demographic and contextual data on only CRKP positive patients. Thus, all inferences are conditional on both members of a pair being colonized. Accordingly, the epidemiologic risk factors identified in our results should be interpreted as driving genomic similarity between isolates from colonized individuals, as opposed to an individual’s risk of colonization. This aspect of our study, however, makes it more accessible, as in many community settings, public health datasets only contain case data, and thus methods that necessitate uninfected controls would be infeasible.
As WGS pathogen data become more widely available and inexpensive to obtain, genomic data have an important role to play in routine surveillance of pathogens such as CRKP. Our analysis shows that these data can provide insights in high-prevalence settings which would not be accessible otherwise. Our results also underscore that the choice of genomic relatedness may be important but is also somewhat flexible and is likely to vary by context and goal of the surveillance activity. For example, infection prevention efforts targeted at mitigating the spread of novel drug-resistant variants may utilize different outcomes than those focused on identifying generic transmission risk factors of endemic pathogens. The approach outlined in this analysis requires few assumptions including no arbitrary SNV cutoffs, no uninfected controls, and only modest computational power and suggests that routine WGS-based surveillance may allow for earlier detection and facility-specific intervention in nosocomial outbreaks of CRKP and other pathogens causing significant morbidity and mortality in vulnerable, hospitalized populations.
Data Availability
Sequence data and limited meta-data are available under National Center for Biotechnology information BioProject PRJNA603790.
Acknowledgements
We thank Josh Warren for his consultation on genomic pairwise regression modeling for infectious disease transmission studies. This research was supported by the National Institutes of Health [R01 AI148259]. H.S., E.S, and J.Z. conceptualized the study and developed methodology. H.S. conducted analyses, created visualizations, and wrote the original draft. M.H. provided data and guidance in interpreting results. T.A. consulted throughout the project. All authors reviewed and edited this manuscript. Authors have no competing interests to report.