Detection of ESKAPE pathogens and Clostridioides difficile in Simulated Skin Transmission Events with Metagenomic and Metatranscriptomic Sequencing
aSignature Science, LLC, 8329 North Mopac Expressway, Austin, Texas, USA
bSignature Science, LLC, 1670 Discovery Drive, Charlottesville, VA, USA
#Address correspondence to Krista L. Ternus, kternus@signaturescience.comAbstract
Background
Antimicrobial resistance is a significant global threat, posing major public health risks and economic costs to healthcare systems. Bacterial cultures are typically used to diagnose healthcare-acquired infections (HAI); however, culture-dependent methods provide limited presence/absence information and are not applicable to all pathogens. Next generation sequencing (NGS) has the capacity to detect a wide variety of pathogens, virulence elements, and antimicrobial resistance (AMR) signatures in healthcare settings without the need for culturing, but few research studies have explored how NGS could be used to detect viable human pathogen transmission events under different HAI-relevant scenarios.
Methods
The objective of this project was to assess the capability of NGS-based methods to detect the direct and indirect transmission of high priority healthcare-related pathogens. DNA was extracted and sequenced from a previously published study exploring pathogen transfer with simulated skin containing background microorganisms, which allowed for complementary culture and metagenomic analysis comparisons. RNA was also isolated from an additional set of samples to evaluate metatranscriptomic analysis methods at different concentrations.
Results
Using various analysis methods and custom reference databases, both pathogenic and non-pathogenic members of the microbial community were taxonomically identified. Virulence and AMR genes known to reside within the community were also routinely detected. Ultimately, pathogen abundance within the overall microbial community played the largest role in successful taxonomic classification and gene identification.
Conclusions
These results illustrate the utility of metagenomic analysis in clinical settings or for epidemiological studies, but also highlight the limits associated with the detection and characterization of pathogens at low abundance in a microbial community.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Footnote Group
3Introduction
The estimated number of annual deaths due to infections from multidrug resistant organisms is upwards of ∼70,000 for individuals within inpatient hospital care and ∼80,000 for those in outpatient care in the United States, based on 2010 mortality rates (Burnham et al., 2019). The ESKAPE pathogens, consisting of Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, and Enterobacter species (Boucher et al., 2009), are responsible for many drug-resistant healthcare-acquired infections (HAIs) (Boucher et al., 2009; Santajit and Indrawattana, 2016). Along with Clostridioides difficile (Slimings and Riley, 2014), these pathogens are the leading causes of nosocomial infections (Boucher et al., 2009; Santajit and Indrawattana, 2016). Culture-based methods within clinical laboratories are typically utilized to identify and track HAI transmission, such as the nosocomial infections caused by ESKAPE pathogens and C. difficile (ESKAPE+C), (Didelot et al., 2012), but cultures have multiple drawbacks. Dead or unculturable pathogens will be overlooked by culture-dependent methods, even though usable biochemical signatures (e.g., DNA) persist. Culturing is primarily a method for identifying viable pathogens amenable to growth under certain conditions, aiming to confirm the presence of known pathogens at the species level. Once a putative pathogen species has been identified, multiple rounds of culturing and biochemical assays may be necessary to further characterize pathogens at the strain level or to identify antibiotic resistance activity.
Metagenomic and metatranscriptomic analyses of samples collected in a healthcare setting provide compelling alternatives to traditional culture-based pathogen identification. These analyses do not require pathogen viability or culturability; instead, collected cells are lysed and the nucleic acids are collected for sequencing. These approaches permit species or even strain level identifications of pathogens present within a sample without multiple rounds of culture analysis. Perhaps most importantly, sequencing approaches can provide valuable insights into gene content and expression, identifying components of the resistome and elements contributing to virulence in a clinical sample. Previous studies have evaluated the relationship between culture and metagenomic analysis, highlighting both successes and challenges for this technology (Didelot et al., 2012). Challenges of unbiased metagenomic or metatranscriptomic sequencing methods include complexities in developing standardized analysis protocols and databases, and pathogen concentrations falling below the limit of detection in relation to other organisms in the sample. In the current study, we constructed customized databases based on the known mock microbial community genome and gene content to explore the impact of different ESKAPE+C concentration levels and simulated HAI transfer scenarios on pathogen detection from metagenomic and metatranscriptomic sequence data.
Our research expands upon previously published data from a study establishing an in vitro method to model ESKAPE+C transmission using a synthetic skin surrogate (Weber et al., 2020). This prior study enabled the investigation of both direct (skin-to-skin) and indirect (skin-to fomite-to skin) pathogen transmission scenarios using VITRO SKIN® N-19 to mimic human skin, including a simulated commensal skin flora (Figure 1). The commensal skin flora was included on both the pre-transfer and post-transfer coupon to simulate pathogen transfer from skin containing a mix of pathogen and commensal organisms to a second piece of skin containing only the existing commensal community. Different transfer scenarios of ESKAPE+C species, including multiple wash or decontamination steps and high or low spike-in concentrations, were evaluated using culture analysis. Additionally, nucleic acids were extracted from all sample replicates to compare sequence data with culture results. The resulting sample set had a wide range of relative pathogen abundance in comparison to the commensal community, which was ideal for evaluating metagenomic and metatranscriptomic analysis methods. Here, we present the results, contrasting the utility of metagenomic and metatranscriptomic analysis across a range of pathogen abundance within simulated clinical samples.
4Materials and Methods
4.1Bacterial Isolates and Sequence Data
Microorganisms used for this effort were sourced from American Type Culture Collection (ATCC) or the Centers for Disease Control and Prevention (CDC) Antimicrobial Resistance (AR) Isolate Bank as described previously (Weber et al., 2020). Any isolates that did not have existing published whole genome sequencing reference data were sequenced internally on an Illumina MiSeq® FGx System. The new isolate sequence data produced by this study included Enterococcus faecium (CDC AR Bank #0579), Clostridioides difficile (ATCC 43598), Brevibacterium linens (ATCC 9172), Corynebacterium matruchotii (ATCC 14265), Cutibacterium acnes (ATCC 11827), Escherichia coli (ATCC 9637), Lactobacillus gasseri (ATCC 33323), Micrococcus luteus (ATCC 4698), Staphylococcus epidermidis (ATCC 12228), and Streptococcus pyogenes (ATCC 19615). All raw FASTQ sequences were submitted to NCBI SRA and subsequently processed for further analysis in this study. This included pre-processing to ensure a high quality of reads, mapping to reference genomes with 21x – 352x average genome coverage, and undergoing downstream assembly, gene identification, and taxonomic classification (Supplementary Tables 1-6).
Methods for bacterial culture, laboratory mixes, VITRO-SKIN® coupon preparation for transfer experiments, and analysis of the culture data were previously published (Weber et al., 2020). Briefly, VITRO-SKIN® N-19 (IMS Inc.) coupons were inoculated with “background organisms” to simulate a commensal skin microbial community at a constant concentration, and ESKAPE+C pathogens were added at two different high and low concentrations. Coupons were then allowed to dry before being touched to a second coupon containing only the commensal organisms, followed by a simulated handwashing step or non-handwashing step. Coupons were also transferred to a fomite surface (i.e., cotton, nitrile, stainless steel coupon), which was then subject to washing, decontamination, or no treatment before transfer to a skin coupon containing only commensal organisms.
4.2Nucleic Acid Extractions
Each sample from the previous study was split, with one fraction used for culture analysis and the remaining sample used for nucleic acid extractions (Weber et al., 2020). Bacteria suspended in phosphate buffered saline (PBS) recovery buffer were transferred into 15 mL conical tubes and centrifuged at 5,000 x g for 20 minutes. The supernatant was removed, and DNA was extracted from the cell pellet using the ZymoBIOMICS DNA/RNA Miniprep Kit per manufacturer instructions. Total DNA yield was quantified using the Qubit™ dsDNA BR Assay Kit (Thermo Fisher Scientific) or Qubit™ dsDNA HS Assay Kit (Thermo Fisher Scientific) as appropriate per manufacturer instructions.
For the RNA control mixtures, a culture mixture was prepared that contained equal amounts (∼106 CFU/mL) of all background bacteria and pathogens, representing the high pathogen concentrations for subsequent transfer events. Additional control of culture mixtures of all pathogen and background bacteria were made by titrating in equal amounts of each pathogen at ∼102, ∼104, or ∼106 CFU/mL with a consistent amount of background bacteria at equal amounts (∼106 CFU/mL) to mimic relative abundances commonly reported from the human hand in healthcare settings (WHO Guidelines on Hand Hygiene in Health Care: First Global Patient Safety Challenge Clean Care Is Safer Care, 2009). Samples were centrifuged at 5,000 x g for 20 minutes. The supernatant was removed, and RNA was extracted from the cell pellet using the ZymoBIOMICS DNA/RNA Miniprep Kit as per manufacturer instructions. Total RNA yield was quantified using the Qubit™ RNA HS Assay Kit (Thermo Fisher Scientific) as per manufacturer instructions.
4.3RNA Preparation and cDNA Conversion
All RNA samples were converted to cDNA for subsequent library preparation and sequencing. First, sample mRNA was enriched using the MICROBExpress™ Bacterial mRNA Enrichment Kit (Thermo Fisher Scientific), followed by additional sample clean-up using the MEGAclear™ Transcription Clean-Up Kit (Thermo Fisher Scientific). Samples were then converted to cDNA using the SuperScript™ Double-Stranded cDNA Synthesis Kit (Thermo Fisher Scientific). All kits were used according to the manufacturer instructions. After each process, nucleic acid quantification was measured using either Qubit™ RNA HS Assay Kit or Qubit™ dsDNA HS Assay Kit (Thermo Fisher Scientific), as appropriate.
4.4Library Preparation and Sequencing
Prior to library preparation, all samples (DNA for genomic analysis or cDNA for transcriptomic analysis) were diluted to 0.2 ng/µL based on measured Qubit™ values. For samples below 0.2 ng/µL, no dilution was performed. All samples were prepared for sequencing using the Nextera® XT DNA Library Prep Kit (Illumina) as per the manufacturer’s protocol using AMPure XP (Beckman Coulter) bead-based normalization. Samples were diluted and denatured based on the recommended Nextera® XT DNA bead-based normalization loading concentration. PhiX control (Illumina) was added at a final concentration of 1% to the denatured library, and libraries were loaded on a MiSeq® Reagent v3 (Illumina) cartridge for sequencing. Sequencing was performed on a MiSeq® FGx System (Illumina) in research use only (RUO) mode using the MiSeq® Reagent Kit v3 (Illumina) with paired-end read lengths of 75 base pairs and results generated in FASTQ files. A bioinformatics quality control pipeline verified that all FASTQ files were of sufficient quality for downstream processing and analysis.
4.5Bioinformatics Analysis
All bioinformatics tools and databases used in this analysis were open source, and more information, including the commands used with each of the bioinformatics tools is available in Supplementary Table 7. Quality control for sequence data analysis was performed with FastQC (Andrews, S, 2010), Trimmomatic (Bolger et al., 2014), and MultiQC (Ewels et al., 2016). Mash distance was used to identify the closest available reference genome with the trimmed read data (Ondov et al., 2016). SPAdes assemblies were generated from the high-quality genome sequence data for each isolate (Bankevich et al., 2012), evaluated with QUAST (Gurevich et al., 2013), and the best preexisting genome assembly was identified with the Mash distance from all genomes available in NCBI RefSeq (O’Leary et al., 2016) and GenBank® (Clark et al., 2016) at the time of this study (Supplementary Tables 3 and 4). Multiple genome alignments were performed with progressiveMauve to compare gene gain, loss, and rearrangement of the E. faecium isolate (Darling et al. 2010, Supplementary Figure 1). Genes were annotated de novo with prokka (Seemann, 2014) from assembled contigs or detected by alignment with ABRicate (Zankari et al., 2012) using a custom database of expected genes based on prior annotations from known strains. Taxonomic analysis of metagenomes and metatranscriptomes was performed by Bowtie2 (Langmead and Salzberg, 2012), SAMtools (Li et al., 2009), and Qualimap2 (Okonechnikov et al., 2016) mapping of reads to the expected genomes, and Mash Screen (Ondov et al., 2019) containment estimations with a custom reference genome database. When equivalent hits were identified for one genome, the larger number was selected to simplify final reporting and analysis. Bowtie2 was used to map reads to a custom database of genes present within the isolates. Metagenomic assembly was performed with metaSPAdes (DNA) (Nurk et al., 2017) or rnaSPAdes (RNA) (Bushmanova et al., 2019) before downstream gene annotation with prokka and ABRicate. Metagenomic reads derived from DNA extracted from contact scenarios and RNA from controlled spike-ins of pathogen to background organisms aligned to genes within the custom database using Bowtie2. Data analysis figures were generated with R Studio, all sequence data were submitted to NCBI BioProject 530203, and intermediate analysis results from the tools can be found at OSF project https://osf.io/3qwps/.
5Results
5.1Modeling Transmission
Bacterial transmission can occur through either a direct skin-to-skin situation, such as between an infected patient and a health care worker, or an indirect skin-to fomite-to skin scenario, where an infected patient touches an object (i.e., cotton, nitrile, stainless steel surface) followed by a healthcare worker touching the same surface. To simulate these two contact scenarios, approaches were developed in Weber et al. 2020 to investigate direct and indirect ESKAPE+C pathogen transfer utilizing a synthetic human skin material VITRO-SKIN® N-19 with or without including a relevant handwashing step. As pathogen transfer does not occur in isolation, a set of background bacteria were included on the VITRO-SKIN® coupon to represent the native skin microbiota that could be present on human hands. Supplementary Tables 1-6 and Supplementary Figures 1-2 describe features of the ESKAPE+C pathogen and background organisms analyzed, as well as their closest matching reference genome in NCBI databases. While culture data was successfully collected from all direct and indirect scenarios previously described (Weber et al., 2020), the indirect wash scenarios were not sequenced in this study because it was anticipated that they would fall well below the sequencing limit of detection (Figure 1).
5.3Impact of Simulated Handwashing on DNA Yield
ESKAPE+C pathogens were cultured together in three different mixes based on their media and growth requirements (Weber et al., 2020). Samples were split between the culturing and metagenomics sequencing experiments, and the results were compared downstream. Two or three pathogens were included in each cultured mix, and eight background microorganisms were included along with the pathogens before each metagenome was sequenced (Figure 1). Therefore, the reads within a metagenome that did not map to an ESKAPE+C reference genome in Table 1 and Figure 4 had originated from either 1) another pathogen in the mix or 2) one of the background microorganisms. The CFU/mL values calculated from the culture data were compared to the percentage of metagenomic reads mapped to one of the ESKAPE+C reference genomes. Pathogens were only detected from metagenomics data in the high spike-in direct contact scenarios (Figure 3), such as direct transfer events with and without handwashing. There was no observable correlation between CFU/mL and reads mapped to the reference genomes (Supplementary Figure 3).
While the simulated handwashing events on VITRO-SKIN® decreased the number of viable pathogen cells, they resulted in higher overall pathogen detection relative to the non-handwashing scenarios rates within metagenomes in the high spike-in (∼106 CFU/mL) direct contact scenarios (Table 1, Figure 4). Most ESKAPE+C species yielded more DNA after the handwash step compared to no handwashing, suggesting that the handwash helped to lyse the bacterial cells and release more DNA for sequencing. Gram-negative Klebsiella aerogenes (formerly known as Enterobacter aerogenes), A. baumannii, and K. pneumoniae, as well Gram-positive E. faecium, showed the largest maximum increases in DNA yield after the simulated handwashing scenarios. Gram-positive S. aureus and C. difficile endospores also showed a relatively smaller impact of handwashing on increased DNA yield. Gram-negative P. aeruginosa and Enterobacter cloacae were the exceptions, with no handwash scenarios showing a larger maximum or average DNA yields compared to handwash scenarios. The relatively thin layers of peptidoglycan in the Gram-negative cell walls of P. aeruginosa and E. cloacae may have allowed the microorganisms to be more effectively lysed and DNA recovered with or without the handwash step, although this trend was not observed for all Gram-negative bacteria in Table 1.
Discussion
Culturing of bacteria is commonly performed to identify infectious pathogens (Nekkab et al., 2017), but culture-dependent methods have inherent limitations. Established nucleic acid-based detection approaches like polymerase chain reaction (PCR) or 16S rRNA sequencing overcome some of these limitations, including the requirement for pathogen viability or culturability; however, these targeted methods are only able to detect limited, known genomic regions and fall short of identifying pathogens at the strain or even species level. Metagenomic analysis methods hold substantial promise in overcoming these hurdles. By drawing upon an established, well-characterized data set, we have illustrated many of the strengths and weaknesses posed by unbiased metagenomic and metatranscriptomic sequencing. Although unbiased sequencing methods have enormous potential for pathogen detection, this study has highlighted limitations in failing to detect low abundance pathogens in mixed samples (< 104 CFU/mL), challenges with determining the species of origin for gene sequences that are partially or fully shared among multiple species (e.g., AMR genes), and variable rates of nucleic acid extraction efficiency in different bacteria.
After mapped sequences of DNA extracted from different contact scenarios were compared to previously reported ESKAPE+C colony counts through selective culturing, the mapped reads in direct transfer scenarios demonstrated considerable variation in which pathogen species were detected and how the conditions of simulated handwash or no handwash impacted the results. Only three replicates were performed per condition in this study, and it is possible that additional technical replicates could reduce the amount of observed variability. However, additional technical replicates would not overcome variation due to inherent biological characteristics, like differences in cell wall and membrane structures influencing the amount of DNA extracted (e.g., C. difficile endospores). Differences in NGS library preparation methods can also introduce bias in the detection of different bacterial species (Morgan et al., 2010; Van Dijk et al., 2014), and the results from handwash vs. non-handwash scenarios may be conceptually similar to different sample processing and cell lysing methods before sequencing. Bias can also be introduced in CFU measurements, as selective media utilize intrinsic attributes of a bacteria to isolate and differentiate from other bacteria, and this selective pressure may also eliminate viable or injured bacteria that are unable to recover once plated (Apajalahti et al., 2003).
It was promising that genes were detected within microbial mixtures despite the lack of detection with culturing techniques or genome-level taxonomic calls. This suggests that functionally informative genes (e.g., virulence, AMR) could have lower the limits of detection than culturing or standard taxonomic identification methods. A. baumannii was at or near the limit of detection for the culturing method (left of dotted line in Supplementary Figure 5), while specific genes from these species were detected by sequencing. Compared to whole genome techniques, the method of assembly and annotations of organism specific genes led to greater retention of pathogen signal, especially from the no handwash scenarios. The comparison of ABRicate annotations of assembled contiguous sequences with Bowtie2-mapped bases demonstrates that while alignment of short reads allowed for more sensitive detection of genes, the annotation of assembled reads provided better precision in gene identifications. High precision was especially true when at or near 100% coverage was achieved for a species-specific gene within a contiguous sequence.
Bacterial nucleic acids from skin and nonsterile specimens may result in a stronger signal than that of the pathogen (Gu et al., 2019), and in this study that was simulated by a constant level of “background” bacterial species. The lack of signal from ESKAPE+C pathogen sequences within the indirect scenarios using unbiased NGS methods compared to selective culturing techniques could be attributed to the lack of selection for pathogen signal over background microorganism signal in greater quantity, highlighting a key challenge of NGS in clinical settings with nonsterile specimens and complex sample types. Enrichment techniques to amplify the pathogen signal before sequencing could have improved the limits of detection for ESKAPE+C in all scenarios, but enrichment techniques often come at the price of biased sequencing in search of known targets. Such methods have limited applicability to emerging or novel pathogens, as well as situations when the infectious agent is unknown to the physician and fails standard clinical tests.
6Conclusions
Metagenomic and metatranscriptomic analyses promise an unbiased approach to pathogen species-level detection and functional gene characterization within clinical samples; however, the limitations of this technology must be fully evaluated before traditional culturing methods can be supplemented or replaced by this new methodology. This study makes significant progress toward this goal, capitalizing on a large, well-curated data set from a previously published study and generating complimentary NGS analyses for comparison. In doing so, we illustrate both the strengths of this type of analysis, such as the ability to identify pathogens and characterize elements of virulence or the resistome in a given sample, as well as the limitations of unbiased sequencing, predominantly highlighted by low sensitivity when pathogens are present at low abundance within a complex mixture. Variability in the loss of signal from different bacterial species also lends support to how laboratory and bioinformatics methods are impacted by the intrinsic nature of the organisms, as it relates to nucleic acid extraction and the uniqueness of genome content. These results will inform and aid the healthcare and epidemiological community as they evaluate the appropriate scenarios to utilize metagenomic analysis.
Supporting information
Abbreviations
- HAI
- Healthcare-acquired infections
- NGS
- Next Generation Sequencing
- ESKAPE+C
- Enterococcus faecium, Staphylococcus aureus, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Enterobacter species and Clostridioides difficile
- CDC
- Centers for Disease Control and Prevention
- AR
- Antibiotic resistance
- AMR
- Antimicrobial resistance
- ATCC
- American Type Culture Collection
- CFU
- Colony forming units
- PBS
- Phosphate buffered saline
- mRNA
- messenger ribonucleic acid
- mL
- milliliter
- polymerase chain reaction
- PCR
- Enterococcus faecium
- E. faecium
- Enterobacter cloacae
- E. cloacae
- Staphylococcus aureus
- S. aureus
- Klebsiella pneumoniae
- K. pneumoniae
- Klebsiella aerogenes
- K. aerogenes
- Acinetobacter baumannii
- A. baumannii
- Pseudomonas aeruginosa
- P. aeruginosa
- Clostridioides difficile
- C. difficile
- Brevibacterium linens
- B. linens
- Corynebacterium matruchotii
- C. matruchotii
- Cutibacterium acnes
- C. acnes
- Escherichia coli
- E. coli
- Lactobacillus gasseri
- L. gasseri
- Micrococcus luteus
- M. luteus
- Staphylococcus epidermidis
- S. epidermidis
- Streptococcus pyogenes
- S. pyogenes
Acknowledgements
The authors would like to thank Mr. Jim Gibson for his assistance creating the contact scenario figure and Dr. Madeline Roman for her review of this manuscript. The authors would also like to thank Drs. Alison Laufer Halpin and Rachel Slayton of the CDC for their constructive feedback throughout this project.
10Funding Disclosure
This work was supported by the Centers for Disease Control and Prevention’s investments to combat antibiotic resistance under award number 200–2018-75D30118 C02922 (https://cdc.gov).
11Availability of data and materials
All data generated or analyzed during this study are included in this published article, its supplementary files, our OSF project (https://osf.io/3qwps/), or the NCBI BioProject 530203.
12Conflict of Interest Statement
The authors declare no personal, professional or financial relationships that could potentially be construed as a conflict of interest.