Web-based Tool Validation for Antimicrobial Resistance Prediction: An Empirical Comparative Analysis
1Center for Biotechnology, Siksha ‘O’ Anusandhan Deemed to be University, Bhubaneswar, India-751003, Email.Id: sweta.routray6@gmail.com
2Center for Biotechnology, Siksha ‘O’ Anusandhan Deemed to be University, Bhubaneswar, India-751003, Email.Id: prabhaswayam94@gmail.com
3Department of Computer Science & Engineering, Institute of Technical Education and Research, Siksha ‘O’ Anusandhan Deemed to be University, Bhubaneswar, India-751030, Email.Id: swapnesh.nayak@gmail.com
4Department of Computer Science and Bioscience, Faculty of Technology, Marwadi University, Rajkot, Gujarat, India-360003, Email.Id: sejalshah9580@gmail.com
5Department of Computer Application, Institute of Technical Education and Research, Siksha ‘O’ Anusandhan Deemed to be University, Bhubaneswar, India-751030, Email.Id: triptiswarnakar@soa.ac.in
*Corresponding author, Tripti Swarnkar, Email: triptiswarnakar@soa.ac.in, Mob: +918249096553Abstract
Global public health is seriously threatened by Antimicrobial Resistance (AMR), and there is an urgent need for quick and precise AMR diagnostic tools. The prevalence of novel Antibiotic Resistance Genes (ARGs) has increased substantially during the last decade, owing to the recent burden of microbial sequencing. The major problem is extracting vital information from the massive amounts of generated data. Even though there are many tools available to predict AMR, very few of them are accurate and can keep up with the unstoppable growth of data in the present. Here, we briefly examine a variety of AMR prediction tools that are available. We highlighted three potential tools from the perspective of the user experience that is preferable web-based AMR prediction analysis, as a web-based tool offers users accessibility across devices, device customization, system integration, eliminating the maintenance hassles, and provides enhanced flexibility and scalability. By using the Pseudomonas aeruginosa Complete Plasmid Sequence (CPS), we conducted a case study in which we identified the strengths and shortcomings of the system and empirically discussed its prediction efficacy of AMR sequences, ARGs, amount of information produced and visualisation. We discovered that ResFinder delivers a great amount of information regarding the ARGS along with improved visualisation. KmerResistance is useful for identifying resistance plasmids, obtaining information about related species and the template gene, as well as predicting ARGs. ResFinderFG does not provide any information about ARGs, but it predicts AMR determinants and has a better visualisation than KmerResistance.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Introduction
The world’s most severe public health concern is the rapid growth of resistant superbugs, as well as the ongoing battle between bactericidal drugs and extensively drug-resistant bacteria. Antibiotic overuse in both medical and agricultural contexts has aided in the development of multidrug-resistant (MDR) bacteria. Unfortunately, many antibiotics lack specificity, killing pathogenic and non-pathogenic bacteria indiscriminately and leading to antibiotic-associated illnesses (1). The innovation of new anti-bacterial to treat resistant infections is a major goal in healthcare, yet no obvious answer to this problem has been identified. These antibiotic-resistant genes allow bacteria to resist antibiotics in a variety of ways, including the activation of efflux pumps, antibiotic molecule degradation by enzymes, and chemical alteration (ribosome and cell wall) to protect antibiotic-targeted cellular targets. These resistance mechanisms, when combined, represent a danger to the therapeutic effectiveness of antibiotics (2,3). The World Health Organization (WHO) lists AMR as one of the top ten global public health threats to humanity (4). Drug-resistant diseases are also expected to kill ten million people every year by 2050. This indicates that drug-resistant diseases will result in more fatalities than road accidents, diabetes, and cancer combined (5). Antibiotic formulation, testing, and screening are resource-intensive and expensive, limiting possible treatment choices for resistant bacterial species.
ARGs have the ability to threaten public health (6). ARGs are commonly present on transposons or plasmids, and they can be transferred from one cell to another by transduction, transformation, or conjugation. Resistance spreads rapidly within a bacterial population and among different types of bacteria due to gene transfer. This is known as horizontal gene transmission (7). The detection of these genes is essential for identifying resistant strains, validating non-susceptible phenotypes, and better understanding resistance epidemiology (8). Phenotypic assays have traditionally been used to detect AMR. The criterion for determining antibiotic sensitivity is diffusion-based or standardised dilution in vitro Antibiotic Susceptibility Test (AST), and much research and testing has been done to link AST findings with treatment response. Resistance surveillance and in some cases clinical therapeutic guidance are increasingly using molecular approaches. These approaches include everything from PCR-based resistance element detection to mass spectrometry-based methods(9). Sequencing has become a viable approach for routine bacterial characterization due to the enhanced accessibility and cheaper cost of NGS. Over the past few years, it has significantly improved our ability to combat AMR. Although NGS-based technologies can detect practically any known AMR gene or mutation and explore new variations of known AMR determinants, they have largely replaced traditional genotypic approaches for AMR identification (10–12).
Furthermore, as new phenotypic AMR determinants are discovered, sequence data can be continuously stored and re-analyzed. The primary drawback of any genotypic AST approach is that it can detect only proven AMR mechanisms, while resistance induced by novel mechanisms and/or gene expression regulation (heteroresistance, increased efflux pump expression, etc.) cannot be detected (13).
The prediction of AMR in NGS data using a simple and quick method is essential for source tracking, infectious disease detection, diagnostics, and epidemiological surveillance, and additional research is highly required(14). The development of tools for continuously monitoring AMR globally has become more important because these tools may provide useful information to assist healthcare professionals in developing stewardship programmes and implementing public health measures. It is also required to create an infrastructure and network to support this huge amount of data (15). AMR is detected computationally by querying input DNA or amino acid sequence data for the existence of a pre-determined set of AMR determinants provided in AMR reference databases using a search algorithm (Figure 1) (16). Numerous drug resistance prediction tools have been developed and made publicly available online over the last couple of years (17). There is however, a requirement for the standardization of Tools. Our research has the ability to determine the relative advantages and drawbacks of the tools of the predictions generated by the various tools.
Results
This study provides better insights into the advantages and drawbacks of the tools, such as user-friendliness, to determine which tool has the best visualisation and which tool offers the maximum information.
Results have been discussed with the following aspects
- ➢Comparison based on the output resistant plasmid sequences
- ➢Comparison based on the output Antibiotic Resistance Genes (ARGs)
- ➢Comparison based on the amount of information provided and result visualisation
- ➢In detailed result of each tool
Comparison based on the output resistant plasmid sequences
ResFinder, KmerResistance, and ResFinderFG were each provided 25 plasmid sequences to analyse, and these three tools examined all sequences. All three tools predicted AMR for the five common sequences, whereas ResFinder and KmerResistance both predicted 13 sequences that are shared by both. ResFinderFG predicted AMR in only six sequences (Figure 4).
Comparison based on the output Antibiotic Resistance Genes (ARGs)
KmerResistance detected 15 of the 16 AMR genes predicted by ResFinder, along with four additional resistance genes. KmerResistance did not predict one of the genes predicted by ResFinder (Figure 5). ResFinderFG, on the other hand, cannot detect any AMR genes and can only predict the AMR determinants.
Comparison based on the amount of information provided and result visualisation
The results of all three tools are made available through email. The outcomes of the analysis were displayed in a tabular manner in ResFinder, with the first table containing Antimicrobial, Class, WGS-predicted phenotype, and Genetic background. Other tables in ResFinder are organised by drug class, with each table containing various drug class information such as Resistance gene, Identity, Alignment Length/Gene Length, Position in reference, Contig or Depth, Position in contig, Phenotype, PMID, Accession no., and Notes. There is an extended output option at the bottom of the result page that displays the alignment result. The results for KmerResistance are provided in a single table which contains the following columns: Template, Score, Expected, Template Length, Q Value, P Value, Template Id, Template Coverage, Query Id, Query Coverage, Depth, and Depth Corr. The ResFinderFG result is displayed as a tabular box with numerous columns containing Hit name, Identity, Query/HSP, Contig, Position in contig, Drug treatment, and Accession no. Each of the three tools offers a variety of downloadable files containing various types of information (Table 3).
In detailed result of each tool
As a case study, we took 26 plasmid sequences and ran them through three different prediction tools to identify AMR factors.
ResFinder 4.1
ResFinder 4.1 predicted that plasmid X60321.1 contains the majority of AMR genes. It carries two AMR genes, aac(6’)-Ib3 resistance to aminoglycosides and aac(6’)-Ib-cr resistance to fluoroquinolone classes of antibiotics. The presence of AMR genes was found in 18 of the 25 plasmid sequences tested in this study, with no resistance gene found for thirteen of the drug classes: Colistin, Disinfectant, Fosfomycin, Fusidic acid, Glycopeptide, MLS (Macrolide, Lincosamide, and Streptogramines), Nitroimidazole, Oxazolidinone, Phenicol, Pseudomonic acid, Rifampicin, Tetracycline, and Trimethoprim. A fair number of AMR gene homologues were identified in the other four drug classes (Supplementary Table 1). The drug classes in which AMR genes were detected were beta-lactams (13 out of 19; 68.42%), followed by aminoglycosides (21.05%), fluoroquinolones (5.2%), and sulphonamides (5.2%). There were 16 number of different acquired AMR genes found to be resistant to the four drug classes, with aminoglycosides and beta-lactams showing the highest frequency (Table 4). The most common genes are aadA10 (2/19), blaNPS (2/19), and blaPAU-1 (2/19), each of which was found to be present in 10.52% of the outcomes observed (Figure. 6)
KmerResistance 2.2
For an in-silico analysis of AMR genes, the sequences were uploaded to the KmerResistance database. Eight plasmid template genes were linked to other Gram-negative bacteria, while one plasmid template gene was linked to a Gram-positive bacteria, indicating that they could be the source of acquisition (Supplementary Table 2). Two Escherichia coli, one each of Klebsiella pneumoniae, Comamonas testosterone, Acinetobacter sp., Pseudomonas putida, Achromobacter xylosoxidans, and Providencia sp. linked with Gram-negative strain. Glutamicibacter nicotianae Gram-positive strain is associated with one of the plasmids. Out of 25 input plasmid sequences, it predicted 19 ARGs in 13 plasmids.
ResFinderFG 2.0
In ResFinderFG 2.0 three distinct ARD families were identified in six input plasmid sequences. The most common was beta-lactamase, which was found in four different input plasmid sequences. Three of the four plasmids with beta-lactamase ARDs conferred resistance to ampicillin and one conferred resistance to piperacillin. Others include one dihydropteroate synthase (dpr) and one aminoglycoside acetyl-transferase (AAC), as both are resistant to Sulfamethazine and Amikacin respectively (Supplementary Table 3).
Discussion
Currently, the use of sequencing technology is revolutionizing practically every component of biological study. In the field of infectious diseases, scientific discoveries, as well as diagnostic and outbreak investigations, are developing at a rapid pace. The ability to interpret sequencing data and the benefit of quick development, on the other hand, is not evenly distributed between institutions and countries (39,40). We choose CGE tools because its goal is to give access to bioinformatics tools for those with limited knowledge, allowing all countries, institutions, and individuals to benefit from new sequencing technology. It is believed that by doing so, it will encourage more open data sharing around the world and give equal advantages to all. CGE is fully non-profitable and provides a variety of free online bioinformatics services (19).
The AMR prediction methods have been built using DNA or amino acid sequence data. The presence or absence of software for searching within an AMR determinant database, which can be precise to a tool or replicated from other resources, the type of input data accepted, and the search approach used, which can be alignment or mapping, are all factors that distinguish bioinformatics resources. Each tool has its own set of capabilities and limitations when it comes to AMR prediction, which includes the identification of AMR sequences, ARGs, the volume of information produced, and visualisation. As the web-based application is one that uses a website as its interface or front-end. It has the potential to provide competitive benefits over traditional software-based systems by providing researchers to streamline data and information at a lower cost, time, and maintenance (19). Using a regular browser, users can quickly access the application across any computer with internet access. Functionality and features were the two main elements that were prioritized. From this study, a non-bioinformatician or non-technician can gain subject-specific knowledge as well as determine which tool is appropriate for their specific work. The table below compares the three tools used for identifying drug resistance genes and resistance plasmids as well as the amount of information they provide and the way their results are displayed. (Table 5).
The description and significant observations of each selected web-based tools have been discussed below
ResFinder is a web-based tool that finds chromosomal alterations that promote antibiotic resistance in bacteria’s whole or partial DNA sequence and identifies acquired antimicrobial resistance genes in the whole-genome data using BLAST. As input, this tool accepts pre-assembled, whole or partial genomes, along with fragmented sequence reads from four distinct sequencing technologies. It is accessible at (https://cge.food.dtu.dk/services/ResFinder/). It is also constantly being updated whenever different resistance genes are discovered (41). Here, we found that the database showed a restriction that stated it only looked for acquired genes and did not detect chromosomal mutations. Since new resistance genes are continuously discovered, it may be necessary to confirm the presence or absence of identified AMR genes phenotypically. According to a study (42), genotyping using aligned whole-genome sequences is a practical substitute for surveillance based on phenotypic antimicrobial susceptibility due to the high concordance (99.74%) between phenotypic and predicted whole-genome sequence antimicrobial susceptibility. One of the three most common genes, the aadA10 gene, was found to be resistant to the drug class aminoglycoside, while the other two genes, blaNPS and blaPAU-1, were found to be resistant to the beta-lactam drug class. AadA10 is a class 1 integron containing gene cassette, which suggests that rather than transposition, the resistance determinants from one plasmid to another plasmid were moved by recombination between two class 1 integrons (43). The ability of integrons in bacteria to acquire new cassettes and recombine cassette rows emphasises the adaptability of integron diversity. It is necessary to be aware of what other integron-mediated traits, such as increased resistance to antimicrobials, virulence, or pathogenicity, might affect human health in the future given their capacity to quickly spread resistance phenotypes. There is an urgent need for control integrons and cassette formation (44). Tauch et al., 2003, found that blaNPS can rapidly transfer from one species to another (45). According to Subedi et al., 2018, environmental resistance gene pools contain blaNPS, which can be acquired and maintained in clinical isolates. (46). In a prior study, it was discovered that a clinical isolate of Pseudomonas aeruginosa contained a transferable plasmid containing the gene known as blaPAU-1, which is connected to the mobile genetic element (47). Considering the high ubiquity of beta-lactam resistance genes, it is preferable to monitor the level of antibiotic resistance and resistance genes in patients with Pseudomonas aeruginosa infection (48). In this study, we found that Pseudomonas aeruginosa plasmids contain many beta-lactam resistance genes, a number of which are found in a single plasmid. This leads us to the conclusion that beta-lactam resistance genes are rapidly spreading through plasmids, and the need of the hour is to control their spread.
KmerResistance (https://cge.food.dtu.dk/services/KmerResistance/) is primarily based on KmerFinder, which has been developed for typing of bacteria with raw WGS data. KmerResistance and KmerFinder search for co-occurring k-mers between such a query genome and a resistance gene database. KmerResistance, like KmerFinder, biases the threshold based on the quality of the data, as shown by the coverage as well as the depth of the detected species genome. Since this k-mers in this scenario are dispersed across the total sample, we can estimate both depth and coverage. Unlike KmerFinder, KmerResistance may generate an outcome for species prediction in addition to getting acquired antimicrobial resistance genes.(49). This database improves on poor-quality assembly by using k-mers to map raw whole-genome sequence data against reference databases and species. (Fragments of a DNA sequence of length k) (28). Additionally, it can find host or template genes. The KmerResistance database displays the resistance genes but not the drug classes as an analysis output, even though being claimed to be more precise than ResFinder. As a result, comparisons were restricted to resistance genes that were present in both databases rather than an overall assessment of their sensitivity. ResFinder and ResFinderFG accept input sequences in a single input file and provide results based on each sequence, while KmerResistance requires us to provide the sequence file separately and perform separate executions. If we provide the input sequence to KmerResistance in a single file, the results are very perplexing and unrelated to the input sequence.
The ResFinderFG (https://cge.food.dtu.dk/services/ResFinderFG/) approach is based on databases containing sequences detected by functional metagenomics but not represented in existing databases constructed mostly from antibiotic-resistant genes in clinical isolates. It identifies a resistant phenotype in general (50). One of the main causes of beta-lactam resistance in Pseudomonas aeruginosa is the increased prevalence of beta-lactamase (48). Beta-lactamase enzymes render beta-lactam antibiotics ineffective by hydrolyzing the peptide bond of the characteristic four-membered beta-lactam ring. The bacterium gains resistance after the antibiotic is rendered inactive. Over 300 beta-lactamase enzymes have been described so far, with numerous kinetic, structural, computational, and mutagenesis studies. The threat posed by more and more powerful beta-lactamases to antimicrobial therapy is only going to increase (51). The scientific and medical communities are only one step ahead and must continue to put forth diligent effort to prevent being overcome by the difficult and rapidly rising nosocomial pathogen resistance.
Conclusion
A variety of freely accessible tools are available for the prediction of AMR determinants. A Comparison of these tools will aid in the expansion of global pathogen monitoring and AMR tracking based on genomic information. Since AMR tools are used to accomplish various tasks, it would be unfair and difficult to compare them directly. We provided a case study using bioinformatics tools that allow researchers to predict AMR. Here we discovered that ResFinder is very advantageous for the prediction of resistant plasmid identification and provides a large amount of information with better visualisation, whereas KmerResistance is quite good for resistance plasmid identification, information regarding related species and the template gene, as well as predicting ARGs. ResFinderFG does not provide any information about ARGs and provides relatively lesser information than the other two tools, although the visualisation is superior to KmerResistance. Each has its own unique strengths and weaknesses. Our research provides a thorough overview of the results obtained by three different tools, allowing users to select the best tool for their needs.
Materials and methods
The framework employed in this study is described in the subsequent sections, and the comprehensive methodology is depicted in the figure below (Figure 2).
Validation dataset
Pseudomonas aeruginosa is a pathogenic Gram-negative bacterium that is becoming more common in infections caused by Multidrug-resistant (MDR) and extensively drug-resistant (XDR) strains, limiting available effective treatments. Plasmids contribute significantly in antibiotic resistance because they are the means by which resistance genes are captured and then disseminated. Pseudomonas aeruginosa plasmid nucleotide sequences were acquired using the key phrase “Pseudomonas aeruginosa” from the nucleotide database of the NCBI (https://www.ncbi.nlm.nih.gov/nuccore/). The organism was then chosen as Pseudomonas aeruginosa, the species as bacteria, the molecular type as genomic DNA/RNA, the sequence type as nucleotide, genetic compartment as plasmid, and the length from 900bps to 1100bps.
In order to obtain a standard range, the length of the sequence was taken into account by the Romaniuk et al., 2019 article (18). We obtained 26 Pseudomonas aeruginosa plasmid sequences from the aforementioned screening, which are further considered for the AMR prediction. The table (Table 1) below includes NCBI accession numbers, plasmid length and species of the retrieved sequences.
Segregation of in silico AMR determination tools
The development of online databases and bioinformatics tools has been required for drug-resistance gene prediction. After conducting a scientific literature search from 2012 to 2022, Forty-seven freely accessible bioinformatics tools for identifying AMR determinants were discovered, including the most commonly used tools listed below (Table 2).
Center for Genomic Epidemiology (CGE) is fully non-profitable and provides a variety of free online bioinformatics services. The Technical University of Denmark (DTU) provides core funding, as well as funding from a variety of public and commercial sources (19). The Center for Genomic Epidemiology offers 38 services, including nine phenotyping tools: ResFinder, ResFinderFG, LRE-finder, KmerResistance, PathogenFinder, VirulenceFinder, Restriction-ModificationFinder, SPIFinder, and ToxFinder (Figure 3). From the mentioned nine phenotypic tools, we selected three tools with versions (ResFinder 4.1, KmerResistance 2.2, and ResFinderFG 2.0) based on similarities of their objective i.e., to find out the resistance factors. Here we did not include LRE-Finder in our study because the input data format is FASTQ sequence, it only predicts AMR for a single species i.e, Enterococcus faecalis, and only identifies acquired linezolid resistance genes without considering other antibiotics.
Steps and parameters setup for segregated tools
ResFinder 4.1 searched the database for all seventeen classes of antibiotic drugs, regardless of the target region. The sequences were entered into the database, and the acquired resistance gene testing parameters were adjusted to predict resistance genes for all seventeen drug classes provided by the server. The minimum percentage identity was set to 90%, with perfect alignment set to 100%. The percentage of identity was computed by counting the number of identical nucleotides between the best-matching resistance gene in the database and the equivalent sequence in the plasmid. The tool was run with the aforementioned parameters, and the results were recorded.
The scoring method in KmerResistance 2.2 was species identification on maximum query coverage, and also the host database was set to the bacterial plasmid. The gene database was set to resistant genes, and the identity threshold was left at 70%, with a depth correction threshold of 10%. The AMR genes identified in the resulting output were recorded and the host organism and template sequence were also noted.
The functional genomics database ResFinderFG 2.0 uses functional metagenomic antibiotic resistance determinants to identify resistance phenotypes. The percentage identity setting was set to 98%, along with the minimum query length was set to 60%. The read type used was ‘assembled contigs/genomes,’ and the sequences were screened for all the thirteen antibiotic resistance determinant (ARD) families present in the database.
Acknowledgement
The author acknowledges and thank the Center for Genomic Epidemiology in Denmark for providing web tools.
Conflict of interest
Authors declare no conflict of interest.