Ultraviolet Rate Constants of Pathogenic Bacteria: A Database of Genomic Modeling Predictions
1Sanuvox; Montreal, Quebec, Canada
2The Pennsylvania State University, University Park, PA, USA
3Infectious Diseases Translational Research Laboratory, Weill Cornell Medicine of Cornell University, New York City, NY, USA
*Correspondence: Wladyslaw Kowalski, Sanuvox Technologies, Inc., wkowalski@sanuvox.comAbstract
A database of bacterial ultraviolet (UV) susceptibilities is developed from an empirical model that correlates genomic parameters with UV rate constants. Software is used to count and evaluate potential ultraviolet photodimers and identifying hot spots in bacterial genomes. The method counts dimers that potentially form between adjacent bases that occur at specific genomic motifs such as TT, TC, CT, & CC. Hot spots are identified where clusters of three or more consecutive pyrimidines can enhance absorption of UV photons. The model incorporates nine genomic parameters into a single variable for each species that represents its relative dimerization potential. The bacteria model is based on a curve fit of the dimerization potential to the ultraviolet rate constant data for 92 bacteria species represented by 216 data sets from published studies. There were 4 outliers excluded from the model resulting in a 98% Confidence Interval. The curve fit resulted in a Pearson correlation coefficient of 80%. All identifiable bacteria important to human health, including zoonotic bacteria, were included in the database and predictions of ultraviolet rate constants were made based on their specific genomes. This database is provided to assist healthcare personnel and researchers in the event of outbreaks of bacteria for which the ultraviolet susceptibility is untested and where it may be hazardous to assess due to virulence. Rapid sequencing of the complete genome of any emerging pathogen will now allow its ultraviolet susceptibility to be estimated with equal rapidity. Researchers are invited to challenge these predictions.
Importance
This research demonstrates the feasibility of using the complete genomes of bacteria to determine their susceptibility to ultraviolet light. Ultraviolet rate constants can now be estimated in advance of any laboratory test. The genomic methods developed herein allow for the assembly of a complete database of ultraviolet susceptibilities of pathogenic bacteria without resorting to laboratory tests. This UV rate constant information can be used to size effective ultraviolet disinfection systems for any specific bacterial pathogen when it becomes a problem.
Introduction
In spite of the widespread application of ultraviolet disinfection technologies in healthcare and other industries, less than 25% of pathogenic bacteria have been studied in laboratories to assess their UV rate constant. A complete database of bacteria UV rate constants would be useful for estimating the required UV dosage for bacteria when using UV disinfection equipment. Laboratory testing can be expensive, time-consuming, and even hazardous but the use of genomic modeling can enable rapid estimation of required UV dosage for any new or emerging pathogen once the genome is sequenced. Genomic predictions of UV rate constants can potentially be extremely accurate, more accurate than typical laboratory test results. The accuracy achievable by genomic modeling may be such that predictions could serve as a standard by which to judge the accuracy of laboratory testing, in much the same way as laws of physics are used to assess the efficiency of motors and the strength of structures. The present research is based on an empirical model that is fit to existing data and represents a step towards a fundamentals-based model of ultraviolet susceptibility that promises accuracy beyond six decimal places. The present research addresses bacteria because they have long been the largest and most persistent infection problem in healthcare. Although this view is being challenged by the current pandemic, the success of this bacteria genomic model will also lead to a model for predicting UV rate constants for viruses as well.
Methods
The methods used to develop the model and make predictions are purely analytical and empirical and involved firstly a compilation of UV rate constant studies for bacteria and computation of the UV rate constants of k values. The measured UV rate constants are averaged when there are multiple studies and are applied in an empirical model that correlates the k values with a theoretical dimerization value, Dv, which quantifies a series of genomic parameters into a single value. The dimerization value is computed directly from the genomes for all subsequent bacteria to produce the predicted UV rate constant. The details of these computations are discussed in the following subsections.
Rate Constant Studies
Studies on ultraviolet (254 nm) survival of bacteria on surfaces or in water were reviewed. Criteria for selection included having interpretable results that would define the first stage of exponential decay independent of any survival tail (the second stage) or shoulder in the survival curve, since the presence of a tail can confuse interpretation of the first stage rate constant. Table 1 summarizes the 92 bacteria that form the basis of the model and the UV rate constants averaged from the indicated studies. Every attempt was made to be inclusive of all available studies but invariably there were a few studies that defied interpretation or for which a genome was lacking. Shown in the columns of Table 1 are the UV rate constant k computed as the average of all studies, the Mean Aerodynamic Diameter (MAD), the D90 values, and the genome identifier or accession number. In some cases the identifier is the first in a series of shotgun sequence identifiers. The number of data sets reflects the number of studies up to a maximum of 3 beyond which no further weight is given to the species in the model. This avoids bias for the small number of species that have a large number of studies and 3 is an appropriate number of studies from which to obtain an average value. In cases where an excessive number of studies exist many were omitted due to redundancy. Four species formed outliers in this model and were excluded. Nocardia asteroides was also a low outlier and it is a bacteria that more resembles a fungus and may not belong to this data set, but more studies on Nocardia species are needed to resolve this issue. Enterobacter cloacae was a low outlier. Mycobacterium abscessus, Burkholderia mallei, and B. pseudomallei were all high outliers. These outliers were each single studies except for B. pseudomallei which had a second study which was not an outlier and which is included in the model in Table 1. Micrococcus radiodurans was expected to be a low outlier due to its unique ability to resist radiation and was excluded from consideration altogether except that it is shown for reference.
The Average UV Rate Constant k shown in Table 1 is computed from information provided in the referenced studies. If only a D90 value is provided and there is no significant shoulder or tail evident then the value is converted to a rate constant according to equation (1). In most cases the original data is inspected to determine the UV rate constant separate from any shoulder or tail. The range of data points within which the value of k1 can be accurately computed is illustrated in Figure 2. Any points above and beyond the shoulder region, if any, and any points beyond the survival tail, if any, are suitable for use in computing the UV rate constant. In all case there are multiple studies available, the rate constant is computed as the average and not the D90 value.
The Mean Diameter listed in Table 1 represents the logmean diameter of spherical cells and for rods or oblong cells it represents the logmean diameter computed with the Cauchy Theorem, which represents the average area of the profile of the rod. The profile area accounts for the area upon which the UV rays impinge and this determines the amount of UV scattering (Kowalski 2020). The mean diameter is used to compute the shielding constant Sc, which represents the amount of shielding a cell has from UV due to its physical size. The shielding constant is used in turn to compute the intrinsic UV rate constant which is the value correlated by the empirical model. The advantage of using the intrinsic rate constant, however theoretical it may be, is seen in the fact that the MAD exponent in this model converged to zero and ceased to have any predictive value. This simplification separates the only physical parameter (MAD) from the genomic parameters and renders the core model purely genomic in nature.
Calculations of the UV rate constants from the studies in Table 1 will be made available as supplemental files online at XXXX.later.
The last five microbes shown in gray cells are the four outliers and D. radiodurans.
The Genomic Parameters
Theoretically, the exact sequence of adjacent base pairs are the primary determinants of the intrinsic UV sensitivity and the frequency of dimer formation is in the order TT>TC>CT>CC (Setlow 1966, Becker 1989, Khoe 2018). Specifically, adjacent pyrimidines have the potential to form dimers, and some purines flanked by pairs of pyrimidines may also form dimers (Pan 2011). Pyrimidine doublets, or pairs of pyrimidines, may have dimer formation suppressed when they are flanked by purines (Law 2013). Table 2 summarizes the genomic motifs that are likely to dimerize and those that are not.
There are five groups of potential dimers, TT, TC, CT, CC, and YR (where R represents purines and Y represents pyrimidines). The sites are counted inclusively, wherein, for example, the sequence TT counts as 2 dimer sites and the sequence CTC counts as 3 dimer sites. Where sequences overlap, dimers are categorized based on the priority TT>TC>CT>CC>YR.
When multiple pyrimidines occur sequentially they may create a hot spot where the likelihood of dimerization is increased by a factor of up to 2 (Becker 1989). This phenomenon is known as hyperchromicity. Not enough is known about the hyperchromic effect to quantify it precisely but a parameter is included to add to the sum of dimers of any doublet or triplet whenever 3 or more pyrimidines are found in sequence (Kowalski et al 2009). This factor, H, increases the probability of dimerization for doublets and triplets based on how many adjacent pyrimidines are present in the genome, up to a value of 8 in a row where the effect levels off. The value of H may typically double the value of an isolated doublet, based on studies by Becker and Wang (1989).
The value of TTh represents the additional effect from a local hotspot, a value from 0 – 1.0 which represents the increase in dimerization probability due to being part of a cluster. The increased photoreactivity is based on an assumed sigmoid function relating the hyperchromicity factors (TTh, etc.) to the number of pyrimidines in sequence, as shown in Table 3.
It was found in the course of this research that the hyperprimers (labeled TTh, TCh, CTh, CCh & YRh) were better predictors than the specific sequences of base pairs (TT, TC, CT, CC & YR). This innovation marks a departure from previous genomic models by the author (Kowalski 2009, 2009a, 2009b, 2009c). Since the hyperprimer values represent hot spots where dimerization is amplified, this approach is called a Hot Spot model. The values of the hyperprimers were simply substituted in place of the dimer counts in the equation for computing the Dimerization value.
The Dimerization Value
The nine genomic parameters that proved most significant are identified in Table 3. The Relative Dimerization Value Dv is computed by multiplying all parameters together, with each having an adjustable exponent, as follows: where a thru h are adjustable exponents
The exponents were adjusted through multiple iterations to maximize the least squares curve fit of ki to Dv. If the exponent of a parameter is zero, then the dimer contribution becomes unity, and the factor makes no contribution to the model. All parameters except bp are fractions since they are individually normalized as a fraction of the total bp. Since all the variables in equation (2) are dimensionless except bp, the dimerization value Dv has the units of bp.
Table 4 summarizes the final iterated values of the genomic parameter exponents after curve-fitting to maximize the Pearson Correlation coefficient.
In this model, as opposed to previous models by the author, the intrinsic UV rate constant is used as the predictor (Kowalski 2020). The intrinsic UV rate constant is theorized as a parameter that depends solely on the genomic sequences and not on any physical parameter. The physical size of a microorganism plays a role in determining its ultraviolet sensitivity via Mie scattering effects. The intrinsic UV rate constant is defined as: Where k1 = first stage UV rate constant, m2/J
Sc = shielding constant
The shielding constant is computed as the complement of the ratio of the scattering cross-section to the extinction cross-section per Kowalski (2020): Computations of the shielding constants are provided in the supplementary materials. UV rate constants from Table 1 are converted to intrinsic rate constants for use in the genomic model. The resulting predictions are then converted back to first stage rate constants in the final summary.
The Empirical Model
The final empirical model correlating measured UV rate constants with those predicted by the model is summarized by the parameters in Table 3 and plotted in Figure 2.
The complete equation for the intrinsic rate constant ki can be written as follows: Equations (2) and (3) can then be combined into a complete closed form equation to determine ki.
The Bacteria Database
Table 5 presents a summary of all human pathogens and zoonotic pathogens that could be identified that have a published genome. The Type of microbe (H=Human, Z = Zoonotic) is specified for each pathogen. The “Genome RefSeq/NCBI#” shows either the complete genome number or in some cases it shows the first chromosome or segment identification number from a shotgun sequence. The UV k is given along with the D90 range (for a 98% Confidence Interval) and the Mean D90 is computed based on the UV k. No account is taken for shoulders in this summary and therefore the UV k value and the D90 value represent the absence of any shoulder.
The genomic parameters used in the model were further studies to evaluate their relative contribution to the model. The contributions were analyzed as stand-alone correlations and the results are shown in Figure 3. Only the Hot Spots were used for this model and no correlations can be made for the individual dimers. The results indicate that the combined Hot Spots (TTh+TCh+CTh+CCh+YRh) were the single strongest predictor with a 70% Pearson correlation coefficient.
A graph of the Hot Spots was made and the results are presented as a circular genome for the bacteria Borrelia afzeli in Figure 4. In this graph only the highest hyperprimer values for the first 4000 bp are shown for simplicity. Figure 4 shows the pyrimidine hot spots in red and the purine hot spots in yellow. It is worth noting that of the hundreds of indicated Hot Spots only a few would be necessary to inactivate the pathogen. The purine hot spots are significant because photoreactivation enzymes may repair pyrimidine dimers but they cannot repair purine dimers. Purine dimers may therefore be more fatal to the pathogen than pyrimidine dimers.
Discussion
The present genomic model improves upon the results presented previously by the author (Kowalski 2009, 2009a, 2009b, 2009c). The current Hot Spot bacteria model incorporates many more studies than previous models and more scrutiny was placed on calculating the UV rate constants but the results are not drastically different from previous attempts. While there are clearly a great many variations possible in the modeling approach and in the selection and adjustment of parameters, the final predictions of UV rate constants have changed little, suggesting they limited sensitivity to the approach.
Other attempts to develop genomic models for viruses have been published and these have employed much the same genomic motifs as specified in the current model and have also produced similar predictions of UV rate constants (or D90 values) as the author’s virus genomic models (Rockey 2021, Pendyala 2020). Although a fundamentals-based genomic model has been elusive, it appears that empirically based genomic models are adequate to generate predictions of UV sensitivity for all types of pathogens. Genomic modeling has also been used by the author to successfully predict the UV sensitivity of individual antibiotic resistant genes (Destiani 2017).
We invite researchers to challenge these predictions, especially those bacteria that have no prior studies, and also for the four outlying bacteria. The predicted values for UV rate constants can act as a target for the experimental design and should enable researchers to seek more accurate results. Because of the fact that the UV rate constant must be absolutely defined by the species genome, the predictions provided here may act as a standard by which to gauge results from future experiments.
It was clear from reviewing all studies that there was a lack of consistency in the test methods and this surely resulted in scattering of the results and reduced predictive performance of the model. Problems were observed in many studies related to how the UV dose was established. Researchers should take care to use photosensors properly. In general, if a collimated beam apparatus is not used, the distance from the photosensor to the lamp should be on the order of one lamp length. In addition, many studies used low levels of UV irradiance and this produced a shoulder in the survival curve. The shoulder will distort any D90 value estimated from the data. It is suggested that researchers use irradiance levels of at least 1 W/m2 (100 µW/cm2) to avoid shoulders in the results and allow more accurate determination of the first stage rate constant.
A standardized test for UV rate constant evaluation is needed. Genomic modeling makes it possible to compute UV rate constants to six decimal places for bacteria, which have genomes on the order of 1M bp. A standard test capable of resolving UV rate constants to six decimal places, which represents six logs of reduction, may require the use of bacteria monolayers to avoid tailing in the survival curve, which reduces measurement accuracy.
The current method makes it possible to quickly evaluate the UV susceptibility of any emerging pathogen once the genome is published. The calculations could conceivably be incorporated into the NCBI Genome database to provide instant results once a genome is sequenced. Future research will address improved genomic modeling of viruses and fungi and investigation will be conducted into whether the Hot Spot model may have applications in predicting susceptibility to skin cancer.