Integrative Metabolomic and Proteomic Signatures Define Clinical Outcomes in Severe COVID-19
1Department of Physiology and Biophysics, Weill Cornell Medicine, New York, NY, USA
2Meyer Cancer Center and Caryl and Israel Englander Institute for Precision Medicine, Weill Cornell Medicine, New York, NY, USA
3Department of Medicine, Division of Pulmonary and Critical Care Medicine, Weill Cornell Medicine, New York, NY, USA
4Proteomics Core, Weill Cornell Medicine – Qatar, Doha, Qatar
5Department of Population Health Sciences, Division of Biostatistics, Weill Cornell Medicine, New York, NY, USA
6Proteomics and Metabolomics Core Facility, Weill Cornell Medicine, New York, NY, USA
7Department of Physiology and Biophysics, Weill Cornell Medicine – Qatar, Education City, 24144 Doha, Qatar
8Department of Medicine, Division of General Internal Medicine, Weill Cornell Medicine, New York, NY, USA
9Department of Pathology and Laboratory Medicine, Weill Cornell Medicine, New York, NY, USA
10Division of Nephrology and Hypertension, Joan and Sanford I. Weill Department of Medicine, New York, NY, USA
#corresponding authors Correspondence to: Soo Jung Cho, sjc9006@med.cornell.edu Jan Krumsiek, jak2042@med.cornell.eduAbstract
The novel coronavirus disease-19 (COVID-19) pandemic caused by SARS-CoV-2 has ravaged global healthcare with previously unseen levels of morbidity and mortality. To date, methods to predict the clinical course, which ranges from the asymptomatic carrier to the critically ill patient in devastating multi-system organ failure, have yet to be identified. In this study, we performed large-scale integrative multi-omics analyses of serum obtained from COVID-19 patients with the goal of uncovering novel pathogenic complexities of this disease and identifying molecular signatures that predict clinical outcomes. We assembled a novel network of protein-metabolite interactions in COVID-19 patients through targeted metabolomic and proteomic profiling of serum samples in 330 COVID-19 patients compared to 97 non-COVID, hospitalized controls. Our network identified distinct protein-metabolite cross talk related to immune modulation, energy and nucleotide metabolism, vascular homeostasis, and collagen catabolism. Additionally, our data linked multiple proteins and metabolites to clinical indices associated with long-term mortality and morbidity, such as acute kidney injury. Finally, we developed a novel composite outcome measure for COVID-19 disease severity and created a clinical prediction model based on the metabolomics data. The model predicts severe disease with a concordance index of around 0.69, and furthermore shows high predictive power of 0.83-0.93 in two previously published, independent datasets.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
JK is supported by the National Institute of Aging of the National Institutes of Health under award 1U19AG063744. SJC is supported by National Heart, Lung and Blood Institute under award K08HL138285. KS is supported by 'Biomedical Research Program' funds at Weill Cornell Medical College in Qatar, a program funded by the Qatar Foundation and multiple grants from the Qatar National Research Fund (QNRF).
Introduction
The novel coronavirus disease 2019 (COVID-19) has a broad spectrum of clinical features that range from asymptomatic disease to acute respiratory distress syndrome (ARDS)(1, 2). COVID-19 ARDS can lead to refractory hypoxia, mechanical ventilation, prolonged intensive care unit (ICU) stay and increased mortality(3). Previous studies have shown a high incidence of concomitant organ failure in COVID-19, including acute kidney injury (AKI)(4), acute liver injury(5), thromboembolic events(6, 7) and secondary infections contributing to a fatal outcome(8).
Massive investigative efforts by multiple scientific groups have used proteomic and metabolomic approaches to begin to unravel disease mechanisms relevant to SARS-CoV-2 infection such as inflammation, coagulation, and metabolism(9). However, how COVID-19 specific protein-metabolite interactions relate to the severity of disease and clinical outcomes remains poorly understood. Key study limitations have included relatively small sample sizes, absence of protein-metabolite network analysis and the focus on dichotomous outcome measures such as death and survival. These limitations have been difficult to overcome and restrict our understanding of COVID-19 pathogenesis.
Here, we report the largest study to integrate targeted metabolomic and proteomic analyses of serum samples obtained from hospitalized COVID-19 patients during SARS-CoV-2 infection compared to patients admitted during the same time period with symptoms related to COVID-19 and negative RT-PCR for SARS-CoV-2 as controls (Figure 1). Through this work, we uncovered COVID-19-specific metabolite and protein profiles, identified novel protein-metabolite modules, and defined the molecular signatures of several clinical indices (CRP, ferritin, platelet count, AKI, and death). Additionally, we developed the first clinical composite outcome prediction model in COVID-19, where the input of discrete metabolic profiles significantly improves our clinical insight into a broad range of outcomes known to plague many survivors of severe COVID-19.
Results
Study cohort and dataset
Our cohort was comprised of 330 patients with confirmed SARS-CoV-2 RT-PCR, and 97 non-COVID-19 controls with negative RT-PCR results who were hospitalized at the NewYork-Presbyterian Hospital/Weill Cornell Medical Center between March and April 2020. Serum samples were obtained within the first 3 days of admission. The majority of COVID-19 patients had samples drawn at two or three different time points resulting in a total of 582 serum samples from the 330 COVID-19 patients, while all 97 controls had one sample drawn. Metabolomics was measured for all available samples. Proteomics was measured for fewer samples (n=189), also across different time points for some patients. Notably, there were only minor time effects across the three days (Supplementary Figure 1), and the repeated samples were thus treated as replicates using a linear mixed effect model, see Methods. A detailed description of the clinical and demographic characteristics of the cohort can be found in Table 1, Supplementary Table 1 and Supplementary Table 2. Of note, we excluded samples collected after intubation because we found that the clinical act of intubation significantly alters a patient’s metabolic profile (Supplementary Table 3).
Metabolic profiles were assessed for all samples using liquid chromatography coupled with mass spectrometry (LC/MS). After quality control and data preprocessing, 125 metabolites were available for comparative analysis. Targeted proteomic profiling was performed on a subset of 227 samples (173 from COVID-19 patients and 54 controls) using the Olink inflammation, cardiovascular II and cardiovascular III panels, which cover 266 unique protein biomarkers. These panels were selected since it has previously been shown that inflammation and cardiovascular pathways are essential during COVID-19 pathogenesis(10).
Conclusion
The novel coronavirus has ravaged the global healthcare system due to its high transmissibility and unpredictable clinical course that often affects multiple organ systems. Moreover, the long-term consequences of COVID-19 infection remain poorly understood. A full understanding of the pathogenesis of COVID-19 will require an unraveling of the mechanisms of inflammation, immune dysfunction, endothelial cell injury and dysregulated coagulation that underlie this disease.
In our study, we used an integrative proteomic-metabolomic analysis to identify global molecular signatures specific to the acute illness of COVID-19, while many prior metabolomic and proteomic studies have not assessed the interplay between proteins and metabolites. Our analyses establish associations of specific inflammation and vascular injury-related proteins with various metabolites during COVID-19, which appear to link inflammation with mitochondria-dependent energy metabolism and viral replication, as well as coagulation with fibrogenesis and glycolysis.
Our discovered network modules not only provide a better understanding of disease pathogenesis, but also facilitate novel potential therapeutic targets for COVID-19. The modules identified various proteins and metabolites involved in inflammatory and vascular injury processes, such as MMP 12, Cathepsin D and RAGE which, to the best of our knowledge, have not yet been studied as targets for therapeutic intervention in COVID-19. Of note, our modules contained IL-6, which is already a mainstay of treatment for severe disease(50), and several other molecules such as carnitine, niacinamide and IFN-gamma which others have been studying in the context of COVID-19 therapies(51-53).
There is increasing evidence that evaluating symptoms and multiple clinical outcomes during acute disease is crucial in determining the risk of long COVID-19(54). To the best of our knowledge, we are the first group to develop a composite outcome measure in COVID-19 using multiple clinical indices in a prediction model that assesses not only COVID-19 disease severity but also the sequelae of COVID-19 that characterize post-acute COVID-19 syndrome (PACS). Compared to dichotomous outcome measures such as death and survival, our composite outcome score reflects a broader, more holistic assessment of COVID-19 morbidity in the hospital setting. Moreover, we were able to validate the model in two independent studies, thereby demonstrating its generalizability and translational potential.
Several key strengths underlie our study cohort. As opposed to the use of healthy controls reported in other COVID-19 studies(9, 20, 43, 44, 55), our use of non-COVID patient samples, in the same hospital during the same period between March and April 2020, allowed us to investigate the interactions highly relevant to COVID-19 pathogenesis and clinical course. Additionally, we analyzed a relatively larger cohort compared to other studies, with hundreds of samples available for both metabolomic and proteomic analysis.
Our study has several limitations. As alterations in proteome and metabolome were analyzed in sera but not in lung tissues or bronchoalveolar lavage fluid, our results may not reflect what occurs at tissue-specific cellular levels. Furthermore, based on the current study design and methodology, the correlative relationships we report between metabolomic, and proteomic alterations and SARS-CoV-2 outcomes should be interpreted as purely correlative rather than causal in nature. Additional studies are required to define the mechanistic roles of individual molecules highlighted in this paper. Finally, as our study was only a single center investigation, our results will need to be validated in other cohorts.
In conclusion, our investigation has sought to not only define the metabolomic and proteomic signatures of COVID-19, but also to explore interactions between metabolites and proteins that can serve as a roadmap for future mechanistic studies. We have furthermore proposed a novel clinical composite outcome score that can be used in a clinical prediction model for COVID-19. Ultimately, a better understanding of the pathophysiology of COVID-19 at the molecular level may lead to short-term and long-term targeted therapies.
Methods
Cohort description
This is a single-center prospective analysis of one cohort comparing hospitalized COVID-19 patients and non-COVID-19 controls. Our cohort was comprised of 330 patients with confirmed SARS-CoV-2 RT-PCR, and 97 non-COVID-19 controls with negative RT-PCR results who were hospitalized at the NewYork-Presbyterian Hospital/Weill Cornell Medical Center between March and April 2020. This study has been approved by the Weill Cornell Medicine (WCM) IRB with protocol #19-10020914. Remnant serum samples were matched with selected patients after which patients were deidentified. Controls were randomly selected patients admitted to the hospital with symptoms suspicious for COVID-19, but with negative SARS-CoV-2 RT-PCR. Seventy nine percent of control group patients had shortness of breath, fever, cough or chest pain which are commonly seen in COVID-19. Children (less than 18 years old) and pregnant women (confirmed by a positive beta-HCG test and/or medical records) were excluded.
Sample handling
Standard practices for serum collection and storage at the NewYork-Presbyterian/Weill Cornell Medical College include collecting venous blood into a serum-separating tube (SST), and serum is obtained by centrifuging at 1,500g for 7 minutes as soon as possible with a maximum time limit of 2 hours from the time of collection. The specimens are typically stored at 4°C for 1 to 5 days before coded/de-identified and then transferred into a -80°C freezer. Samples were thawed and inactivated in different ways: for the metabolomic analysis, x3 sample volume of HPLC grade ethanol were added; for the proteomic analysis, the samples were heat-inactivated in a water bath of 56°C for 15 minutes. After these processes, the samples were stored again at -80°C until the analyses were performed.
Data Collection
Data were obtained from the Weill Cornell Medicine COVID Institutional Data Repository (COVID-IDR), which is a high-quality registry of COVID-19 patients at NewYork-Presbyterian - Cornell with laboratory confirmed SARS-CoV-2 RT-PCR. The COVID-IDR houses both manually and automatically extracted Electronic Health Record (EHR) data. Demographics, comorbidities, and important dates of patients’ hospital course (admission, intubation, extubation, discharge, death) were extracted by a team of medical professionals and stored in the COVID-IDR. Laboratory tests, ventilation parameters, vital signs, and respiratory variables were additionally available via automated extraction through the Weill Cornell-Critical Care Database for Advanced Research (WC-CEDAR) within the COVID-IDR. WC-CEDAR(56) is a critical care database originally designed to automatically extract, transform, and store EHR data on Intensive Care Unit (ICU) patients; it was expanded to include all hospitalized patients during New York City’s COVID-19 surge. Data not available within WC-CEDAR were manually extracted and recorded in REDCap.
Proteomic profiling
Proteomics analysis was performed using the Olink platform (Uppsala, Sweden) at the Proteomics Core of Weill Cornell Medicine-Qatar, according to manufacturer’s instructions. We used the Inflammation, Cardiovascular II and Cardiovascular III panels. High throughput real-time PCR of reporter DNA lined to protein specific antibodies was performed on a 96-well integrated fluidic circuits chip (Fluidigm, San Francisco, CA). Each sample was spiked with quality controls to monitor the incubation, extension, and detection steps of the assay. Additionally, samples representing external, negative, and inter-plate controls were included in each analysis run. From raw data, real time PCR cycle threshold (Ct) values were extracted using Fluidigm reverse transcription polymerase chain reaction (RT-PCR) analysis software at a quality threshold of 0.5 and linear baseline correction. Ct values were further processed using the Olink NPX manager software (Olink, Uppsala, Sweden). Here, log2-transformed Ct values from each sample and analyte were normalized based on spiked-in extension controls and scale-inverted to obtain Normalized log2-scaled Protein Expression (NPX) values. NPX values were adjusted based on the median of inter plate controls (IPC) for each protein and intensity median scaled between all samples and plates.
Each metabolite and protein was annotated with pathways from the Kyoto Encyclopedia of Genes and Genomes (KEGG) database(59).
Data preprocessing
Children, pregnant women, and samples after intubation were excluded from all analyses. The metabolomics data was measured in three different batches. For each batch, data was preprocessed by filtering out samples with more than 50% missing values, followed by filtering out metabolites with more than 25% missing values, probabilistic quotient normalization(60), and log2 transformation. Two extreme outlier metabolites were manually removed (phosphorylcholine and adenosine monophosphate). The next step was to merge the different batches into a joint dataset. Batch 3 contained only control patients and could thus not be simply added by batch correction. To avoid issues created by this imbalanced experimental design, batches 2 and 3 contained an overlapping set of samples which were used for an anchor-based normalization by dividing each metabolite in batch 3 by the mean fold change of the overlapping samples. The anchor samples from batch 3 were then deleted and the batches combined using median-based batch correction(61). Overall, this procedure eliminates batch effects and allows for a batch that only contains control samples. Missing values were then imputed using the k-nearest neighbor approach(61).
Proteomics data preprocessing included the same steps of filtering, quotient normalization, logging, and missing value imputation with identical parameters as for the metabolomics data. Ten proteins were measured as duplicates on the Olink platform, so their expression values were averaged.
Statistical analysis
Differential expression of metabolites and proteins for both the COVID-19 vs. control analysis as well as the clinical parameter analysis within the COVID-19 cohort was assessed using the following linear mixed effect model: where met_prot is each individual metabolite or protein, outcome is either COVID-19 yes/no or the value of a clinical parameter, time is the day of sample taking as a factor, and (1|patient) is a random effect per patient to account for repeated measurements. P-values are reported for the significance of the outcome term.
Data preprocessing and statistical analysis was performed using the “maplet” toolbox for R(62) (https://github.com/krumsieklab/maplet).
Network construction
The dataset was first reduced to the samples that overlap between metabolomics and proteomics (n=227), and corrected for age, sex, BMI, and COVID-19 status (yes/no). A Gaussian Graphical Model (GGM) based network was then constructed using the GeneNet algorithm(63) and drawing an edge for all partial correlations with an FDR smaller than 0.2. In a second step, this network was condensed to highlight the connections between molecules that were significantly different between COVID-19 and controls. To this end, a shortest-path distance matrix between all molecules was constructed and subset to the significant molecules. A minimum spanning tree(64) of this matrix was then constructed to visualize a simplified network.
Composite Outcome
A detailed description of the construction of the composite outcome along with patient numbers in each group can be found in Supplementary Figure 3.
Regularized linear mixed effect ordinal regression model
A new regression model was developed to deal with an ordinal outcome, repeated measurements, and feature selection with regularization. Repeated measurements are handled as a random effect, while age, sex, BMI, and metabolites are treated as fixed effects. The metabolites are penalized using an L1 LASSO-type regularization to obtain a sparse solution. The model is fitted using an mixed-effect ordinal regression model with complementary log-log link function(65), using maximum likelihood (ML) estimation as proposed by Ripatti and Pamgren(66). An optimal LASSO penalty parameter was estimated through an iterative algorithm for maximizing the Laplace approximation of the integrated model likelihood. This approach was adapted from Therneau(67), which was originally developed for mixed effect Cox models. To obtain an unbiased estimate of the model performance, leave-one-out-cross-validation across the entire dataset was performed. The added value of metabolomics data over baseline clinical data was assessed by comparing the final model with a model only consisting of age, sex, and BMI.
Validation datasets
Metabolomics data were downloaded from Su et al.(43) (n=121) and Shen et al. (44) (containing two validation sets, n=10 and n=19). The Su dataset contained all six metabolites from our reduced model as well as age, sex and BMI as baseline parameters. The Shen study also covered age, sex and BMI, and the first test dataset (“C2”) contained all six metabolites. The second test dataset (“C3”) was measured using a targeted assay of only 7 metabolites and 22 proteins, and the metabolites did not overlap with our model metabolites. Thus, we had to follow a more complex procedure for validation. We first applied our risk score in their training cohort, which contained all metabolites. In the training cohort, this score was then regressed on the available measurements in C3, i.e. modeling ‘score ∼ metabolite1 + … + metabolite7 + protein1 + …+ protein22’. The coefficients from this model were then used in C3 to derive a surrogate severity score, which we evaluated in Figure 5.
Supporting information
Data Availability
There are legal and ethical restrictions on data sharing because the Institutional Review Board of Weill Cornell Medicine did not approve public data deposition. The data set used for this study constitutes sensitive patient information extracted from the electronic health record. Accordingly, it is subject to federal legislation that limits our ability to disclose it to the public, even after it has been subjected to deidentification techniques. To request the access of the de-identified minimal dataset underlying these findings, interested and qualified researchers should contact Information Technologies & Services Department of Weill Cornell Medicine support@med.cornell.edu.
Data Availability
The data used in this study can be downloaded at https://doi.org/10.6084/m9.figshare.19115972.v1.
Code Availability
Code to reproduce all the statistical results presented in this paper is available at https://github.com/krumsieklab/covid-omics.
Funding
JK is supported by the National Institute of Aging of the National Institutes of Health under award 1U19AG063744. SJC is supported by National Heart, Lung and Blood Institute under award K08HL138285. KS is supported by ‘Biomedical Research Program’ funds at Weill Cornell Medical College in Qatar, a program funded by the Qatar Foundation and multiple grants from the Qatar National Research Fund (QNRF).