Clinical Study Protocol of the ‘Biomarkers of Severity of COVID-19 Patients’ (BIOMARCOVID) Project
1Univ. Grenoble Alpes, CNRS, UMR 5525, VetAgro Sup, Grenoble INP, CHU Grenoble Alpes, TIMC, 38000 Grenoble, France
2Univ. Grenoble Alpes, Data Engineering Unit, Public Health Department, Grenoble Alpes University Hospital, 38000 Grenoble, France
3Univ. Grenoble Alpes, Inserm, CHU Grenoble Alpes, CIC, 38000 Grenoble, France
4Université Clermont Auvergne, INRAE, UNH, Plateforme d’Exploration du Métabolisme, MetaboHUB Clermont, Clermont-Ferrand, France
5Plateforme de Cytométrie, CHU Grenoble Alpes, 38000 Grenoble, France
6Université Paris Saclay, CEA, INRAE, Médicaments et Technologies pour la Santé (MTS), MetaboHUB, Gif-sur-Yvette, France
7MetaToul-Lipidomic MetaboHUB Core Facility, Inserm U1048, Toulouse Cedex 4, France
8Univ. Grenoble Alpes, Département des Maladies Infectieuses, CHU Grenoble Alpes, 38000 Grenoble, France
9Centre International de Recherche en Infectiologie (CIRI), Équipe VirPath, INSERM U1111, CNRS UMR 5308, ENS Lyon, Université Claude Bernard Lyon 1, Lyon, France
#Correspondence: Dr Audrey Le Gouellec; alegouellec@chu-grenoble.fr;, +33(0)4.76.76.63.76, Prof. Olivier Epaulard; oepaulard@chu-grenoble.frABSTRACT
Introduction
The coronavirus disease 2019 (COVID-19) pandemic has challenged health care systems worldwide, in certain areas exceeding hospital capacities and human resources. This has underscored the importance of having better tools to predict the outcome of potentially severe respiratory infections such as SARS-CoV-2. Predicting COVID-19 severity may allow physicians to better manage ICU beds and increase the chances of patient survival through appropriate management. During the toughest months of the pandemic, most physicians tried to identify patients that might develop severe forms based primarily on clinical features on admission (e.g., BMI, age). In this context, significant research has focused on identifying comorbidities, clinical manifestations, and routine blood biomarkers to predict disease severity. However, despite the demonstrated value of untargeted metabolomics in assessing severity, limited data exist on its use for identifying novel metabolite biomarkers that could improve both the sensitivity and specificity of outcome prediction. Our goal is to identify metabolite biomarkers that could enhance the predictive accuracy of standard medical biology data and clinical parameters.
Methods and analysis
This is a retrospective, observational, monocentric cohort study conducted at the Centre Hospitalier Universitaire Grenoble Alpes (CHUGA). The maximum number of eligible patients admitted for PCR-confirmed COVID-19 between March and December 2020 will be included. Severity outcome is defined using the WHO 10-category ordinal scale (mild: categories 4–5; severe: >5). Blood samples were collected within 48 hours of admission and analyzed for 62 routine blood tests and untargeted multiplatform LC-MS/MS metabolomics across four national platforms. Statistical analysis will include logistic regression with variable selection for the primary aim, and multi-block chemometric integration of clinical, biological, and metabolomics data as a secondary aim.
Ethics and dissemination
A study steering committee has been formed to ensure the accuracy of the collected data by thoroughly reviewing it prior to the data lock. All aspects of the study comply with ethical standards, including approval by the CHUGA institutional review board and adherence to CNIL Reference Methodology MR004 for the protection of participants’ rights, privacy, and confidentiality. This study is registered on the French Health Data Hub (number F20210218154851). Results will be disseminated through peer-reviewed publications, presentations at national and international scientific and clinical conferences, and reports shared with key healthcare system stakeholders.
ARTICLE SUMMARY
Strengths and limitations of this study
- Blood samples were collected within 48 hours of admission and before any severe symptom onset, aliquoted within 4 hours of collection and stored at −80°C, ensuring high pre-analytical quality.
- This study uses a multiplatform untargeted LC-MS/MS metabolomics approach across four complementary analytical platforms (CEA Saclay, INRAE Clermont, MetaToul, GEMELI), providing broad metabolome coverage.
- The integration of three heterogeneous data blocks (Metabolome, Biologicome, Clinicome) via multi-block chemometrics enables a systems-level view of COVID-19 severity.
- The sample size is limited by the complexity and cost of metabolomics analyses, and no formal sample size calculation was performed; findings should be considered exploratory and hypothesis-generating.
- As a monocentric, retrospective study conducted during a single epidemic wave (March–December 2020), generalizability may be limited by incomplete data across some routine blood tests and by differences in settings, subsequent variants, or treated populations.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Clinical Protocols
INTRODUCTION
The COVID-19 pandemic profoundly disrupted hospital operations, placing physicians worldwide in unprecedented circumstances. Patient stratification for those hospitalized with COVID-19 can be accomplished through clinical characteristics, 1·2 clinical symptomatology, 3 and routine blood biomarkers analyzed in medical biology laboratories. □However, the ability to predict disease severity using such data has proven to lack both specificity and sensitivity. This underscores the importance of ongoing research to identify more effective biomarkers for predicting disease progression.
At the Centre Hospitalier Universitaire Grenoble Alpes (CHUGA) in March 2020, when admitting patients without severe symptoms, predictions were primarily based on clinical features and emerging evidence from the literature, including age, sex, comorbidities (e.g., arterial hypertension, body mass index (BMI)1), and clinical symptoms.3 Nevertheless, due to the lack of sensitivity and specificity of these factors, an important “grey area” remains: physically fit and relatively young patients without known pre-existing conditions may develop severe COVID-19. Validated prognostic biomarkers include C-reactive protein (CRP) and IL-6,□as well as platelet count, N-terminal prohormone of brain natriuretic peptide (NT-proBNP), troponin I, lactate dehydrogenase (LDH), aspartate transaminase (AST), alanine transaminase (ALT), albumin, creatinine, urea, creatine kinase (CK), and white blood cell count (WBC).□□1□ However, the prognostic value of several other parameters remains insufficiently established, including agranular neutrophils (AGRAN), hemoglobin, red blood cell count, cholesterol, triglycerides, high-density lipoprotein (HDL), alkaline phosphatase (ALP), sodium, potassium, chloride, pH, red cell distribution width, gamma-glutamyl transferase (GGT), mean corpuscular hemoglobin concentration (MCHC), calcitriol (25-OHD), and total proteins.
Untargeted metabolomics offers significant advantages for biomarker discovery. It enables a comprehensive, hypothesis-free analysis of all metabolites detectable in biological samples, 11 making it well suited for identifying novel markers not captured by targeted methods. To date, untargeted metabolomics has been used primarily to study metabolic perturbations induced by SARS-CoV-2 infection, 12□1□ or to explore biomarkers of COVID-19 severity, 1□ with promising but preliminary results. Adding metabolite biomarkers identified by untargeted metabolomics to standard prognostic tests has the potential to improve accuracy, sensitivity, and specificity, and to reveal specific metabolic pathways involved in the response to infection.
This project is one of the few to combine prognostic biomarker discovery via untargeted metabolomics with routine blood test data in hospitalized COVID-19 patients and aims to provide a methodology adaptable to future emerging infectious diseases.
METHODS AND ANALYSIS
Study design and aims
BIOMARCOVID is a monocentric, retrospective cohort study of patients hospitalized for COVID-19 at CHUGA. The primary aim is to build a logistic regression model integrating clinical and biochemical parameters collected at hospitalization to predict disease severity outcome. Secondary aims are: (1) to identify differential metabolites and metabolic pathways between mild and severe outcomes using untargeted metabolomics; and (2) to perform an integrated multi-block analysis combining metabolomics, routine blood test, and clinical data. A biobank of clinical, biological, and metabolomic data will be constituted to support future sub-studies.
Participants
Eligible participants must meet all inclusion criteria (Figure 1)
- Admitted to CHUGA for suspected COVID-19
- Confirmed positive PCR test for SARS-CoV-2
Exclusion criteria
- Age <18 years
- Transferred from another hospital
- Expressed opposition to research
- Admitted directly to the ICU on day 1
- Blood samples collected more than 48 hours after admission to CHUGA
Data collection
Clinical evaluation (Clinicome)
Demographic data (age, sex, BMI), medical history (comorbidities including arterial hypertension, obesity, diabetes), and disease progression data (date of first symptoms, date of hospitalization, date of ICU admission and discharge, date of oxygen initiation, date of corticosteroid initiation, date of return home, live/death status) will be recorded for all patients. Involvement of ICU and infectious disease medical teams will support complete clinical data acquisition.
Patient-reported outcome measures
Patients will be classified retrospectively using the WHO 10-category ordinal progression scale: 1□ 0 (uninfected) to 10 (dead). For this study, patients with WHO scores of 4 or 5 (hospitalized, not requiring or requiring low-flow oxygen) are classified as mild COVID-19, and those with scores >5 as severe COVID-19.
Biological samples (Biologicome)
Sixty-two routine blood tests will be included (see online supplemental data 1). These tests were performed according to CHUGA standard procedures upon patient admission. Pre-analytical steps and pricing are described in CHUGA’s biological catalogue (http://www.monkiosquesante.org/KS_EXP/LivretBio). For routine blood tests, samples were analysed prospectively as part of standard clinical care; no randomisation was applicable given the retrospective nature of the study.
Statistical analysis
General considerations
This is a pilot study; no formal sample size calculation was performed. All eligible patients admitted to CHUGA during the study period will be included. Quantitative variables will be described using median and interquartile range; categorical variables using count and percentage. Complete case analysis will be used for the primary outcome. Additional targeted assays were performed to reduce missing data across routine blood tests. Statistical significance is set at α=0.05.
Primary analysis: clinical and biochemical factors
The primary aim is to build a logistic regression model with LASSO variable selection, using clinical and routine biochemical parameters at admission to predict severity outcome. Model performance will be evaluated using accuracy, sensitivity, specificity, positive predictive value, and negative predictive value, as well as the cost of biological parameters (Figure 2). Routine blood test variables with more than 80% missing values across samples will be excluded from analysis. Blinding was applied where appropriate. Unsupervised analyses (PCA, multi-block exploratory analyses) were performed without knowledge of severity group allocation, allowing unbiased exploration of data structure. For supervised analyses (logistic regression with LASSO variable selection, discriminant analyses), severity class labels were provided to the model as the outcome variable, as required by the analytical approach.
Multi-block analysis
The study is structured around three data blocks: Metabolome (untargeted LC-MS/MS data), Biologicome (62 routine blood tests), and Clinicome (clinical characteristics and symptoms). Single-block analyses will first be performed to compute block-specific scores and loadings. A multi-block modelling approach will then be applied to integrate all data sources, provide a global view of their combined information, assess the relative contribution of each block, and identify cross-block variable associations. 21·22
ETHICS AND DISSEMINATION
This study is conducted in accordance with International Conference on Harmonization Good Clinical Practice guidelines. It is a non-interventional, monocentric, retrospective study involving data and samples from human participants, conducted at CHUGA in compliance with French regulations. The principal investigator (Prof. Olivier Epaulard) has signed a commitment to Reference Methodology No. 004 (MR004) issued by the French Data Protection Authority (CNIL). The study is registered on the French Health Data Hub under number F20210218154851. All participants were informed; none expressed opposition. Written informed consent was not required under French national legislation and institutional guidelines. This protocol will be made available upon publication via BMJ Open. Any substantial amendments to this protocol will be submitted to the CHUGA IRB and the CNIL for approval prior to implementation and communicated to all co-investigators.
Patient and public involvement
Patients were not involved in the design, conduct, or reporting of this study. This was a retrospective study using residual biological samples and routine clinical data collected during standard care. Results will be disseminated to the public through lay summaries and conference presentations.
Data availability
The statistical code and analytical pipeline supporting this study are available from the Zenodo repository (DOI: 10.5281/zenodo.20558738). The metabolomics data and metadata are available via MassIVE upon reasonable request, in compliance with the General Data Protection Regulation.
A steering committee comprising intensive care physicians, infectious disease specialists, medical biologists, and statisticians will validate data accuracy prior to data lock and oversee sub-study project selection. Results will be published in peer-reviewed journals and presented at national and international scientific and clinical conferences.
DISCUSSION
Numerous studies have investigated blood biomarkers and clinical risk factors to predict the severity and mortality of COVID-19.1□1□ Elevated systemic cytokine levels, immune activation, and organ injury markers have been documented, driven in part by the angiotensin-converting enzyme 2 (ACE2) receptor, which is expressed in multiple organs and serves as the functional entry point for SARS-CoV-2.23
Understanding ACE2-mediated pathophysiology is critical because many factors — including age, sex, ethnicity, comorbidities, and medication — influence both ACE2 expression and COVID-19 severity. However, ACE2-expressing organs do not equally participate in COVID-19 pathophysiology, implying that additional mechanisms contribute to tissue damage and disease severity. 23
BIOMARCOVID addresses this gap by coupling 62 routine blood tests with untargeted multiplatform metabolomics in a well-characterized cohort of hospitalized patients. The use of plasma samples collected within 48 hours of admission — before severe symptom onset — is a key methodological strength, as it ensures that biomarkers reflect early pathophysiological processes rather than late-stage complications.
Metabolomic profiling has the potential to differentiate mild from severe SARS-CoV-2 infection and to identify metabolic pathways dysregulated in early disease. Combined with routine blood test data and clinical parameters through a multi-block analytical framework, this approach may enable more precise patient stratification, better anticipation of ICU bed requirements, and identification of early therapeutic targets. 11 Importantly, the methodology developed here is designed to be transferable to other emerging respiratory infections.
The main limitations of this study are its monocentric, retrospective design; the absence of a formal sample size calculation; incomplete data for some routine tests; and the potential influence of self-medication or variable time from symptom onset to admission on metabolomic profiles. These limitations are addressed in part by the multi-platform metabolomics approach, data imputation strategies, and the prospective constitution of a biobank for future validation studies.
Supporting information
FUNDING
ALG was supported by “Vaincre la Mucoviscidose” (VLM) and “Association Grégory Lemarchal” (AGL) (Grant number RF20230503289), ANR-15-IDEX-02, FINOVI, and Fondation Université Grenoble Alpes. Part of this work was performed at the GEMELI-GExiM metabolomics platform. The funders had no role in study design, data collection, analysis, decision to publish, or preparation of the manuscript. This work was supported by the MetaboHUB infrastructure funded by the Agence Nationale de la Recherche under the France 2030 program (MetaboHUB ANR-11-INBS-0010; MetEx+ ANR-21-ESRE-0035; MetaboHUB (JVCE) ANR-24-INBS-0012).
LICENCE STATEMENT A
This article is submitted under the Creative Commons Attribution licence (CC BY). This licence permits reuse, distribution, and reproduction in any medium, provided the original work is properly cited.
COMPETING INTERESTS
All authors declare no conflict of interest.
ETHICS APPROVAL AND CONSENT TO PARTICIPATE
This research was approved by the CHUGA institutional review board and authorized following filing with the CNIL under French procedure for a monocentric study (MR004 study). All patients were informed and none expressed opposition to research use of their data; written consent was not required under national legislation.
ACKNOWLEDGEMENTS
We warmly thank Z. Baidi for assistance with this research.