Discovering Novel Genome Mutations Associated with Liver and Pancreas Disorders: Insights from a Radiogenomics Study
1Department of Diagnostic Radiology, Li Ka Shing Faculty of Medicine, The University of Hong Kong, Pok Fu Lam, Hong Kong SAR, PR China
2Division of Applied Oral Sciences and Community Dental Care, Faculty of Dentistry, The University of Hong Kong, Hong Kong SAR, PR China
3Liver Transplant Center, Queen Mary Hospital, Department of Surgery, The University of Hong Kong, Hong Kong SAR, PR China
*Corresponding author: koohi@hku.hkAbstract
Radiogenomics provides a powerful approach to identify genetic variants linked to imaging-derived phenotypes, offering insights into the genetic underpinnings of organ function and disease susceptibility. The liver and pancreas are critical to metabolism, digestion, and detoxification, with their dysfunction leading to conditions such as pancreatitis, diabetes, and cancer. This study used a genome-wide association study (GWAS) to identify genetic variants associated with liver and pancreas radiomics phenotypes derived from magnetic resonance imaging (MRI). We conducted a cross-sectional study using data from 38,844 unique subjects in the UKBiobank, each with available MRI scans and genome sequences. We identified several novel single nucleotide polymorphisms (SNPs) associated with liver and pancreas characteristics. Notable findings include associations with rs1800562 (HFE gene), rs58542926 (TM6SF2 gene), rs738409 (PNPLA3 gene), and rs855791 (TMPRSS6 gene). These findings may contribute to our understanding of the genetic basis of organ function and could potentially help in future preventive healthcare.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
Seed fund for HKU staff
Introduction
The liver and pancreas are vital organs responsible for numerous physiological processes, including metabolism, digestion, and waste detoxification 1-3. Dysfunction in these organs can lead to severe conditions such as non-alcoholic fatty liver disease (NAFLD), pancreatitis, diabetes, and various cancers 4,5. Among the factors contributing to these disorders, abnormal iron metabolism, fat deposition, and changes in organ volume have emerged as significant risk factors 6-9. Understanding the genetic underpinnings of these traits could provide valuable insights into the pathogenesis of liver and pancreas-related diseases and offer potential avenues for personalized healthcare interventions.
Advances in radiogenomics, which integrate imaging data with genomic information, have opened new opportunities to explore the genetic basis of complex traits 10,11. Magnetic resonance imaging (MRI), in particular, offers a non-invasive method to measure radiomics traits such as iron levels, proton density fat fraction (PDFF), and organ volume in the liver and pancreas 12,13. These imaging-derived phenotypes provide high-resolution, quantitative data, enabling the identification of genetic variants associated with specific organ characteristics. Genome-wide association studies (GWAS) have proven to be a powerful tool for uncovering genetic variants linked to complex traits, offering a comprehensive approach to identifying single nucleotide polymorphisms (SNPs) associated with various phenotypes 14,15.
While previous GWAS have focused on traditional biomarkers and clinical outcomes, the integration of radiomics and genomics remains relatively underexplored. This study aims to bridge this gap by conducting a large-scale GWAS of MRI-derived radiomics traits in the liver and pancreas, leveraging data from the UKBiobank 16, a vast repository of genetic and imaging data. By identifying genetic variants associated with liver and pancreas iron levels, PDFF, and organ volume, this study seeks to uncover novel genomic insights that could inform our understanding of organ-specific pathophysiologies, with a particular focus on metabolic disorders and iron homeostasis.
Our study identified several significant genetic variants associated with MRI-derived traits in the liver and pancreas. These findings highlight the genetic complexity of liver and pancreas traits, offering insights into metabolic disorders and potential pathways for new treatment. Beyond clinical implications, such discoveries could potentially help our understanding of the molecular mechanisms driving organ dysfunction, thus aiding in the development of personalized medical strategies 17,18. By integrating radiogenomics data, we hope to provide new insights that could potentially inform future diagnostic and therapeutic strategies for liver and pancreas disorders.
Materials and methods
Study Population and Data Sources
This cross-sectional study utilized data from the UK Biobank, a large population-based cohort comprising over 500,000 participants aged between 40 and 69 years at recruitment 19. We extracted data for 38,844 unique participants who had undergone both abdominal magnetic resonance imaging (MRI) and whole-genome sequencing. The population cohort in this study was from the UK Biobank [Application Number 78730] which received ethical approval from the North West Multicentre Research Ethics Committee (REC reference: 11/NW/03820). All participants gave written informed consent before enrolment.
MRI Data Acquisition and Processing
Abdominal MRI scans were performed using Siemens Aera 1.5 T scanner (Syngo MR D13) (Siemens, Erlangen, Germany) following standardized protocols detailed in the UK Biobank imaging documentation 20. Briefly, MRI sequences included the acquisition of liver and pancreas images to assess iron content, PDFF, and organ volumes. Processed MRI-derived radiomics features for liver and pancreas were obtained through the UK Biobank Research Analysis Platform (UK Biobank RAP, ukbiobank.dnanexus.com). The radiomics traits extracted for each participant included, liver iron content, liver PDFF, liver volume, pancreas iron content, pancreas PDFF and pancreas volume. These measurements were generated using standardized image analysis pipelines provided by UK Biobank, which utilized advanced algorithms for quantification of organ-specific parameters.
Radiomics Trait Thresholding and Cohort Selection
We focused on six MRI-derived radiomics traits, including liver iron levels, liver PDFF, liver volume, pancreas iron levels, pancreas PDFF, and pancreas volume. For each radiomics trait, we calculated the mean and standard deviation (SD) across the entire dataset. Participants were categorized into case and control groups based on whether their measurements exceeded a threshold defined as the mean plus one SD (mean+SD) of the respective radiomics feature. Participants with values higher than this threshold were assigned to the case group, representing elevated levels of the trait, while those with values below the threshold were assigned to the control group. Subjects lacking data for any of the radiomics features were excluded from the respective analyses. The number of participants in each group for each trait shown in Table 1.
Genotype Data Preparation and Genome-Wide Association Study
Whole-genome sequencing data for the participants were obtained from the UK Biobank, which provided high-quality genotypic data generated using standard protocols 21. Data extraction focused on the exome sequences with PLINK-formatted files provided by UK Biobank. We used a combination of Python packages, including Pandas, PySpark, and dxpy, to retrieve and process the genome data. For the GWAS, we used both PLINK 22 and REGENIE 23 to analyze the association between genetic variants and MRI-derived radiomics traits. The genomic data were filtered to include only the participants with available radiomics traits, and necessary quality control (QC) steps were applied to ensure data integrity. The QC process included the following steps 24: 1) Individual and SNP Missingness: We applied filters to remove individuals and SNPs with high missingness rates. Specifically, we used PLINK2 with the --geno 0.1 option to exclude SNPs with a missing call rate > 10%, and the --mind 0.1 option to exclude individuals with a missing genotype rate > 10%. 2) Sex Discrepancy: We addressed inconsistencies between assigned and genetic sex by filtering participants where the reported sex (field p31) matched their genetic sex (field p22001). Participants with discrepancies were excluded from the analysis. 3) Minor Allele Frequency (MAF): To ensure sufficient power for detecting associations, we filtered out rare variants with a minor allele frequency (MAF) < 0.01 using the --maf 0.01 option in PLINK2. Additionally, we used the --mac 100 option to exclude variants with a minor allele count below 100. 4) Hardy-Weinberg Equilibrium (HWE): SNPs deviating from Hardy-Weinberg equilibrium were filtered out using the threshold --hwe 1e-15 in PLINK2. This helped ensure that only variants in equilibrium were retained for the analysis. 5) Heterozygosity Rate: We monitored the heterozygosity rate to detect outliers that could suggest underlying genotyping errors or sample contamination. This was handled as part of the QC metrics in the UK Biobank data pipeline. 6) Relatedness: To remove participants with close genetic relatedness, we excluded individuals with a kinship coefficient > 0.125 (field p22021). This ensured that only unrelated individuals were included in the analysis, reducing potential biases from familial relationships. 7) Population Stratification: We limited the analysis to participants of White British ancestry (field p22006), which helped mitigate potential confounding due to population stratification. Further, we included the top 10 principal components (PCs) as covariates in the association models. This adjustment was done to further correct for subtle population structure within the White British subset. All analyses were performed on the UK Biobank Research Analysis Platform (RAP) using large-scale computing resources suitable for handling genomic and MRI data.
Post-GWAS Analysis and Plotting
After running the GWAS, we applied post-processing steps to identify significant genetic associations. First, we merged the GWAS results across all chromosomes into a single comprehensive results file to streamline analysis. To isolate the most promising genetic variants, we set a genome-wide significance threshold of p < 5×10-7. Variants meeting or exceeding this threshold were considered statistically significant and were subjected to further analysis. For visualization, we used LocusZoom 25 to generate both Manhattan Plots and QQ plots. Additionally, LocusZoom facilitated fine-mapping of the associated regions, allowing us to pinpoint potential causal variants within complex loci and providing deeper insights into the genetic architecture influencing liver and pancreas radiomics phenotypes.
Network and Enrichment Analysis
To further understand the biological significance of the identified SNPs, we conducted SNP enrichment analysis using the web-based tool snpXplorer 26. This tool allowed us to perform enrichment analysis and generate plots to visualize the enriched pathways and biological processes associated with the SNPs identified in our study. For tissue enrichment analysis, we utilized the web-based tool Enrichr 27. This platform enabled us to analyze the genes linked to the significant SNPs, providing insights into their expression in different human tissues. Besides, we used Enrichr-KG 28 to perform network analysis to find the connection between the SNPs and the reported traits and diseases.
Results
Genetic variants significantly associated with liver and pancreas MRI-derived traits
Our genome-wide association study identified multiple single nucleotide polymorphisms (SNPs) significantly associated with liver and pancreas traits derived from MRI phenotypes (Fig S1-S6). Several SNPs showed strong associations with liver iron levels. The most significant was rs1800562 (-log10 p= 60.437), located within the HFE gene on chromosome 6, which is known to be involved in iron metabolism. Other notable associations included rs13219787 near H2BC17 (-log10 p= 19.966), rs3130253 near MOG (-log10 p= 14.288),
rs3094093 near MDC1 and MDC1-AS1 (-log10 p= 11.195), rs855791 within TMPRSS6 on chromosome 22 (-log10 p= 9.805), rs2233952 near PSORS1C1 and PSORS1C2 (-log10 p= 8.992), and rs11752919 near ZSCAN23 (-log10 p= 6.356), all on chromosome 6.
For liver PDFF, the strongest association was observed with rs58542926 (-log10 p= 45.923) near TM6SF2 and AC138430.1 on chromosome 19. Additionally, rs738409 within PNPLA3 on chromosome 22 (-log10 p= 37.377), rs429358 near APOE on chromosome 19 (-log10 p= 8.298), and rs72785308 near TMC5 on chromosome 16 (-log10 p= 7.05) showed significant associations. For liver volume, the most significantly associated SNP was rs4665972 (-log10 p= 7.7), located near the SNX17 gene on chromosome 2. Regarding pancreas PDFF, the most significant SNP was rs1341982 (-log10 p= 7.131), located near CDKN2C on chromosome 1. For pancreas volume, the top associations were rs2287990 near CTRB1 on chromosome 16 (-log10 p= 9.976), rs77581903 near AGAP12P on chromosome 10 (-log10 p= 8.925), and rs3734626 near RSPO3 and AL356534.1 on chromosome 6 (-log10 p= 6.322). The complete results can be found in Table 2 and Fig 1A.
Genomic landscapes analysis highlight diversity of SNP associations with MRI traits
Our results show a variety of SNP types, categorized by their genomic locations and functions. Specifically, we identified five positional SNPs, four coding SNPs, and six eQTLs (expression Quantitative Trait Loci) (Fig 1B). To understand the impact of these SNPs on gene function, we analyzed the number of genes affected by each variant (Fig 1C). This analysis helps to highlight the pleiotropic effects of certain SNPs, demonstrating that a single variant can influence multiple genes. For instance, the variant rs4665972 on chromosome 2 is associated with several genes, including CAD, PPM1G, EIF2B4, GCKR, FNDC4, and KRTCAP3. Similarly, rs3130253 on chromosome 6 affects multiple HLA genes.
We also examined the distribution of genes affected by SNPs across different chromosomes (Fig 1D). Our findings indicate that certain chromosomes, such as chromosome 6, harbor a higher density of SNPs associated with multiple genes. This chromosome-specific analysis provides insights into the genetic architecture and potential hotspots for functional variants related to liver and pancreas disorder. The minor allele frequency (MAF) of the identified SNPs varies significantly, ranging from very low frequencies to more common variants. For example, rs72785308 on chromosome 16 has a MAF of 0.02596, while rs3734626 on chromosome 6 has a MAF of 0.4531. The sources of annotation for these SNPs are diverse, including coding regions, regulatory elements, and non-coding regions. This information is visually summarized in a circular chart, providing a comprehensive view of the allele frequencies and their functional annotations (Fig 1E).
SNP enrichment analysis reveals key pathways in iron homeostasis
In addition to our initial SNP analysis, we conducted further investigations to understand the association of genes with specific traits. Our analysis revealed that certain genes are strongly associated with various reported traits (Fig 1A). Notably, the highest fraction of genes is associated with hemoglobin measurement, with nine genes implicated. Following closely, seven genes are linked to type II diabetes mellitus. Hematocrit and total cholesterol measurement each have five associated genes. Additionally, four genes are associated with aspartate aminotransferase measurement, liver fat measurement, liver disease biomarker, mean corpuscular hemoglobin concentration, reticulocyte count, and serum alanine aminotransferase measurement. This distribution highlights the significant role of these genes in various physiological and pathological processes. Importantly, the association with traits related to iron metabolism, such as hemoglobin measurement and liver disease biomarker, supports our hypothesis that these genes play critical roles in iron-related traits of MRI images.
To further understand the biological significance of the identified SNPs, we performed an enrichment analysis. The results show that several pathways are enriched with our SNPs (Fig 1B). The most significantly enriched pathway is “Hfe effect on hepcidin production” (WP:WP3924). This pathway is crucial for iron regulation and homeostasis. Additionally, other pathways such as “Iron metabolism disorders” (WP:WP5172) show moderate enrichment. These findings suggest that the identified SNPs play important roles in various biological processes, with a notable emphasis on iron metabolism.
Network analysis of SNP-associated genes shows connections to multiple traits and diseases
To further investigate the functional implications of the SNPs identified in our study, we conducted a network analysis to explore the connections between the genes associated with these SNPs and various diseases and GWAS traits. We used the GWAS catalog 29 and the Jensen Diseases database 30 for our network analysis (Fig 3A). Our analysis revealed significant associations between the identified genes and several traits and diseases. The top traits associated with the genes include hematology traits, iron status biomarkers, autism spectrum disorder or schizophrenia, hemoglobin concentration, and platelet count (Fig 3B). Additionally, several diseases were found to be significantly connected with the identified genes. These diseases include iron metabolism disease, fatty liver disease, peeling skin syndrome, hemochromatosis, and systemic scleroderma. The gene-trait and gene-disease network illustrates the complex interconnections between the identified genes and the various traits and diseases.
Tissue enrichment analysis highlights liver as a key organ with high expression of genes associated with identified SNPs
To further understand the biological relevance of the tissue associated with the identified SNPs, we performed a tissue enrichment analysis. This analysis aimed to determine if these genes exhibit meaningful expression in specific tissues, particularly the liver and pancreas. We utilized the GTEx database 31 for this analysis. The results of our analysis revealed significant expression of these genes in various tissues and demographic groups (Fig 3C). Notably, the highest expression was observed in the liver of males aged 50-59 years, with a p-value of 1.47e-05. These genes also showed expression in the liver of females aged 60-69 years, with a p-value of 0.00164, suggesting that these genes are consistently expressed in liver tissue across different genders and age ranges. Additionally, gene expression was observed in the liver of males aged 30-39 years, with a p-value of 0.0025, further supporting their presence across various demographics.
Our findings also highlighted gene expression in the stomach across both genders and different age groups, although with slightly higher p-values, such as 0.0010 in males aged 50-59 years and 0.0014 in females aged 40-49 years. Similarly, the colon of males aged 40-49 years showed gene expression with a p-value of 0.0014. These results indicate a meaningful tissue-specific expression pattern, particularly in the liver, which could have implications for understanding biological functions or disease mechanisms in these tissues.
Discussion
In this study, we performed a comprehensive genome-wide association analysis integrating radiomics phenotypes derived from MRI imaging with genetic data from the UK Biobank. Our aim was to identify genetic variants associated with liver and pancreas iron levels, PDFF, and organ volume. The radiogenomic approach allowed us to uncover several novel and known SNPs significantly associated with these traits.
Pancreas-specific associations highlight genes involved in exocrine and endocrine functions
In the pancreas, we identified rs1341982 near CDKN2C (-log10 p = 7.131) associated with pancreas PDFF. CDKN2C is a cyclin-dependent kinase inhibitor involved in cell cycle regulation 43. Its role in pancreatic β-cell proliferation and function suggests that variants in CDKN2C may influence pancreatic fat deposition, potentially affecting insulin secretion and glucose homeostasis.
For pancreas volume, significant associations were found with rs2287990 near CTRB1 (-log10 p = 9.976), rs77581903 near AGAP12P, and rs3734626 near RSPO3. CTRB1 is a digestive enzyme precursor produced by pancreatic acinar cells. Variants near CTRB1 have been linked to type 2 diabetes and pancreatitis, suggesting a connection between exocrine function and pancreas size 44.
RSPO3 is involved in the Wnt signaling pathway, which plays a role in cellular proliferation and differentiation 45. Its association with pancreas volume may reflect its influence on pancreatic development and maintenance. These findings indicate that genetic factors affecting both exocrine and endocrine components contribute to pancreatic morphology and function.
Network and tissue enrichment analyses emphasize the multifaceted roles of identified genes
Our network analysis revealed that the genes associated with the identified SNPs are connected to a variety of traits and diseases, particularly hematological traits and iron status biomarkers. The strongest associations were with hemoglobin measurement and type II diabetes mellitus, highlighting the interplay between iron metabolism, glucose regulation, and organ function.
The identified genes also showed significant expression in liver tissue across different age groups and genders, as indicated by our tissue enrichment analysis using Genotype-Tissue Expression (GTEx) data. The highest expression levels were observed in the liver of males aged 50-59 years (p = 1.47E-05), underscoring the liver-specific relevance of these genes. The expression patterns support the functional significance of these genes in hepatic physiology and their potential impact on liver-related disorders.
Conclusion
By integrating MRI-derived radiomic phenotypes with genomic data, our study provides a nuanced understanding of how genetic variation influences organ-specific traits. The use of non-invasive imaging biomarkers allows for precise quantification of organ characteristics. This approach enhances the detection of genetic associations that may be missed when relying solely on traditional clinical phenotypes. Our findings highlight the value of radiogenomics in uncovering the genetic determinants of complex traits. The identified SNPs not only confirm known associations but also reveal novel variants that contribute to the variability in liver and pancreas characteristics. This integrative methodology paves the way for more personalized risk assessment and targeted interventions based on an individual’s genetic profile.
Limitations and future directions
While our study provides valuable insights, there are limitations to consider. The cross-sectional design can’t reveal causality or temporal relationships. Longitudinal studies are necessary to determine how these genetic variants affect organ traits over time and their role in disease progression. The study population was derived from the UK Biobank, which predominantly includes individuals of European descent. Therefore, the generalizability of the findings to other ethnicities may be limited. Future studies should include more diverse populations to explore potential differences in genetic associations across ethnic groups.
Functional validation of the identified SNPs is also needed to elucidate their precise biological effects. Experimental studies using cellular and animal models can help confirm the mechanisms by which these variants influence liver and pancreas physiology. Finally, integrating additional omics data, such as transcriptomics and proteomics, could provide a more comprehensive understanding of the pathways involved and identify potential therapeutic targets.
Conflict of interest
The authors of this study declare that they do not have any conflict of interest.
Data availability statement
The data supporting the findings of the study are available to researchers upon approval of an application to the UK Biobank (https://www.ukbiobank.ac.uk/researchers/).
Funding
This research was funded by the University of Hong Kong Seed Fund for Research – Translational and Applied Research, awarded to MK.