Sampling bias in large healthcare claims databases
1Quantitative Sciences Unit, Department of Medicine, Stanford University School of Medicine, Stanford, CA
2Department of Pathology, Stanford University School of Medicine, Stanford CA
*Authors for correspondence: Alex Dahlen, PhD, Quantitative Sciences Unit, Department of Medicine, Stanford University School of Medicine, Stanford, CA, adahlen@stanford.edu, Vivek Charu, MD PhD, Quantitative Sciences Unit, Department of Medicine, Department of Pathology, Stanford University School of Medicine, Stanford, CA, vcharu@stanford.eduAbstract
Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum's de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Introduction
Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence1–4. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum’s de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.
Methods
Optum’s CDM consists of administrative claims derived from several large commercial and Medicare Advantage health plans. For our primary analysis, we count the number of individuals with CDM coverage on a single day (June 1, 2018) in each zip code, and compare it to the 2018 census estimates of the total population in that zip code. In sensitivity analyses, we also consider one strictly larger cohort (individuals with at least one day of coverage at any point during 2018) and one strictly smaller cohort (individuals with continuous coverage during the entire year of 2018).
To explain the variation in zip-code level CDM sampling, we fit an inverse-variance weighted multivariable linear regression model with 29 socioeconomic and demographic features extracted from the 2018 census, and state-level fixed effects. See Table 1 for the list of features, and the Appendix 1 for more details about the model and methods.
Results
16.4 million distinct individuals were captured in CDM on June 1, 2018, representing 5.4% of the US population. The median zip code sampling fraction was 4.4% (interquartile range: 2.5-7.1%), with clear geographic variation in sampling (Figure 1). At the state level, Alaska had the lowest sampling rate (0.9%) and Colorado had the highest sampling rate (10.9%); in multivariable regression models, state-level fixed effects explained 34.6% of the zip-code level variation in sampling.
Associations between socioeconomic/demographic features and CDM sampling fraction, after adjusting for state-level variation, are provided in Table 1. Estimated partial correlations and regression model coefficients demonstrate that inclusion in CDM is associated with zip codes that are wealthier, older, more educated, and disproportionately White. These patterns are robust to the choice of cohort definition, and across the 10 most populous states individually (Appendix 2) Socioeconomic and demographic features explain an additional 19.4% of the zip code level-variation in sampling on top of state-level variation, for a total adjusted R2 of 54.0%.
Discussion
To interpret results generated from healthcare claims databases, it is essential understand which patients are represented in them. We demonstrate that inclusion in Optum’s CDM at the zip-code level in 2018 varies spatially and along socioeconomic and demographic lines. We analyzed data at the smallest geographic scale available in CDM, the zip code; given our findings, there is likely to be additional bias within zip codes as well.
The same socioeconomic and demographic features that correlate with over-representation in claims data have also been shown to be effect modifiers across a diverse spectrum of health outcomes, including the burden of disease and access to healthcare5-6. This combination of heterogenous sampling and effect modification—both driven, in this case, by social determinants of health—gives rise to external validity bias, where results generated from the claims data will fail to generalize to the underlying population7-8. This external validity bias can affect both studies that estimate disease incidence or prevalence and comparative effectiveness studies that use contemporary causal inference methods: when there is meaningful heterogeneity in treatment/policy effects along socioeconomic or demographic lines, then claims-derived estimates will be biased.
Our study highlights the importance of investigating sampling heterogeneity in analyses of large healthcare claims data to evaluate how sampling bias might compromise the accuracy and generalizability of results9. Importantly, investigating these biases or accurately correcting for such bias will require external data sources outside of the claims database itself. Healthcare claims databases offer enormous promise for medical research; characterizing and overcoming sampling bias in these datasets is essential.
Conflict of Interest
Neither of the authors have any conflicts of interest relevant to the subject matter.
Acknowledgement
We thank Kate Miller and Steve Goodman for conversations about the manuscript. Data for this project were accessed using Stanford Center for Population Health Sciences (PHS) Data Core. The PHS Data Core is supported by a National Institutes of Health National Center for Advancing Translational Science Clinical and Translational Science Award (UL1TR003142) and from Internal Stanford funding. The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH. Dataset DOI: 10.57761/y4af-nd44.
Appendix Group
Detailed Methods
Here we describe additional details about the methodology used.
Optum’s de-identified Clinformatics® Data Mart Database (CDM) is a de-identified database derived from a large adjudicated claims data warehouse. We considered three cohort definitions: (1) all unique patients in CDM on June 1, 2018; (2) all unique patients with at least a single day of coverage from January 1, 2018 to December 31, 2018 and (3) all unique patients with coverage during the entire year, January 1, 2018 to December 31, 2018. Cohort 1 parallels a common use of the that computes incidence or prevalence rates by normalizing counts by patient-days of coverage; cohort 2 is the maximal count of covered patients in 2018; and cohort 3 parallels a different common use of the data where cohorts with longer periods of continuous coverage are compared.
For each cohort definition, we estimated the zip-code level CDM sampling fraction as the number of unique patients in each zip code in CDM divided by the number of unique persons in the American Community Survey (ACS) 2018 5-year census. Of note, CDM aggregates some smaller zip-codes into groups; all analyses are performed at the level of CDM’s zip-code clusters definition.
Socioeconomic/demographic information was derived from the ACS 2018 5-year census, and included population density; the fraction of persons identifying as Black, Asian, White, Hispanic, or other races; the fraction of persons aged <18, 18-39, 40-59, 60-19, and >80; the fraction of households earning (in US dollars) less than $15,000, $15,000-$30,000, $30,000-$45,000, $45,000-$60,000, $60,000-$100,000, $100,000-$125,000, $125,000-$200,000, and over $200,000; the fraction of individuals unemployed and the fraction of individuals without health insurance; the fraction of individuals with who completed less than high school, high school, some college, college, and graduate school; the fraction of houses owner-occupied; and the median house price. This data was read in at the census track level and then to get the data on equal footing, we rolled the census data up using a population-weighting or square mileage-weighting, as appropriate.
The multivariable model is a linear regression of the form: where i indexes each zip-code, Yi is the CDM sampling fraction in each zip code, Xi represents a vector of 25 census features for each zip code (29 total socioeconomic and demographic features with 4 comparisons groups) and Zi represents a vector of 50 state-level fixed effects (51 total states including the District of Columbia with one comparison group). Each zip code is weighted by the inverse of the variance associated with the zip code level estimate of CDM sampling. We used the standard binomial estimate of variance, modified by Winsorizing the estimate of the sampling fraction at the 5th percentile, to prevent unstable weighting. Thus, the weights used scale with the total census population in each zip-code, down-weighting the contribution of very small and high variance zip codes. We produced standard diagnostic plots for this model, which demonstrated a reasonable fit.
Partial correlation coefficients were computed using the analogous short model with only a single census variable, in addition to the 50 fixed effects. The analysis was carried out in Python version 3.8.5, and the code to reproduce this analysis is publicly available at: https://github.com/alex-dahlen/who_is_in_optum.
Results of Sensitivity Analyses
Figure 3. provides unadjusted zip-code level correlation coefficients between socioeconomic and demographic features and CDM sampling for the 10 most populous states. Table 2 provides results of sensitivity analyses in which we vary the cohort definition, as described above.