Forecasting the Diabetes Burden Across USA-Mexico Border States Through 2030: A Multi-Indicator, Multi-Model Analysis
1Department of Population Health Sciences, School of Public Health, Georgia State University, Atlanta, GA, 30303, USA
2Department of Applied Mathematics, Kyung Hee University, Yongin 17104, Korea
3Centre for Research on Pandemics & Society (PANSOC), Oslo Metropolitan University, Oslo, 0130, Norway
*Corresponding authors: Kaustubh Wagh kwagh2@student.gsu.edu; Svenn-Erik Mamelund masv@oslomet.no; Gerardo Chowell gchowell@gsu.eduAbstract
Objectives
The USA-Mexico border area faces an elevated diabetes burden, with prevalence 2-3 times higher than national averages in both countries. Yet, no forward-looking projections exist beyond 2021, limiting healthcare system planning. We provide the first multi-indicator, multi-model forecasts of diabetes burden through 2030 across border states.
Methods
Diabetes mortality, prevalence, and disability-adjusted life years (DALYs) data were extracted from the Global Burden of Disease Study 2021 for all ten border states in both countries (1990-2021). Eight different time-series models including ARIMA, generalized additive models, generalized linear models, Prophet, and n-sub-epidemic framework were used to project burden from 2022-2030. Models were evaluated and compared using weighted interval scores, with forecasts stratified by age, sex, and geography.
Results
Multi-model ensembles project 15-17 million people with diabetes, 23,000-26,000 annual deaths, and 1.6-1.8 million Disability-Adjusted Life Years (DALYs) due to diabetes in USA-Mexico border states by 2030. Mexican border states face accelerating mortality among adults aged 20-59 years (15-58% increases), while USA states show prevalence increases of 35-52% in youth and 47-49% in older adults. Males experienced consistently higher projected burden than females. ARIMA demonstrated superior performance for mortality and prevalence forecasting, while ensemble methods performed better for DALYs projections.
Conclusions
Diabetes burden is projected to increase 30-50% across border states by 2030, driven by distinct regional patterns requiring divergent public health responses. Mexican states require community health worker programs, subsidized medications, and mobile clinics for mortality prevention, while US states must expand school-based screening and nutrition program capacity to address rising youth prevalence. These projections provide actionable intelligence to support coordinated cross-border efforts in diabetes prevention and management.
Strengths and Limitations of This Study
Strengths
- – Multi-model ensemble approach (eight models) with rigorous comparative evaluation using standardized metrics, providing transparency on performance variability across outcomes and demographics.
- – Comprehensive stratification by age group (five categories), sex, and geography enables demographically targeted healthcare planning, addressing a gap in prior border health forecasting studies.
- – Explicit quantification of forecast uncertainty through 95% prediction intervals and formal calibration assessment, identifying outcome-specific and geographic instances of substantial uncertainty.
Limitation
- – Forecasts use statistically modeled disease burden estimates (GBD 2021) rather than primary surveillance data, potentially introducing measurement error and limiting precision in capturing true state-level variation.
- – Time-series models assume historical trends continue unchanged, unable to account for policy interventions, technological innovations, or epidemiological shocks. Forecasts are most reliable in short term.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
The authors received no external funding.
1.Introduction
The USA-Mexico border area faces a high diabetes burden, with a 2-3 times higher prevalence compared to the national average in both Mexico and the United States (1). Ten USA-Mexico border states—four in the USA (California, Arizona, New Mexico, and Texas) and six in Mexico (Baja California, Sonora, Chihuahua, Coahuila, Nuevo León, and Tamaulipas)—form a binational corridor home to over 100 million people (2,3). The border populations experience a diabetes prevalence of approximately 15.7% compared to 8.2% nationally for Mexico and 5.2% for the United States (1). With mortality rates reaching nearly 50% higher than the national averages, diabetes represents a critical and persistent public health challenge in the region (1,4,5). Border states share distinct epidemiological characteristics: higher concentrations of Hispanic/Latino populations, high poverty rates, and lower health insurance coverage compared to the national averages (6,7), alongside limited healthcare infrastructure, genetic susceptibility (e.g., TCF7L2 polymorphisms among Hispanic/Latino populations), and environmental determinants such as limited access to nutritious food, and the built environment for physical activity (8–10). This unique profile, characterized by cross-border mobility and binational health challenges, demands cross-state and country-level analysis to inform regional policy and resource allocation. Accurate and timely projections of future diabetes burden are therefore essential for state health system planning and binational coordination.
Despite extensive research on diabetes at national levels (11–13), critical gaps persist in understanding border states’ specific trajectory. To date, most diabetes forecasting studies are focused on national or global projections (14–16), with only a few examining state-level patterns in border regions. Existing forecasting studies rely on single indicators such as prevalence or mortality, limiting comprehensive burden assessment.(17,18). Moreover, few studies provide age-sex stratification for state-level interventions or compare multiple forecasting methods to quantify prediction uncertainty. While national analyses documents rising diabetes burden (17), these estimates often obscure within and between-country variation, particularly the elevated burden in border states. Our previous work documented consistently elevated historical trends in deaths, prevalence, and disability-adjusted life years (DALYs) across border states from 1990-2021, revealing consistently higher burden compared to national averages (19). However, no projections are available for these specific border states beyond 2021, leaving state policymakers without any future estimates needed to prepare healthcare systems for the coming decades. This work aims to address that and to our knowledge, this is the first state-level, multi-indicator, multi-model diabetes burden forecasting study for the USA–Mexico border region that integrates diabetes mortality, prevalence, and DALYs. This study advances prior work by comparing multiple modeling frameworks and to produce age- and sex-stratified forecasts with quantified uncertainty.
Our study employs state-level data from all ten border states (six Mexico and four USA) as the appropriate geographic unit for several reasons. First, state-level age–sex–stratified diabetes data are consistently available across decades from the Global Burden of Disease Study (20,21), ensuring comparability across countries. Second, state boundaries represent defined functional and administrative units through which health policy is formulated, healthcare systems are organized, and budgets are allocated, making state-level projections directly actionable (22). Third, the epidemiological influence of the border is broad and extends throughout border states due to population mobility, economic integration, and cultural continuity that transcends proximity to the international boundary (23,24). It is also worth emphasizing that border states exhibit distinct diabetes risk profiles compared to interior states in both countries. Fourth, state-level analysis provides sufficient data for robust age-sex stratified forecasting while enabling meaningful binational comparisons (25). Finally, sister-state relationships and regional health commissions operate at the state level, making state-level burden estimates critical for binational coordination and future collaborations. Thus, state-level analysis aims to provide optimal balance between administrative relevance, data quality, and forecasting reliability for informing border health policy.
We sought to forecast diabetes burden in all ten USA-Mexico border states from 2022 to 2030 using three complementary indicators: deaths (fatal burden), prevalence (healthcare demand), and DALYs (comprehensive disease impact combining premature mortality and disability) (26). Specific objectives include: generating projections for Mexico and USA border states to identify which region face the greatest future burden in terms of numbers; producing age-sex stratified forecasts to inform demographically targeted interventions; comparing eight time-series forecasting models (ARIMA, generalized additive models, generalized linear models, Prophet, and n-subepidemic models) to assess prediction uncertainty and identify optimal approaches for forecasting (27,28); evaluating models calibration through coverage analysis to ensure projection reliability; and comparing projected burden between U.S. and Mexican border states to inform binational coordination. These forecasts provide evidence-based estimates to support planning, binational health coordination, and targeted diabetes prevention policies.
2.Methods
2.1.Data Collection and Preparation
We obtained data for this study from the Global Burden of Disease Study 2021 (GBD2021), a publicly accessible database maintained by the Institute for Health Metrics and Evaluation (IHME) (29). We extracted diabetes-related mortality (both type 1 and type 2), prevalence, and disability-adjusted life years (DALYs) for all USA-Mexico border states from 1990 to 2021 (yearly). Border states included California, Arizona, New Mexico, and Texas on the US side, and Baja California, Sonora, Chihuahua, Coahuila, Nuevo León, and Tamaulipas on the Mexican side. We stratified all estimates by age group and sex to enable related demographic comparisons. We aggregated data for six Mexico states (referred to as MX), four USA states (referred to as US) and all ten USA-Mexico border states (referred to as USMX). The GBD 2021 data incorporate uniform and consistent methods for data collection, processing, and statistical modeling to ensure comparability across locations and time periods (30). Publicly available, de-identified aggregate data were used; institutional review board approval was not required. All data use agreements and terms of service were followed.
We used all 32 years of (yearly) historical data and set the calibration period to 32 years and the forecasting horizon between 2022 and 2030 (yearly forecasting for 9 years). This nine-year horizon aligns with the 2030 Sustainable Development Goal (SDG) targets (31), ensuring policy-relevant projections for the current global health agenda. Data was preprocessed using Microsoft Excel pivot tables. All forecasting analyses were conducted using RStudio, Posit Software, Boston, MA, USA (32) and MATLAB, MathWorks Inc., Natick, MA, USA (33). Specifically, we used the StatModPredict dashboard (developed in RStudio v. 4.5.0) to generate forecasts from the ARIMA, GAM, GLM, and Prophet models (34). For the ensemble n-sub-epidemic modeling framework, we employed the SubEpiPredict toolbox (SubEpiPredict toolbox) implemented in MATLAB (R2025) (35). Both toolboxes support reproducible workflows, automated parameter selection, and standardized output generation.
2.2.Forecasting Models
2.2.1.ARIMA (Autoregressive Integrated Moving Average)
ARIMA is a time series model that builds forecasts by identifying patterns in past diabetes burdens (deaths, prevalence, DALYs) data (36). It combines three components: trends over time (autoregression), smoothing of random fluctuations (moving average), and differencing to stabilize the data to ensure necessary stationarity. ARIMA works the best for short-term forecasting when historical trends are stable, linear and non-seasonal (28), while it is less effective when periodic patterns shift suddenly due to interventions or emerging risk factors (28).
2.2.2.GAM (Generalized Additive Model)
GAMs extend traditional regression models that allow for flexible, nonlinear relationships between time and diabetes burdens (deaths, prevalence, DALYs) (37). Instead of fitting a straight line, GAMs use smooth curves to model long-term trends in diabetes burdens (38) which make such models well suited for capturing gradual changes in diabetes burdens that may result from evolving lifestyle behaviors, treatment access, or demographic shifts.
2.2.3.GLM (Generalized Linear Model)
GLMs are classic statistical models used to relate a discrete response variable (like mortality count) to one or more predictors (such as time). They assume a specific linear relationship between the predictors and the outcome, transformed by a so-called link function, and use a distribution suited for the type of data used in the analysis (e.g., Poisson or normal) (39). GLMs have straightforward formulations and easily interpretable and effective when trends are relatively smooth.
2.2.5.n-Sub-Epidemic Model
This model interprets diabetes burdens’ (deaths, prevalence, DALYs) trends as a series of overlapping “sub-epidemics”, each representing a surge or phase in diabetes deaths (42,43). Each sub-epidemic has its own timing and intensity, allowing the overall model to capture complex mortality dynamics such as plateaus or multiple peaks often seen in real-world chronic disease progression. It is particularly powerful in reflecting regional differences or the cumulative effects of repeated disruptions.
2.2.6.Ensemble Sub-Epidemic Models (Unweighted and Weighted)
These models combine forecasts from the top-performing sub-epidemic models. The unweighted ensemble gives equal importance to each model, while the weighted version gives more influence to models with better fit (44). This ensemble approach improves robustness by balancing strengths across models and helps produce more reliable forecasts with quantified uncertainty. Under the n-sub-epidemic modeling framework, we developed several models, including the n-subepidemic ensemble (of top 2 ranked models) unweighted [NSE-ensemble(2)(UW)], the n-subepidemic ensemble (of top 2 ranked models) weighted [NSE-ensemble(2)(W)], and the top two individually ranked models [NSE-ranked(1) and NSE-ranked(2)]. The ensemble models [NSE-ensemble(2)(UW), and NSE-ensemble(2)(W)] are superior to individual ranked models [NSE-ranked(1), and NSE-ranked(2)] as they improve forecasting performance by systematically incorporating the predictive accuracy of individual models (44). Although the utilized n-sub-epidemic framework was initially validated primarily on infectious-disease time series, its flexible structure allows application to chronic disease burden data, where overlapping sub-trends may reflect multiple epidemiological drivers or intervention phases. However, for transparency and to show the individual models that informed our ensemble approaches, we also presented the results of NSE-ranked (1) and NSE-ranked (2) models.
2.3.Model Evaluation
Models were evaluated using a set of performance metrics that measure both point and probabilistic forecast accuracy simultaneously (44). In particular, the utilized models were evaluated based on Weighted Interval Score (WIS) and mean absolute error metrics, which quantify the deviation between predicted and the observed values for each given point in time and aggregated across time points. We used data from 2013 to 2021 (9 years) to measure the performance of models. In this context, smaller values of those metrics indicate better model performance. The reliability of uncertainty estimates was evaluated based on the proportion of observations falling in the predicted intervals, with 95% P.I. coverage where higher values ideally estimate better calibrated uncertainty (45). It is important to note that WIS provides a balance between sharp and out-of-range predictions due to penalizing prediction intervals that are too wide and observations that fall outside predicted range, thus sustaining an overall balance between forecast sharpness and calibration (42).
3.Results
3.1Performance Metrics
3.1.1Deaths
Mortality forecasting showed uniformly strong performance with minimal model differentiation. ARIMA, GAM, and GLM achieved comparable accuracy (MAE 2.3-3.1, WIS 2.0-3.0) across all demographic strata and geographic regions (Fig 1A, Supplementary Fig S1A). Ensemble methods showed no systematic improvement. Error magnitudes remained consistent across Mexican border states (MAE 2.5-3.0), US border states (MAE 2.8-3.3), and the combined USMX region. Model rankings were stable across age and sex strata.
3.1.2Prevalence
Prevalence forecasting revealed marked model performance differences among models with substantially higher errors than mortality outcomes. ARIMA achieved the lowest MAE values, ranging from 3.2-4.1 in Mexican border states to 4.5-5.0 in US border states, with WIS values following similar patterns (Fig 1B, Supplementary Fig S1B). GAM showed comparable performance. Prophet and ensemble approaches (NSE-Ensemble-Ranked, NSE-Ensemble-UW/W) consistently ranked lowest, with MAE values 15-40% higher than ARIMA across most age-sex combinations. The performance gap between ARIMA and ensemble methods was largest in Mexican border states (MAE difference 1.0-1.5 units) and smallest in US border states (MAE difference 0.5-0.8 units).
3.1.3DALYs
DALY forecasts showed error magnitudes intermediate between deaths and prevalence. ARIMA and GAM maintained lower errors (MAE 2.7-4.3, WIS 2.4-4.2) compared to Prophet and ensemble variants (MAE 3.5-5.2, WIS 3.8-5.0) (Fig 1C, Supplementary Fig S1C). The youngest age groups (<20 years) achieved MAE values 40-50% lower than older populations (60+ years) across all models. Male and female strata produced nearly identical error magnitudes and model rankings within each geographic region for all three outcomes.
3.2Calibration Metrics
3.2.1Deaths
Death forecasts demonstrated the steepest age-dependent performance gradient across all outcomes (Figure 2A, Supplementary Fig S2A). Younger populations achieved lower WIS score and coverage of 95% PI closer to 95% (i.e. better calibration), with models showing WIS values of −0.7 to 0.6 log10 units and coverage of 84-100% for ages <20 years. However, GLM critically failed in middle-aged males (40-59 years), achieving only 46% coverage compared to >92% for ensemble methods. Regional disparities emerged in older populations, where Mexico consistently maintained superior coverage (>92%) across all models, while US border states required ensemble approaches to achieve adequate calibration.
3.2.2Prevalence
Prevalence forecasts exhibited the widest performance variability and most severe model failures (Figure 2B, Supplementary Fig S2B). NSE-Ranked (2) catastrophically under-covered predictions in US border states for both sexes (7.69% coverage), despite acceptable performance elsewhere. This outcome showed pronounced geographic stratification: Mexico border states maintained robust calibration in younger populations (WIS: 0.5-2.9 log10 units, coverage >85%), whereas US border states struggled across all age groups, with older populations (>60 years) requiring ensemble methods to achieve minimally adequate coverage.
3.2.3DALYs
DALY forecasts revealed ensemble superiority across demographic strata (Figure 2C, Supplementary Fig S2C). NSE-Ensemble(2)(W) maintained stable performance in USMX both-sex populations (WIS: 2.5-3.3 log10 units, coverage >92%), while GAM showed age-dependent variability—excellent calibration in youth (<20 years: WIS 1.4, coverage 100%) but substantial degradation in older groups. Sex-stratified analyses showed males experienced moderately elevated WIS in young adults (20-39 years: 3.2-3.7 versus 2.8-3.4 log10 units for females). Mexico and US border states diverged primarily in middle-aged populations, where GLM coverage dropped to 46-54% in US regions.
3.3Forecasting
3.3.1Deaths
Diabetes mortality forecasts from 2021-2030 varied substantially by geography (Fig 3A). MX showed the highest increases in younger adults (20-39 years: GAM 57.8%, ARIMA 14.7%) and middle-aged adults (40-59 years: GAM 40.5%, ARIMA 33.4%). US demonstrated decreases in younger populations (20-39 years: GAM −35.0%, ARIMA −10.0%) but mixed patterns in older ages. USMX showed consistent increases across all age groups (ARIMA 10-21%, GAM 10-49%). Males experienced higher mortality increases than females across all geographies (MX males: ARIMA 28.4% vs females 15.3%; US males: 21.6% vs females 3.2%). NSE models predicted mortality reductions in older US adults (60-79 years: NSE-ranked (2) −21.6%).
Forecasted mortality projections varied by weighting approach (Fig 4). WIS-weighted forecasts projected 23,156 deaths (95% PI: 20,525-25,845) across the combined USA-Mexico border region, while unweighted forecasts estimated 26,419 deaths (95% PI: 22,472-30,528). Geographic distribution showed Mexico forecasted at 13,943 deaths (95% PI: 11,995-15,923) under weighted and 15,959 deaths (95% PI: 13,578-18,458) under unweighted approaches, with the United States at 12,632 (95% PI: 10,263-15,269) and 13,421 deaths (95% PI: 10,772-16,309), respectively. Males were consistently projected to experience higher mortality than females across both methods: 14,066 (95% PI: 12,556-15,631) versus 12,050 deaths (95% PI: 10,622-13,520) for weighted, and 15,368 (95% PI: 13,155-17,716) versus 14,012 deaths (95% PI: 11,896-16,253) for unweighted forecasts. Age-specific projections demonstrated strong concordance between methods, with the 60-79 age group expected to bear the highest burden under both weighted (13,012 deaths) and unweighted (14,848 deaths) approaches, followed by the 80+ and 40-59 age groups. Model-specific projected deaths by 2030 are shown in Table 1, with forecasted trajectories of deaths in Mexico and USA border states are presented in Supplementary Figs S3.1 and S3.2, respectively.
3.3.2Prevalence
Universal prevalence increases characterized most forecasts, but geographic patterns varied considerably (Fig 3B). US showed the steepest rises in youth (<20 years: 35-52%) and older adults (60-79 years: 47-49%), while MX exhibited modest youth increases (11-13%) but substantial middle-aged burden (34% for 40-59 years). USMX demonstrated 22-45% increases across most ages. All geographies showed similar increases among those 80+ years (30-37%). Sex differences remained minimal across all regions, typically under 6 percentage points, with males showing marginally higher prevalence growth than females.
Prevalence forecasts showed substantial differences between weighting methodologies (Fig 4). WIS-weighted forecasts predicted 14.9 million cases (95% PI: 13.0-17.0 million) across the USA-Mexico border region, whereas unweighted forecasts projected higher prevalence at 17.4 million cases (95% PI: 14.8-20.2 million). Geographic patterns remained consistent across methods, with the United States accounting for the majority: 12.5 million cases (95% PI: 10.4-14.8 million) weighted versus 14.9 million cases (95% PI: 12.5-17.5 million) unweighted, compared to Mexico’s 2.5 million (95% PI: 2.1-2.9 million) weighted and 2.9 million (95% PI: 2.4-3.3 million) unweighted. Sex-stratified projections indicated males with slightly higher prevalence than females under both approaches: 7.8 million (95% PI: 6.7-8.9 million) versus 7.1 million (95% PI: 6.2-8.2 million) for weighted, and 9.1 million (95% PI: 7.7-10.5 million) versus 8.3 million (95% PI: 7.1-9.6 million) for unweighted forecasts. The 60-79 age group was consistently projected to demonstrate the highest disease burden across both methodologies, followed by the 40-59 age group. Model-specific projected prevalence by 2030 is shown in Table 2, with forecasted trajectories of prevalence in Mexico and USA border states are presented in Supplementary Figs S4.1 and S4.2, respectively.
3.3.3DALYs
Geographic heterogeneity in DALYs projections was pronounced (Fig 3C). USMX exhibited the highest burden increases among youth and young adults (46% for both age groups), while MX showed near-zero or negative changes in these populations. US demonstrated substantial DALYs increases in the oldest adults (80+ years: 33-46%), contrasting with lower increases in MX (7-10%) for the same age group. Middle-aged adults showed wide geographic variation, with USMX predicting 27-43% increases versus 3-6% in MX. Males experienced 4-10 percentage point greater DALYs increases than females in USMX and US.
Disability-adjusted life year projections exhibited notable variation between weighting schemes (Fig 4). WIS-weighted forecasts estimated 1.6 million DALYs (95% PI: 1.4-1.8 million) across the USA-Mexico border region, while unweighted forecasts projected 1.8 million DALYs (95% PI: 1.5-2.1 million). Geographic contributions differed by method, with the United States projected at 1.1 million DALYs (95% PI: 949,000-1.3 million) weighted versus 1.3 million (95% PI: 1.1-1.5 million) unweighted, and Mexico at 531,000 (95% PI: 451,000-617,000) weighted versus 607,000 (95% PI: 516,000-704,000) unweighted. Sex-specific projections consistently showed males accumulating higher DALYs than females: 857,000 (95% PI: 739,000-982,000) versus 716,000 (95% PI: 617,000-822,000) for weighted, and 990,000 (95% PI: 841,000-1.1 million) versus 833,000 (95% PI: 708,000-966,000) for unweighted forecasts. Both methodologies revealed similar age patterns, with peak burden projected in younger populations (20-39 and <20 age groups), reflecting substantial years of life lost, while older age groups (60-79 and 80+) showed comparatively lower DALY estimates. Model-specific projected deaths by 2030 are shown in Table 3, with forecasted trajectories of DALYs in Mexico and USA border states are presented in Supplementary Figs S5.1 and S5.2, respectively.
Across all outcomes, ARIMA and GAM consistently produced the most reliable forecasts, demonstrating strong point prediction accuracy and similarities across demographic groups. In contrast, weighted ensemble approaches, particularly WIS-weighted models, showed much larger uncertainty in calibration and coverage performance. Together, these results highlight that no single model dominated across all indicators, emphasizing the importance of combining complementing statistical methods and model ensembles to balance accuracy and reliability in diabetes burden forecasting.
4.Discussion
This study provides the first comprehensive state-level forecast of the diabetes burden across all 10 USA-Mexico border states through 2030. Our multi-model ensemble approach projects approximately 23,000-26,000 diabetes deaths, 15-17 million people with diabetes, and 1.6-1.8 million DALYs annually by 2030 across the ten USA-Mexico border states, with highlighted disparities between Mexican and US border states which reflect distinct epidemiological trajectories. Overall, our findings reveal the urgent need for coordinated, binational health strategies to address the escalating diabetes burden in this region.
Our forecasts reveal three critical epidemiological patterns shaping the future diabetes burden across the border region. First, Mexican border states face accelerating diabetes mortality, particularly among younger and middle-aged adults (20-59 years), with projected increases of 15-58% by 2030. This contrasts sharply with US border states, where mortality is projected to decline in younger populations but increase among older adults. This divergence likely reflects Mexico’s ongoing epidemiological transition, characterized by rising obesity prevalence, delayed diabetes diagnosis, and limited access to glycemic control medications (12,46). The USA patterns suggests successful prevention efforts in younger cohorts but accumulated disease burden in aging populations with longer diabetes duration (47).
Second, prevalence projections uniformly indicate substantial increases across all border states, but with markedly steeper growth in US border states (35-52% in youth, 47-49% in older adults) compared to their Mexican counterpart. This pattern may reflect the expected improved survival with diabetes in the US through better treatment access, paradoxically increasing prevalence despite stabilizing incidence (48). The pronounced youth prevalence increase in US border states align with rising childhood obesity and type 2 diabetes onset at younger ages, a phenomenon extensively documented in Hispanic border communities (49,50).
Third, DALY projections highlight large increases in years of life lost among youth in USA border states, while Mexican states show stable or slightly declining DALYs in younger groups. This geographic heterogeneity with USMX projecting 46% increases in youth while Mexican states show near-zero changes suggests differential impacts of early-onset diabetes and contrasting healthcare effectiveness on different sides of the border (51).
Our findings extend and complement previous studies in several important ways. National-level projections by Lin et al. (2018) estimated US diabetes prevalence reaching 54.9 million by 2030 (11), while our border state forecasts (12.5-14.9 million) thus representing 23-27% of this expected national burden despite border states comprising only 15% of the US population, confirming the disproportionate regional impact. Similarly, Meza et al. (2015) projected Mexican national diabetes prevalence at 13.7 million by 2030 (52), with our border state expected estimates (2.5-2.9 million) representing 18-21% of the national total while border states contain approximately 20% of Mexico’s population, suggesting burden comparable to population distribution in Mexico while still elevated in absolute rate values. What is crucial is that our study addresses a few methodological gaps in previous forecasting research since most prior studies employed single models (i.e. typically ARIMA or exponential smoothing) without quantifying prediction uncertainty through detailed model comparisons (53,54). Our ensemble approach revealed substantial differences in model’s performance, particularly for prevalence forecasting where ARIMA outperformed ensemble methods by 15-40%. These finding challenges assumptions that ensemble averaging universally improves forecast accuracy (38) and suggests outcome-specific model selection strategies may be necessary and should not be discarded. The catastrophic under-coverage that was observed with certain models in US border states [7.69% for NSE-ranked (2)] underscores the importance of rigorous calibration assessment beyond point estimate accuracy. While ARIMA and GAM offered strong interpretability and transparent trend estimation, ensemble models provided greater flexibility and improved uncertainty calibration. Recognizing that the optimal model may vary by outcome type (i.e., mortality, prevalence, or DALYs) represents an important methodological contribution of this study, supporting an outcome-dependent approach to model selection.
Our age and sex stratified forecasts also address a critical limitation in existing border health research. Previous border diabetes studies provided descriptive epidemiology (55) or intervention evaluations (56) but lacked forward-looking projections with the desired level of demographic stratification details necessary for healthcare capacity planning. The pronounced sex differences identified in the analysis including males experiencing 4-10 percentage point greater increases in mortality and DALYs align with documented gender disparities in diabetes self-management, healthcare utilization, and cardiovascular complications among Hispanic populations (51).
Since GBD 2023 data were released after our initial analysis, we compared our projections with these independent estimates as a validation check. Our multi-model ensemble forecast 11.6 million (95% PI: 10.7–12.5 million) people with diabetes in the USA-Mexico Border (USMX) region for 2023, which closely aligned with the GBD estimate of 11.4 million (95% uncertainty interval: 10.6–12.3 million). Similarly, our forecast of 13–15 thousand diabetes-related deaths in six Mexico border states (MX) in 2023 was consistent with the GBD estimate of 13–16 thousand deaths.
4.1Public Health and Healthcare System Implications
The diabetes projections for the USA-Mexico border highlighted in this study reveal a pressing public health issue that requires urgent attention. First, it is essential to develop a coordinated binational health policy and allocate resources effectively. Prioritizing initiatives that promote diabetes health by supporting weight management, encouraging regular physical activity, and promoting a balanced diet rich in fiber, fruits, vegetables, and whole grains are warranted.
Second, Mexican border states require immediate strengthening of primary care infrastructure for diabetes screening and early treatment. The projected mortality increase among working-age adults (20-59 years) is expected to generate substantial economic losses through premature death and disability (57). Targeted interventions should include community health worker programs, which have demonstrated effectiveness in Hispanic border communities (58), and improved access to metformin, statins, and antihypertensive medications through public pharmacy networks. Third, US border states must expand healthcare capacity to be ready to accommodate 35-52% prevalence increases in youth populations. This requires school-based screening programs, pediatric endocrinology workforce expansion, and integration of behavioral health services to address psychosocial barriers to diabetes self-management in Hispanic youth (59). The projected burden which is concentrated in the 60-79 age group necessitates geriatric care infrastructure investments and age-appropriate disease management programs. Four, binational coordination mechanisms, including the USA-Mexico Border Health Commission and sister-state partnerships, should consider utilizing these state-specific projections to align policy priorities and share evidence-based interventions. In particular, the divergent epidemiological trajectories we identified suggest opportunities for bidirectional learning, with Mexico potentially adopting similar policies and benefiting from US prevention successes in younger populations while the US could adopt Mexico’s community-based care models that maintain a lower DALY burden in older adults despite resource constraints (55). By taking the above actions, we can significantly enhance diabetes health and outcomes for communities on both sides of the border.
4.2Limitations
Several limitations warrant consideration. First, forecasts assume historical trends continue without accounting for potential policy interventions, technological innovations (e.g. improved medications and treatment), or epidemiological shocks that could alter trajectories (60). Second, the utilized GBD estimates rely on statistical modeling rather than comprehensive surveillance systems in some border states, potentially introducing measurement error (26). Third, the GBD model estimates themselves are not derived from primary surveillance data at the state level, which may limit precision in capturing substate variations. Fourth, our models do not incorporate diabetes type differentiation (i.e. Type 1,and Type 2), fertility and migration patterns, or cross-border healthcare utilization factors affects border regions (61). Fourth, forecasting performance varied substantially by outcome (data type) and geography (state), with prevalence projections in US border states showing wider prediction intervals than mortality forecasts. Finally, our study period ended in 2021, precluding assessment of potential pandemic-related disruptions to diabetes care and screening that may influence future burden (62).
5.Conclusion
Diabetes burden across USA-Mexico border states will increase 30-50% by 2030, with divergent epidemiological trajectories requiring distinct binational responses. Mexican border states demand immediate mortality prevention through expanded treatment access, while US border states must prepare healthcare infrastructure for increases in youth prevalence. These evidence-based projections provide actionable intelligence for resource allocation, healthcare system strengthening, and binational policy coordination to mitigate the escalating diabetes crisis throughout the joint border region. Beyond diabetes, the utilized forecasting framework can be applied to other chronic diseases and cross-border settings to support evidence-based planning and progress toward global health targets.
Supporting information
Data Availability
The datasets generated and analyzed during the current study are available publicly as described in the methods section.
Acknowledgments
The authors would like to thank Hamid Karami, Georgia State University, for sharing his valuable expertise in MATLAB coding through hemera.rs.gsu.edu (interactive online entry point to advanced research computing) which greatly contributed to the success of this research.
Conflict of Interest Statement
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.
Source of Funding
The authors received no funding from an external source.