Bayesian adjustment for trend of colorectal cancer incidence in misclassified registering across Iranian provinces
1Basic and Molecular Epidemiology of Gastrointestinal Disorders Research Center, Research Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
2Department of Biostatistics, Faculty of Paramedical Sciences, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
3Gastroenterology and Liver Diseases Research Center, Research Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
4Department of Infectious Diseases, Istituto Superiore di Sanità, Roma, Italy.
For corresponding: Mohamad Amin Pourhoseingholi, Gastroenterology and Liver Diseases Research Center, Research Institute for Gastroenterology University of Medical Sciences, Tehran, Iran.Abstract
One of the problems in cancer registry of developing countries is misclassification error. This error leads to overestimation and underestimation of cancer rate in different provinces. The aim of this study is to use Bayesian method to correct for misclassification in registering cancer incidence in neighboring provinces of Iran. Incidence data of colorectal cancer were extracted from Iranian annual of national cancer registration reports 2005 to 2008 And Eighteen of the thirty Iranian provinces were selected to enter the Bayesian model and to correct their misclassification. Always a province with appropriate medical facilities is comparable to its neighbor or neighbors. Between years of 2005 and 2008, on the average, 28% misclassification was estimated between the province of East Azarbaijan and West Azarbayjan, 56% between the province of Fars and Hormozgan, 43% between the province of Isfahan and Charmahal and Bakhtyari, 46% between the province of Isfahan and Lorestan, 58% between the province of Razavi Khorasan and North Khorasan, 50% between the province of Razavi Khorasan and South Khorasan, 74% between the province of Razavi Khorasan and Sistan and Balochestan, 43% between the province of Mazandaran and Golestan, 37% between the province of Tehran and Qazvin, 45% between the province of Tehran and Markazi, 42% between the province of Tehran and Qom, 47% between the province of Tehran and Zanjan. Correcting the regional misclassification and obtaining the correct rates of cancer incidence in different regions is necessary for making cancer control and prevention programs and in healthcare resource allocation.
Introduction
Colorectal cancer (CRC) is the third most common cancer among men (10.0% of the total) and the second in women (9.2% of the total) worldwide. Mortality is lower (694,000 deaths, 8.5% of the total) with more deaths (52%) in the less developed regions of the world, reflecting a poorer survival in these regions [1]. In Iran, CRC is the fourth most common type of cancer (the third most common cancer among females and the fifth among males), which accounts for 8.4% of total cancers in the country [2,3].
There is wide geographical variation in incidence across the world; the highest estimated rates being in Australia/New Zealand, and the lowest in Western Africa. Almost 55% of the cases occur in more developed regions. Of course, it is partly because of their advanced diagnostic and registration capabilities [1].
Inflammatory bowel disease, family history of CRC, obesity, dietary habits, smoking, physical inactivity [2,4], and diabetes [5] are well-known risk factors for CRC. Furthermore, environmental risk factors are found to play the most important role in the incidence and development of CRC [4]. So, people living in the same or adjacent areas which are imposed with the same environmental risk factors are expected to have similar cancer incidence rates.
Since cancer is a leading cause of morbidity and mortality worldwide [6,7], population-based and accurate information on its occurrence is extremely valuable as the foundation for identifying risk factors and making purposeful cancer prevention policies [8]. Cancer registries which are known as the main source of epidemiological data, which collects information regarding burden of cancers by recording the incidence, prevalence, survival and mortality of different cancers in a systematic manner [9–11]. Nowadays, their role has expanded into detecting the impact of interventions for cancer control, planning and evaluation of cancer screening programs, and specifying future needs for materials and manpower resources. But the existence of deficiencies in registering individuals information included patient’s permanent residence, primary site of tumor, date of diagnosis, and date of death [8], makes the registered data inaccurate to use in future planning.
In many developing countries such as Iran, health facilities are not distributed evenly throughout the country. So most cancer patients throughout the country prefer to get diagnostic and medical treatment services in capital of the country or in their neighboring facilitate provinces [12]. Some of them don’t mention their permanent residence and are registered in those provinces. It leads to misclassification error in cancer registry data. Misclassification error is the disagreement between the observed value and the true value in categorical data. As the evidence of existence of misclassification error in registering cancer incidence, the expected coverage of new cancer cases in different provinces can be mentioned; that the observed number of incidence is more than expected number in some provinces, and on the other hand, it is less than expected in a neighboring province [13]. It occurs while it is expected that the rate of cancer incidence be about the same in adjacent provinces; since people are adopt very similar lifestyle and traditions and are exposed to same environmental conditions.
There are two approaches to correct for misclassification error; the first approach is validating a small sample of data with rechecking medical records and extending the results to the target population [14]. The second approach is implementing Bayesian method. Bayesian method is a statistical approach that lets us to take our prior evidence into account in the analysis [15] with determining prior information for some of the parameters [16–18].
The aim of this study is to investigate the trend of colorectal cancer provinces of Iran after estimating the misclassification rate in registering cancer incidence by using Bayesian method and re-estimating the incidence rate in each province of Iran.
Material and methods
Registering of cancer reports is obtainable from different references such as pathologies, hospitals, death certificates and etc. National registration programming of cancer cases from Iranian annual of national cancer registration report is extracted during 2005 to 2008 with software which was created by health ministry, until cancer cases are collected, registered and centralized for the past couple of years and is used for data analyses. Hence all new diagnosed cancer cases in temporary information bank are sent from medical universities to ministry of health periodically. Ministry of health after process of duplicating and coding the recorded cancers based on 10th revision of international coding of disease, this information is registered in permanent information bank. And all changes are sent to medical universities on specific duration until permanent information bank of medical universities is equalized with permanent information bank of health ministry. So each medical university has an observed number of cancer cases and also has an expected coverage of cancer cases that are considered to be 100 per 100000 except 2008 that was 113 per 100000. By dividing the observed number to the expected number of cancer cases, the percent of expected coverage for each province is calculated [20].
Since comparison of simple crude rate i.e. comparison of all cancer cases could make false images in total population regardless of age groups, age standardized rates (ASR) is calculated for all provinces of Iran using direct standardization method. The direct method for all provinces of Iran is based on, first selecting a criterion for the population and then calculating the desired outcome rate of this population using age specified rates at each of the two societies. At first, age groups were considered at level of 5 years. World standard population is the most common used standard population (Wi). By dividing number of incident cases to person-years of observations, ASR is calculated per 100000 . Finally for 4 age groups(0-14 years, 15-49 years, 50-69 years and over than70 years old) and for both genders, ASR is calculated in order to compare statistics on cancer internationally
[20–22].
For entering the data to the Bayesian model two vectors Y1 and Y2 were used. Vector for the province, that, has an expected coverage less than 100% with exact ASR and vector for a neighboring province with a more than 100% expected coverage with ASR from the first group incorrectly labeled as being in the misclassified group. Subscript r is the number of covariate patterns for age and sex group combinations. A Poisson distribution was considered for count data and [19,22–24]. Y1 Y2
Y1 ≈ Poisson(Pi μi1) and Y2 ≈ Poisson(Pi μi1) the joint distribution of the count data Y1 and Y2 is proportional to:
An informative beta prior distribution was assumed for θ as the probability of a data from the first group incorrectly registered in the misclassified group; so θ ~ Beta(a, b). For selecting prior value for the parameters of beta distribution, the calculated expected coverage for the medical university which has a less than 100% expected coverage was used as b and a was calculated with subtracting b from 100. Thus a/(a + b) which is the expectation of beta distribution converges to the misclassified rate. Variable U with binomial distribution, i.e. Ui | Y1, Y2, θ, λ1, λ2 ~ Binomial(Yi2, Pi) that was considered as the number of events from the first group that are incorrectly registered in the misclassified group. Now if theta, Y1, Y2 to be unknown; we have:
But since Y1, Y2 have known values of ASR on two neighboring provinces, then just theta is unknown and with employing a latent variable approach to correct the misclassification effect according to Paulino et al. [26,27], Liu et al. [28] and Stamey et al. [19] using a Gibbs sampling algorithm, the posterior appears in the following form:
After estimating the misclassification rate between each two neighboring provinces, the rates of colorectal cancer incidence for each province were re-estimated and the trend of colorectal cancer were carried out during 2005 to 2008. All analyses were performed using R software version 3.3.1.
Results
Registered cases of colorectal cancer have been included in the study for all provinces in Iran from 2005 to 2008. ASR of CRC incidence for men was 8.02 per 100,000 population (2255 cases) in 2005, whereas that year for women 7.4 per 100,000 (1801 cases). In over time, ASR of CRC incidence for men reached 12.7 per 100,000 population (3527 cases) in 2008 and for women to 11.12 per 100,000 (2658 cases) in the same year. The trend of CRC from 2005 to 2008 for both sexes is shown in Fig 1.
Among 30 provinces of Iran, 18 provinces where the number of cancer cases varied from their expected number were selected to correct the misclassification error in the register of CRC incidence in the neighbor provinces, based on the percentage of expected cancer coverage.
For example, the reported percentage of CRC expected coverage for Fars province as a province with suitable medical facilities and services was 120.8% in 2008. it means that Fars province have covered 20.8% of the new cases more than expected, while Hormozgan, which is adjacent to Fars, has a 19% expected coverage of cancer incidence, indicating clear misclassification in registering cancer cases. The expected coverage for all provinces in Iran between 2005 and 2008 is reported in Table 1. Also the estimated misclassification rate for all provinces in 2005-2008 is reported in Table 2.
For example by using the Bayesian method, misclassification rate was estimated 58% between Fars and Hormozgan in 2008. So, after Bayesian correction, ASR and number of cancer incidence decrease for Fars province and increase for Hormozgan province. ASR and number of cancer incidence, before and after Bayesian correction from 2005 to 2008 are reported in tables 3 and 4.
Discussion
It is obvious, neighboring provinces due to having the same food habitation, lifestyle and locating in the same climate, have the same health outcomes [13]. But sometimes when analyzing registered data, it is observed that the neighboring provinces not only do not have the same outcomes but also have inconsistent. This situation implies that there is misclassification in registered data. This problem is a notable matter in medicine issues which may results to deflection in health programing and health resources allocating. Such deflection would make irrecoverable damage in national scale. The aim of the present study was to help to reduce misclassification error in registered colorectal cancer data in Iran. In Iran, facility in accessing health resources in welfare provinces at first and secondly lack of health facilities in their neighboring provinces are elements creating misclassification error. Fortunately, some studies have been conducted in Iran in order to eliminate the misclassification errors for mortality and morbidity registered cancer data in the case of Liver [29], Gastric [22], and colorectal cancers [23]. Since the above studies had re-estimated data and also had produced valid data, employing their results may be more reliable. According to the result of our research, there was a non-ignorable estimated misclassification rate among adjacent provinces. The highest estimated misclassification parameter, was belong to North Khorasan, Hormozgan, and Sistan which are in east and south of Iran. So the real rates of CRC in those provinces are higher than the rates that are reported by cancer registry system.
On the contrary, in studies that are used cancer registry data ignoring the existence of misclassification error, it is reported that, the highest incidence rates of CRC in Iran were found in the central, northern, and western provinces; and the southwest provinces of Iran had the lowest incidence rates of CRC in the country! [2]. So, ignoring the misclassification error in registry data, leads to a wrong image of distribution of CRC incidence across the country. Expected cancer coverage revealed that from 30 provinces, 18 provinces need to misclassification correction. These provinces are those which are different in economic situation and there are some points in them which are welfare and probably patients for better health care, refers to those welfare places, so they have more referring people than their capacity. On the other hand, some provinces due to less facility, have less referring patients. Table 2 is indicating how the data of some provinces are registered in their adjacent locations. For example, some neighboring Tehranian patients such as Qom’s patients, were referred to Tehran.
Identifying the exact distribution of disease in different areas is a good manner for finding the geographic pattern of disease and causations, assessing the influencing factors on disease incidence [30,31], and quantifying the potentials for disease control and prevention [32,33]. But usually spatial analysis is used for this purpose which is based on registered data while existence of misclassification is often ignored. In spatial analysis, the morbidity or mortality rates for each province are combined with locality’s information for the same province and so the result may lead to an integrated geographical map. This type of maps is helpful in comparison between different provinces in aspect of rate of disease incidence or probable risk factors [34]. For accessing such a goal, we have prepared geographical map for evaluating incidence distribution of colorectal cancer registered data in before and after misclassification correction in Fig 2. Fig 2 revealed that after correction the southern provinces have high incidence rate, while in the previous studies which had been not regarded misclassification, southern provinces had low incidence rate [35].
The maps of present study also revealed that a considerable changes happened in some provinces respect to before correction status. Thus major differences in the incidence of CRC, while it is expected that the incidence of cancer be alike in adjacent provinces, can be justified by existence of misclassification error in registering permanent address of patients who are diagnosed in neighboring facilitate provinces. It leads to overestimation of CRC rate in some provinces and underestimation of its rate in some neighboring provinces.
For future researches, to recognizing high risk spatial clusters, using our colorectal cancer valid data, is suggested.
In conclusion, proper planning for cancer control and prevention, and allocating healthcare facilities to different areas, requires an increase in the quality and accuracy of registering system in different provinces, and correcting the existed deficiencies especially misclassification error in registering patient’s permanent residence. It is in need of enhancing hardware and software resources, training more of educated staff in different sectors of cancer registry program, and implementing the opinions of expert researchers in medicine, biostatistics, and epidemiology [36]. In the absence of valid data, Bayesian method can be adopted as a fast and cost effective method to correct the regional misclassification error.
Acknowledgement
Research reported in this publication was supported by Elite Researcher Grant Committee under award number 958696 from the National Institutes for Medical Research Development (NIMAD), Tehran, Iran