Variables associated with owner perceptions of the health of their dog: further analysis of data from a large international survey
Institute of Life Course and Medical Sciences, University of Liverpool, Liverpool, UK
*Corresponding author: E-mail: ajgerman@liverpool.ac.ukAbstract
In a recent study (doi: 10.1371/journal.pone.0265662), associations were identified between owner-reported dog health status and diet, whereby those fed a vegan diet were perceived to be healthier. However, the study was limited because it did not consider possible confounding from variables not included in the analysis. The aim of the current study was to extend these earlier findings, using different modelling techniques and including multiple variables, to identify the most important predictors of owner perceptions of dog health.
From the original dataset, two binary outcome variables were created: the ‘any health problem’ distinguished dogs that owners perceived to be healthy (“no”) from those perceived to have illness of any severity; the ‘significant illness’ variable distinguished dogs that owners perceived to be either healthy or having mild illness (“no”) from those perceived to have significant or serious illness (“yes”). Associations between these health outcomes and both owner-animal metadata and healthcare variables were assessed using logistic regression and machine learning predictive modelling using XGBoost.
For the any health problem outcome, best-fit models for both logistic regression (area under curve [AUC] 0.842) and XGBoost (AUC 0.836) contained the variables dog age, veterinary visits and received medication, whilst owner age and breed size category also featured. For the significant illness outcome, received medication, veterinary visits, dog age and were again the most important predictors for both logistic regression (AUC 0.903) and XGBoost (AUC 0.887), whilst breed size category, education and owner age also featured in the latter. Any contribution from the dog vegan diet variable was negligible.
The results of the current study extend the previous research using the same dataset and suggest that diet has limited impact on owner-perceived dog health status; instead, dog age, frequency of veterinary visits and receiving medication are most important.
Article notes
Competing Interest Statement
AJG is an employee of the University of Liverpool, but his position is financially-supported by Royal Canin. AJG has also received financial remuneration and gifts for providing educational material, speaking at conferences, and consultancy work, all unrelated to the current study. Royal Canin had no involvement in any aspect of the current work including study design, data analysis, drafting the manuscript or the decision to submit the work for publication. RBJ is an employee of the University of Liverpool, whose salary is financially-supported by the Higher Education Funding Council for England. RBJ also receives a stipend from United Kingdom Research and Innovation for work as a panellist for the Biotechnology and Biological Sciences Research Council.
Summary of Updates:
Introduction
There is an increasing interest from owners in feeding unconventional diets, such as those utilising uncooked (so-called ‘raw’ diets) or plant-based (either vegetarian or vegan diets) ingredients. In a recent UK survey, 7% and approaching 1% of pet owners reportedly fed their dog either a raw or vegan diet, respectively [1]. Veganism is an increasingly popular food choice in humans, with recent surveys indicating 2-3% [2], 4% [3] and 3% [4] of people in the UK, Europe and USA, respectively, to be vegan. Although still uncommon (<1% of owners), use of plant-based (vegan or vegetarian) diets for dogs is increasing in popularity [1], not least amongst owners who are vegan themselves [5]. Interestingly, the popularity of feeding raw diets is also increasing, with the same survey suggesting that 7% of the UK dog-owning population use this method [1], a proportion similar to that reported in recent international study, where 9% of dog owners exclusively fed raw food [6]. The main reason that owners cite for considering a plant-based diet is concern about animal welfare [5], whilst owners who feed raw diets believe these to be more ‘natural’ and healthier than commercial food despite an absence of evidence to support this [7]. Despite their increasing popularity, concerns remain regarding the safety of feeding unconventional foods; for both vegan [5,8] and raw [9–11] food, in terms of nutritional adequacy. For raw food, there are additional concerns about contamination, both with pathogenic bacteria and bacteria resistant to multiple antimicrobials [11–13].
In recent studies, the safety of feeding unconventional diets for dogs has been examined, albeit using owner reports of health from questionnaires or surveys. In one study of 16,475 households that fed raw food, only 39 (0.2%) recalled a member of the household acquiring a pathogenic infection during the time they were using the food [14], perhaps, suggesting that risks to owners from bacterial contamination of raw food are uncommon. The health of dogs being fed unconventional diets has also been assessed. In one survey of 1,413 owners conducted across North America, those feeding plant-based diets reported fewer health conditions in their dogs and a longer lifespan, compared with owners of dogs fed a meat-based diet [6]. In a second study [15], owner-reported health of dogs being fed various diet types was assessed. Compared with dogs fed a conventional diet, owners that fed either raw food or a vegan diet reported better health. However, besides being reliant on owner-reported health, rather than more objective health measures, no account was taken of possible confounding from other variables. To the authors’ credit, the original study data were made available for use by other researchers (https://rebrand.ly/2020/pet-food-consumers-study), which not only included the diet and health data they analysed, but also other metadata including owner (e.g., country, career, education, age and gender) and animal (e.g., age, sex, neuter status and breed) variables, and healthcare-related variables (e.g., veterinary visits, receiving medication, use of therapeutic diet). Therefore, the aim of the current study was to extend the earlier findings using different statistical techniques to create models that best predicted owner perceptions of health, and to identify the relative importance of the variables that contributed to the final model. This included the use of binary logistic regression and machine learning predictive modelling using a scalable tree boosting system (XGBoost [16]).
Materials and Methods
Study design
A secondary analysis was conducted using data from a recent international survey of owner opinions on the health of their dog [15]. The open-access dataset from this study is accessible at https://rebrand.ly/2020/pet-food-consumers-study. The reason for conducting this further analysis was that many variables gathered were not analysed statistically, and that the effect of possible confounding was not adequately investigated, both of which were limitations acknowledged by the study authors [15]. Further, we were particularly interested in exploring associations between a wider range of variables and owner perceptions of the health of their dog.
Dataset
Full details of survey design, methodology, information collected and category definitions have been reported previously [15]. Briefly, the original survey was designed using an online survey platform (https://www.onlinesurveys.ac.uk), with owners being asked to record details about their dog (including demographic data, main diet type and health characteristics) and themselves. Owner information used for the current study included location (country or continental region), setting (urban, rural or a mix of urban and rural), education level (qualifications achieved), household income, whether their occupation was in an animal-related career (including veterinarian, veterinary technician or nurse, animal breeder, animal trainer or pet industry worker), age, gender and diet. Animal information used for the current study included: age, sex, neuter status, breed size category (e.g., toy, small, medium, large and giant) and main diet type (including conventional, raw, vegetarian, vegan, in vitro meat, insect, fungal or algal). We also used responses from owners about the healthcare their dog received including number of veterinary visits in the last 12 months, whether the dog had received medication (besides those used in preventive care e.g. routine vaccinations and endo- or ecto-parasite treatments) in the last 12 months, and whether the dog was on a therapeutic diet; when dogs were fed a therapeutic diet, owners were asked to base the information on their dog’s diet on their previous diet [15]. Finally, in the original survey, owners were also asked their opinion about the health of their dog in the last 12 months, with possible responses including: ‘healthy’, ‘generally healthy with minor or infrequent problems’, ‘significant or frequent problems’, ‘seriously ill’ or ‘unsure’. Since we were primarily interested in owner perceptions of health, we used this information to create our outcome variables, as described below.
Statistical methods
Statistical software
All data pre-processing and statistical analyses were conducted using an online open-access statistical language and environment (R, version 4+) [17], with several packages including: ‘aod’ version 1.3.2 [18], ‘broom’ version 0.8.0 [19], ‘car’ version 3.0.13 [20], ‘caret’ version 6.0.93 [21], ‘cluster’ version 2.1.4 [22], ‘clusterSim’ version 0.51.3 [23], ‘corrplot’ version 0.92 [24], ‘data.table’ version 1.14.2 [25], ‘dbscan’ version 1.1-11 [26], ‘DescTools’ version 0.99.45 [27], ‘dplyr’ version 1.0.9 [28], ‘factoextra’ version 1.0.7 [29], ‘fmsb’ version 0.7.3 [30], ‘foreign’ version 8.82 [31], ‘ggforce’ version 0.4.1 [32], ‘ggplot2’ version 3.3.6 [33], ‘ggthemes’ version 4.2.4 [34], ‘glmtoolbox’ version 0.1.3 [35], ‘heatmaply’ version 1.4.0 [36], ‘Hmisc’ version 4.7-0 [37], ‘IHW’ version 1.26.0 [38]. ‘mltools’ version 0.3.5 [39], ‘pROC’ version 1.18.0 [40], ‘RColorBrewer’ version 1.1-3 [41], ‘readxl’ version 1.4.0 [42], ‘reshape2’ version 0.8.9 [43], ‘ROCR’ version 1.0-11 [44], ‘rsample’ version 1.1.0 [45], ‘rstatix’ version 0.7.0 [46], ‘showtext’ version 0.9-5 [47]; ‘smotefamily’ version 1.3.1 [48], ‘stringr’ version 1.4.0 [49], ‘tidyverse’ version 1.3.1 [50], ‘uwot’ version 0.1.16 [51] and xgboost version 1.6.0.1 [16]. Statistical reports of all data pre-processing and analysis are included in the supplementary information (S1-S9 Files), whilst all code is available online (https://github.com/RichardBJ/CanineHealth23).
Initial dataset pre-processing
Initially, data were imported into the statistical environment using the ‘readxl’ package [42], and then manipulated using the ‘dplyr’ package [28]. Columns were renamed for clarity and to improve compatibility with use in R (see below). Counts of all categorical data were examined, and categories with small groups were either removed or combined with others to exclude missing data or to reduce dataset imbalance, as described below. The variables included in the dataset were owner variables and animal variables (collectively owner-animal metadata), as well as healthcare variables and health outcomes variables.
Data pre-processing occurred in two stages: in the first stage (prior to Uniform Manifold Approximation and Projection; UMAP; see below), the original dog sex variable (male sexually intact, male castrated, female sexually intact, female spayed) was separated into two binary variables, dog sex (male, female) and neuter status (yes, no). In addition, data were excluded from dogs where age was reported as ‘unsure’ (n=3) or were aged <1 year (n=26), to ensure a focus on adult dogs.
To facilitate the modelling studies (logistic regression and XGBoost), further data pre-processing steps were undertaken after UMAP. For owner variables, pre-processing included combining categories with small numbers in the location variable (e.g., ‘Africa’ [n=7], ‘Asia’ [n=25], ‘other’ [n=11] and ‘South America’ [n=47]) into a new ‘other’ category, and also removing the ‘other’ category (n=15) from the settings variable. The owner education variable was simplified by combining the ‘basic’ (n=15) and ‘high school’ (n=478) categories and combining the ‘PG’ (n=515) and ‘PhD’ (n=90) categories. A binary animal career variable was created from the career variable by combining the ‘veterinarian’ (n=125), ‘veterinary technician or nurse’ (n=54), ‘animal breeder’ (n=17), ‘animal trainer’ (n=140) and ‘pet industry worker’ (n=205) categories into a ‘yes’ category, whilst the remaining option (‘none of the above’ [n=2,137]) formed the ‘no’ category. The owner income variable was simplified by removing the ‘prefer not to answer’ category (n=184) and ordered, whilst the owner age variable was simplified (by combining the ‘18-19y’ [n=19] and ‘20-29y’ [n=391] categories into a new ‘<30y’ category and combining the ‘60-69y’ [n=364] and ‘>70y’ [n=85] categories into a new ‘>60y’ category) and then ordered. Given small numbers, the ‘prefer not to answer’ (n=17) and ‘other ‘(n=4) categories were removed from owner gender variable, and the ‘other’ category (n=18) was removed from the owner diet variable. Further, binary variables were created for owners who reported consuming a vegan diet (yes vs. no) and owners who reported consuming either a vegan or vegetarian diet (yes vs. no).
For animal variables, in addition to having a continuous dog age variable, an ordinal variable was also created with data grouped into quintiles (1-3y, 3-5y, 5-7y, 7-9y, 9-20y). For the dog diet variable, the ‘unsure’ (n=8), ‘insect-based’ (n=5), ‘meat-based (lab-grown)’ (n=6) and ‘mixture’ (n=13) categories were all removed since numbers were again small. Further, binary variables were created for dogs consuming a vegan (yes vs. no), vegan or vegetarian (yes or no) or a raw (yes vs. no) diet, and for dogs in the ‘giant breed’ category (yes vs. no) of the breed size category variable.
For healthcare variables, the ‘unsure’ category (n=13) was removed from the veterinary visits variable, and two ordinal variables were created; the first comprised 5 categories (‘0 visits’, ‘1 visit’, ‘2 visits’, ‘3 visits’, ‘4+ visits’), whilst the second comprised 4 categories (‘0 visits’, ‘1 visit’, ‘2 visits’, ‘3+ visits’ [n.b., ‘3 visits’ and ‘4+ visits’ categories combined]).
For our outcome variables (owner’s perception of the health of their dog), we first removed data from owners who reported that they were ‘unsure’ (n=5). From this, two binary health status variables were created: the first was an any health problem variable, where dogs classified as ‘healthy’ were classified as ‘no’, and all remaining categories (‘generally healthy with minor or infrequent problems’, ‘significant or frequent problems’ and ‘seriously ill’) were classified as ‘yes’; the second was a significant illness variable, whereby dogs classified either as ‘healthy’ or ‘generally healthy with minor or infrequent problems’ were classified as ‘no’, and those classified as having ‘significant or frequent problems’ or being ‘seriously ill’ were classified as ‘yes’.
As a final pre-processing step, a decision-maker (primary vs. other) variable was created from data about the role owners played choosing a diet for their dog; for this, the ‘primary decision-maker’ category was classified as ‘primary’, whilst the other two categories (‘play no role’, n=15; ‘play some lesser role’, n=96) were classified as ‘other’. Given concerns about the reliability of information from owners who were not primary decision-makers and healthcare received, its effect on the outcome variables used in modelling studies (both binary illness variables) was first assessed by logistic regression, before the other modelling studies, as described below.
Binary logistic regression
Binary logistic regression models were created using the ‘glm’ function in R. Both binary illness variables (any health problem and significant illness) were used separately as outcome variables, whilst all owner, animal and healthcare variables were tested as predictor variables. We used a combination of simple and multiple logistic regression. Simple logistic regression analyses were initially performed using data from all owners (irrespective of decision-maker status), with results shown in the supplementary information (S1 and S2 Tables). However, these models revealed concerns over the reliability of data from owners who were not primary decision-makers (see below) and, as a result, all remaining logistic regression analyses were instead conducted with the primary decision-maker dataset. For these analyses, the dataset was first randomly divided into training (75%) and test (25%) datasets, using the ‘sample’ function in R, a step that was necessary to facilitate model validation by receiver operating characteristic (ROC) curves (see below).
Simple logistic regression models created with the primary decision-maker dataset enabled unadjusted associations between each outcome variable and single predictor variables to be determined. Some variables had been coded in different ways, as described above; for example, dog age was coded as both a continuous variable and an ordinal variable with 5 categories, with the continuous dog age variable also being tested with and without basis (b-)splines; further, veterinary visits were coded as both a 4- or 5-category variable, whilst dog diet was coded both as a 4-category variable [conventional, raw, vegetarian and vegan] or as separate binary variables for raw diet, vegan diet or a combined vegan-vegetarian binary variable [see above]. To determine the coding approach that was most appropriate for such variables, performance of competing models was assessed using the Bayesian Information Criterion (BIC, see below) [60]; the coding approach that fitted the data best was selected for use in the initial multiple regression modelling as long as model assumptions were met (e.g., independence of errors, the requirement for a linear relationship between the logit of the outcome and continuous predictor variables). For example, the vegan diet variable performed better than the vegan-vegetarian diet variable, the continuous dog age variable performed better than the ordinal dog age variable, albeit that b-splines (utilising boundary knots and an internal knot at the median value [6 years]) were required to ensure that model assumptions were met. Further, for the significant illness outcome variable, the giant breed predictor variable performed better than the breed size category variable, and there were also some issues with model convergence for the latter. Models containing two related predictor variables, with and without interaction terms, were also tested when these were clinically relevant (e.g., between dog sex and neuter status, between owner and dog diet categories and between location and setting).
In the first step of the multiple logistic regression modelling, two multiple regression models were created for each outcome variable (any health problem and significant illness); all predictor variables were included in the first of these models (all-variable model), whilst only owner and animal variables were included in the second model (owner-animal metadata model). The performance of all-variable and owner-animal metadata models were then compared with each other and with the best-fit models that were created subsequently. For these best-fit models, we used a supervised, stepwise approach starting with all variables. This initial model was manually refined in both a backwards and forwards stepwise fashion, with the BIC being used to select the model within the same family with the best generalisability (a measure of its goodness of fit compared with its complexity) [60]. Using this approach, variables could be added or removed until the model with the smallest BIC was found, according to previously published rules [61], whereby model generalisability was deemed to be superior if the BIC of the new model was less than the previous model by at least 2 units.
After selecting a best-fit model and, given the importance of diet in the original work [15], both the dog diet vegan and owner diet vegan variables were separately added back, to determine their effect on model performance.
Results are reported as estimates of the regression coefficients (β) with the associated standard error (SE), and with odds ratios (OR) with the associated 99% confidence intervals (99%-CI). Influential datapoints were identified and assessed using Cook’s distance and, in most models, none were identified. In the occasional cases where outliers were identified, we checked them in the original data to make sure there were no obvious errors (e.g., unrealistic results such as dog age >20 years); no such errors were found. Given that we were using data from a secondary source, had no further means of verifying them (e.g., by cross-checking against the original questionnaire response), we decided not to remove them as there was no valid justification to do so (e.g., a typographical error created when the original data were entered). Possible multicollinearity in all models was assessed using the generalised variance inflation factor (GVIF) and GVIF(1/(2×Df); these were deemed to be acceptable when all values were <4 and <2 for GVIF and GVIF(1/(2×Df), respectively [62]. If necessary, multicollinearity was resolved by removing the variable with the greatest GVIF. Goodness of fit was tested by a visual inspection of observed and expected results and the Hosmer and Lemeshow test for large datasets. In addition, the proportion of variance explained in the overall model was assessed by the coefficient of determination (pseudo-R2) based on the method reported by Nagelkerke [63]. Further, the log-odds of the outcome variable and dog age was plotted to enable visual assessment of whether the relationship was linear, and this was further assessed with the Box-Tidwell test. Finally, prediction accuracy was assessed by creating ROC curves, using the ROCR package [44], and then calculating their area under the curve (AUC). Values could range from 0 to 1; a model that performed no better than chance would have an AUC of 0.5, and models predicting better than by chance would have AUC >0.5, with an AUC of 1.0 suggesting perfect prediction. The test dataset was used to generate the ROC curves and AUC for models using data from owners who were primary decision-makers; however, because the all-owner dataset was not subdivided, AUC are based in ROC curves from the original dataset.
Machine learning predictive modelling
A scalable tree boosting system (XGBoost; “eXtreme Gradient Boosted” tree) was implemented in the statistical environment [16]. Data were pre-processed as described above and, to facilitate XGBoost, ranked variables used for the Kendall’s tau correlation (as described above) were coerced to a numeric datatype, using one-hot encoding where necessary, and then subdivided 70:30 into training and test sets. A model was then built on the training data without any involvement of the test set. After pre-processing but before training, training data were further divided into training and validation subsets (70:30), and the training subset augmented as per standard practice for machine learning [64]. We trialled two augmentation methods: SMOTE [48] and a similar method where underrepresented classes were duplicated, but with simple addition of signal noise. Final models used the latter given that it performed better than SMOTE. Test data were left untouched. The model was then trained using the augmented training data with area-under the curve (AUC) as the evaluation metric, against the validation training subset. Finally, the model was built and saved to file. Receiver operating characteristic curves were calculated and plotted with the saved model against the previously unseen test dataset using the packages ‘pROC’ [40] and ‘ROCR’ [44]. Prediction accuracy was estimated by calculating the ROC AUC in a similar manner to that described above for logistic regression. Different measures of the contribution of the variables (‘features’) included in the final models were also calculated and depicted graphically. These included ‘feature importance’ (a metric of the fractional contribution of each feature to the model, based on the total gain from including each feature), ‘cover’ (a metric of the number of observations related to this feature in the model) and ‘frequency’ (a metric that represents the relative number of times a feature has been used in trees). Since these metrics are normalised, the sum of scores for all features in the model is 1.0.
Results
Summary of healthcare variables
Details of healthcare variables are also given in Table 1. Most dogs had visited their veterinarian at least once in the last year (all owner dataset 1,914 [82.4%]; primary decision-maker dataset 1,812 [81.9%]), whilst smaller proportions of dogs had received medication (all owner dataset 936 [40.3%]; primary decision-maker dataset 880 [39.8%]) and or had been switched to a therapeutic diet (all owner dataset 111 [4.8%]; primary decision-maker dataset 100 [4.5%]).
Owner, animal and healthcare variables associated with the any health problem binary variable
Simple regression
Within the any health problem binary variable, 894 (38.5%) and 1429 (61.5%) dogs were classified as ‘yes’ and ‘no’, respectively. Simple binary logistic regression was performed first to assess the unadjusted effects of each predictor variable separately using data from primary decision-makers only (Table 2). Broadly similar results were obtained when logistical regression was performed on the dataset from all owners (irrespective of decision-maker status; S1 Table). For example, estimates of the regression coefficients (β) for the dog diet vegan variable was −0.43 (SE 0.131) using data from all owners and β −0.30 (SE 0.159) using data from primary decision-makers only.
Multiple regression
Next, we created two multiple regression models, again on data from primary decision-makers, with any health problem binary as the outcome variable and that either contained all predictor variables or just the owner-animal metadata. For the all-variable model (S3a Fig), the variables with the strongest associations with the outcome variable were dog age (both the ‘1-5y’ and ‘6-20y’ categories), received medication, veterinary visits (the ‘1 visit’, ‘2 visits’, ‘3 visits’ and ≥4 visits’ categories) and switched to a therapeutic diet. In terms of model performance, pseudo-R2 was 0.439, BIC was 1836 and AUC on the test dataset was 0.852 (S5 Fig). For the owner-animal metadata model (S3b Fig), the variables with the strongest associations with the outcome variable were owner age (‘50-59y’ and ‘≥60y’ categories), dog age (‘6-20y’ category), breed size category (‘giant breed’ category) and dog diet (‘vegan diet’ category). However, performance of this model was considerably worse in terms of generalisability (BIC 2244), the proportion of variance explained (pseudo-R2 0.150) and prediction accuracy (AUC on the test dataset: 0.668).
Supervised backwards and forwards stepwise regression was then performed to create a best-fit multiple regression model for the any health problem binary outcome variable, starting with the model containing all variables. After refinement by backwards and forwards stepwise elimination, the best-fit model contained 4 variables: dog age, veterinary visits, received medication and switched to a therapeutic diet (Fig 3a; model 1, S5 Table). Compared with the all-variable model, generalisability was markedly improved (BIC 1656), with only slight reductions in both the amount of variability explained (pseudo-R2 0.420) and prediction accuracy (AUC for the test dataset:0.842).
As an additional multiple logistic regression step, the effect of adding the dog diet vegan variable to this best-fit model was assessed (Fig 3b; model 2; S4 Table); generalisability was worse (BIC 1660), whilst the amount of variance explained (pseudo-R2 0.422) and prediction accuracy (AUC for the test dataset 0.845) were similar. Equivalent results were obtained when the owner diet vegan diet variable was instead added to the final best-fit model (BIC 1660, pseudo-R2 0.421, AUC 0.844; Fig 3c; model 2; S4 Table).
Machine learning predictive modelling with XGBoost
The ability of owner, animal and healthcare variables to predict the any health problem variable was further assessed using machine learning predictive modelling using the XGBoost algorithm. Two separate models were created using data from primary decision-makers; the first contained all variables, whilst the second was a reduced model containing just the owner-animal metadata (i.e., without healthcare variables). Based on ROC analysis, the prediction accuracy of the all-variable model was reasonable (AUC 0.836; Fig 4a). The XGBoost procedure also enables the most important variables contributing to the prediction to be identified, for the all-variable model (Fig 4b); the variables that most strongly predicted the any health problem outcome variable were (in order, based on the sum of the 3 metrics): received medication (importance 0.393, cover 0.129, frequency 0.065), dog age (importance 0.175, cover 0.182, frequency 0.188) and veterinary visits (importance 0.153, cover 0.181, frequency 0.155), followed by owner age (importance 0.048, cover 0.083, frequency 0.106) and then breed size category (importance 0.046, cover 0.070, frequency 0.087). Other variables followed these including education, dog sex, location, and dog diet; for the latter variable, dog diet raw was 10th in order of importance (importance 0.015, cover 0.026, frequency 0.032), whilst dog diet vegan was 15th in order of importance (importance 0.008, cover 0.021, frequency 0.013).
The order of importance of predictor variables was broadly similar in the reduced (owner-animal metadata) model although, based on ROC analysis, prediction accuracy was poor (AUC 0.674, Fig 5a). In this reduced model, dog age was overwhelmingly most important (importance 0.261, cover 0.196, frequency 0.157), followed by owner age (importance 0.115, cover 0.123, frequency 0.141, education level (importance 0.095, cover 0.089, frequency 0.124), breed size category (importance 0.091, cover 0.099, frequency 0.108) and neuter status (importance 0.060, cover 0.062, frequency 0.047). Other variables including dog diet, urban setting, location and owner income and followed these (Fig 5b); for the dog diet variable, dog diet raw was 7th in order of importance (importance 0.054, cover 0.048, frequency 0.060), whilst dog diet vegan was 12th in order of importance (importance 0.039, cover 0.045, frequency 0.031).
Results were broadly similar when machine learning predictive modelling with XGBoost was used to create both all-variable and reduced (owner-animal metadata) models for the any health problem outcome measure using the all-owner dataset (S5 and S6 Figs).
Owner, animal and healthcare variables associated with the significant illness binary variable
Simple regression
Within the significant illness binary variable, 122 (5.5%) and 2089 (94.5%) dogs were classified as ‘yes’ and ‘no’, respectively. Results of the simple binary logistic regression stage, to assess the unadjusted effects of each predictor variable separately using data from primary decision-makers only, are shown in Table 3, whilst the results of equivalent analyses on the dataset from all owners (irrespective of decision-maker status) are shown in S2 Table. Once again, estimates of the regression coefficients were broadly similar from both data sets. For example, estimates of the regression coefficients (β) for the dog diet vegan variable was −0.46 (SE 0.309) using data from all owners and β −0.44 (SE 0.402) using data from primary decision-makers only.
Multiple regression
Next, we created both an all-variable and an owner-animal metadata model, using data from primary decision-makers, with significant illness binary as the outcome variable. For the all-variable model (S4a Fig), the variables with the strongest associations with the outcome variable were dog age (‘6-20y’ category), giant breed, and received medication. In terms of model performance, pseudo-R2 was 0.403, BIC was 679 and AUC on the test dataset was 0.890. For the owner-animal metadata model (S4b Fig), the variables with the strongest associations with the outcome variable were dog age (‘6-20y’ category) and giant breed. Once again, the performance of this model was considerably worse in terms of generalisability (BIC 799), the proportion of variance explained (pseudo-R2 0.124) and prediction accuracy (AUC on the test dataset: 0.692).
Supervised backwards and forwards stepwise logistic regression was again used to create a best-fit multiple regression model for the significant illness outcome variable, starting with all variables. After refinement by backwards and forwards stepwise elimination, the best-fit model contained 3 variables: dog age, veterinary visits and received medication (Fig 6a; model 1, S5 Table). Compared with the all-variable model, generalisability was markedly improved BIC 517), with better prediction accuracy (AUC for the test dataset: 0.903), and only a small decreased in the amount of variability explained (pseudo-R2 0.342).
As a further multiple logistic regression step, the effect of adding the dog diet vegan diet variable to this best-fit model was assessed (Fig 6b; model 2; S6 Table); generalisability was worse (BIC 524), whilst the amount of variance explained (pseudo-R2 0.342) and prediction accuracy (AUC for the test dataset 0.904) were similar. Equivalent results were obtained when the owner diet vegan diet variable was instead added to the final best-fit model (BIC 524, pseudo-R2 0.342, AUC 0.903; Fig 6c; model 2; S6 Table).
Machine learning predictive modelling with XGBoost
Using machine learning predictive modelling with XGBoost, and data from primary decision-makers, all-variable and reduced (owner-animal metadata) models were created (S2 File) for the serious illness outcome variable. Based on ROC analysis, the prediction accuracy of the all-variable model was reasonable (AUC 0.887; Fig 7a). In order (Fig 7b), the variables that most strongly predicted the significant illness outcome variable were: veterinary visits (importance 0.530, cover 0.473, frequency 0.285), dog age (importance 0.182, cover 0.219, frequency 0.302) and received medication (importance 0.175, cover 0.190, frequency 0.143), followed by breed size category (importance 0.050, cover 0.034, frequency 0.079), education (importance 0.016, cover 0.020, frequency 0.050), owner age (importance 0.012, cover 0.014, frequency 0.036) and dog sex (importance 0.009, cover 0.010, frequency 0.034). Other variables followed these including location, urban setting, owner sex, animal career and therapeutic food. For the dog diet variable, dog diet raw was 17th in order of importance (importance 6.14 x10-5, cover 4.09 x10-5, frequency 1.55 x10-4) relative importance), whilst the dog diet vegan variable was of insufficient importance to be included in the model (i.e., relative importance less than that of the dog diet raw).
As with the any health problem binary, ROC analysis suggested that prediction accuracy of the reduced (owner-animal metadata) model was poor (AUC 0.689, Fig 8a). Once again, dog age (importance 0.304, cover 0.362, frequency 0.224) was the most important predictor, followed by breed size category (importance 0.195, cover 0.226, frequency 0.186), owner age (importance 0.091, cover 0.082, frequency 0.125), urban setting (importance 0.077, cover 0.066, frequency 0.086) and education (importance 0.072, cover 0.060, frequency 0.094); other variables followed these including dog sex, dog diet, location, animal career and neuter status (Fig 8b); for the dog diet variable, dog diet raw was 7th in order of importance (importance 0.059, cover 0.042, frequency 0.047), whilst dog diet vegan was 12th in order of importance (importance 0.018, cover 0.014, frequency 0.016).
Results were broadly similar when machine learning predictive modelling with XGBoost was used to create both all-variable and reduced (owner-animal metadata) models for the significant illness outcome variable using the all-owner dataset (S7 and S8 Figs).
Discussion
In the current study, we examined associations between owner perceptions of the health of their dog and a range of owner-related, animal-related and healthcare variables. Predictor variables that were most strongly associated with our outcome variables were dog age and healthcare variables (including number of veterinary visits and receiving medication), whilst the contribution from other variables was more limited, and the effect of dog diet was negligible. We utilised data from a previous study, where the opinions of owners about the health of their dog was gathered by questionnaire [15]. The primary aim of that study was to examine associations between owner opinions of dog health status and the type of diet that owners predominantly fed, with the key finding being that health status was positively associated with feeding either a vegan or raw diet, compared with other diet types (including conventional and vegetarian). Given the focus of that study, other information gathered in the questionnaire was not assessed and no account was taken of possible confounding amongst variables. Therefore, to extend the findings of the previous work, additional owner, animal and healthcare variables were studied. Previous work has reported associations between such variables and owner decisions about feeding. For example, animal variables, such as age and neuter status, were associated with the owner feeding choice in the previous study of Knight et al. [15]. In the same study, maintenance of pet health was the most-common reason cited by owners for choosing a particular food [15] whilst, in other research, owner characteristics (such as geographic location) are also associated with food choice [65]. Finally, owner characteristics can also be associated with aspects of dog health; for example, both owner age and income are associated with the prevalence of obesity in dogs [66]. Considering these previous findings, our decision to examine variables beyond those of diet type was justified.
To ensure that owner perceptions of dog health were examined, we created two outcome variables from a question in the original survey question where owners were asked to rate their dog’s health over the previous 12 months [15]. Such responses are not objective measures of health; rather, they require owners to formulate a subjective holistic opinion, which will necessarily require them to take many factors into account (Fig 9a). There will be a contribution from the actual health of their dog (which in turn is associated with other factors), but this will be influenced by their knowledge and recollection, whereas general opinions on what constitutes ‘health’ will be affected by owner attitudes, beliefs and behaviours.
This study also differed from the original study because we excluded dogs <1y because numbers were small (n=26) and, arguably, insufficient to reveal insights into how the perceptions of owners with dogs in their growth phase might vary. We were also concerned about the small numbers (n=111) of respondents who were not make the primary decision about their dog’s diet. Some of these owners reported playing “some lesser role”, whilst others reported playing “no role” in diet decision-making [15]. Not only would this raise concerns about the accuracy of the diet information, but these owners might also be less involved in other aspects of the dog’s life, including health and veterinary matters, which might adversely affect the accuracy of other information provided (Fig 9a). There might also be differences in their opinions and attitudes towards health matters in general. Analyses performed using decision-maker status as an outcome variable supported these concerns, not least since the odds of a dog being recorded as having a health problem were less for owners who were primary decision-makers than for those who were not. One way to account for this would be to include decision-maker status as a covariate in all analyses but the marked imbalance in group size (2,211 vs. 111) could adversely affect the performance of procedures such as logistic regression. Therefore, we instead decided to limit further analyses to owners who were primary decision-makers only. Decisions about excluding observations from a dataset should be made with caution, properly justified and consider the possible impact on results. Based upon our flow diagram of associations within in the dataset (Fig 9b), we determined that it would be possible to eliminate all pathways associated with the decision-maker status variable, thereby simplifying the overall structure without affecting any other associations amongst variables. This was because removing all data from owners who are not primary decision-makers effectively ‘collapses’ the decision-maker status variable into a single category such that variance is no longer associated with it. Estimates of regression coefficients and odds ratios were broadly similar in logistic regression models using the all-owner dataset (S1 Table, S2 Table), compared with models using the primary decision-maker dataset (Tables 2 and 3). Any differences are likely to be the result of differences in the characteristics, attitudes and beliefs of the owners that were not primary decision-makers. Of course, one main disadvantage of focusing on data from owners who are primary decision-makers is that it limits the generalisability of findings to the wider pet-owning public, not least to other members of a dog-owning family. However, given the original concern about data reliability, further studies involving all family members would be required to explore such opinions properly.
The approach of the current study also differed by our decision to include healthcare variables (received medication, veterinary visits, switched to a therapeutic diet) as predictors. In the previous study [15], these variables were instead used as proxy measures for dog health and compared amongst dogs fed different diet types. Whilst both owner perceptions of health and healthcare variables will be associated with actual dog health, they are materially different. As already discussed, owner perceptions of health were derived from a question where owners subjectively rated their dog’s health [15], and these perceptions will be influenced by other factors including owner knowledge and recollections, as well as their beliefs and attitudes (Fig 9). Conversely, the healthcare variables were derived from survey questions that required factual recall rather than the formulation of a subjective, holistic opinion about health. For example, owners were asked whether their dog “had received medications in the last year” and how many times the dog had “visited a veterinarian or received a home visit” [15]. Given that such healthcare information will contribute to the knowledge an owner utilises when constructing opinions about the health of their dog, we reasoned that they were valid predictor variables in our statistical modelling. Further, excluding them from statistical models could be problematic as illustrated in Fig 9c; unlike removing observations from owners who were not primary decision-makers (as discussed above), where any variance associated with the variable is removed, simply excluding a variable from a statistical model will not eliminate any associated variance. The ‘hidden’ variance associated with the healthcare variables might then be incorrectly attributed to other variables leading to erroneous estimates of their regression coefficients. Any unassigned variance also adversely affects model performance, which likely explains why the models that only included owner-animal metadata performed poorly as discussed below.
Using simple logistic regression, we identified associations between owner perceptions of dog health and several owner (e.g., location, diet), animal (e.g., age, breed, neuter status) and healthcare (e.g., veterinary visits, received medication, switched to a therapeutic food) variables. Further, owner perceptions of dog health were associated with the dog diet variable in simple logistic regression (both raw diet and vegan diet associated with the any health problem variable; raw diet associated with the serious illness variable). These findings are not surprising, and would be expected because they were the result of univariable analyses (simple logistic regression) and, therefore, were broadly similar to the univariable analyses (e.g., odds ratios, one-way ANOVA) conducted in the previous study [15]. In the current study, and as described in the methods section, data pre-processing was necessary to ensure that our statistical analyses were valid. The approach taken was to remove or combine groups where numbers of dogs was small, to ensure better balance amongst categories within a predictor variable. The fact that the results obtained from our univariable analyses (simple logistic regression) were similar to those of Knight et al. [15] suggests that, despite this pre-processing, the final dataset remained representative of the dataset from which it was drawn.
We next used a combination of multiple logistic regression and machine learning predictive modelling so that multiple variables could be analysed concurrently. This enabled the creation of models that best predicted owner perceptions of dog health, and to determine the relative importance of variables contributing to the final models. We chose this combined methods approach to maximise the benefits of each whilst minimising their disadvantages. Multiple logistic regression is a well-known, widely available technique that is familiar to most people and has been extensively used in veterinary studies including those in dogs [67–70]. Procedures for supervised model selection are well established, such as those based on BIC [60,61], which can reduce the risk of model overfitting and can take relevant prior knowledge into account. For example, in our exploratory analyses, we tested interactions with a possible clinical (e.g., between the veterinary visit and received medication variables), biological (between the dog sex and neuter status variables) or epidemiological (between the location and setting variables) and psychological (between owner diet and dog diet variables). One disadvantage of our manual selection method is that it could introduce bias into the selection process. Other disadvantages include the fact that logistic regression makes certain statistical assumptions (e.g., independence of errors, the requirement for a linear relationship between the logit of the outcome and predictors, absence of multicollinearity and lack of strong influential outliers; [71]) which, if not met, could lead to poor prediction accuracy. Key advantages of machine learning methods are the fact that they make fewer prior assumptions about data distribution and, therefore, can better model complex, nonlinear relationships between the outcome and predictors even when there is significant noise [72]. However, there are disadvantages including requiring a large dataset, the time and resources needed for computation, challenges with interpretation of the models produced and susceptibility to errors [73]. In this respect, any machine learning procedure is susceptible to overfitting, although this depends upon context; models are often intended to be generalisable to a wide range of datasets beyond the dataset that they were trained on. Arguably, this would be less of an issue for the current study because many of the variables of significance were identified using both multiple logistic regression and machine learning predictive modelling, whilst their relative contribution to model prediction were broadly similar (see below). Therefore, despite the limitations of the different modelling techniques, our findings should be valid.
Using multiple logistic regression, the main variables associated with dog health in the best-fit models were dog age, veterinary visits, and whether the dog had been prescribed medication; in some models, whether the dog had been switched to a therapeutic food was also of significance, especially when owners described their dogs as “generally healthy, but with minor or infrequent problems” (e.g., binary logistic regression on the any health problem variable). Depending on the model used, these variables explained between 34% and 52% of variance within the dataset, with very little additional variance (2% to 6%, depending on the model) gained by the addition of any or all other predictor variables. Many of the same variables were also identified using XGBoost, and prediction accuracy (as measured by AUC) was similar, but other predictor variables were also identified including owner age and education. However, in contrast to the findings of the previous study [15], the diet of the dog had little to no effect in all models, except for those that only included owner-animal metadata, albeit that these models performed poorly. Therefore, we could not find convincing evidence that particular food types (conventional, vegan, vegetarian or raw) have any clinically meaningful positive or negative association with dog health status. As with the original study [15], these analyses were limited by the dataset used, including its cross-sectional design, meaning that causality of associations cannot be assumed, and its reliance on subjective owner perceptions of health status, which might not reflect the actual health of the dogs. Future studies should be considered, for example, cohort studies or randomised trials utilising objective measures of health such as information from electronic patient record as or formal diagnoses of disease by a veterinarian.
A key limitation of using questionnaires to ascertain the perceptions of owners about dog health is that such information might be affected by owner recollection leading to errors in classification of disease status, so-called response bias [74]. Responses might be biased for various reasons including acquiescence, courtesy, the order of questions or even the perceived social desirability of the responses [74]. Recall bias is a form of response bias whereby the recollections of the owner are affected by their past experiences [75]. Owners of dogs with health problems might search their memories more thoroughly for possible risk factors (such as diet choice, number of visits to the vet, use of medication etc, diet changes) than owners of dogs that do not have health problems, with the potential to exaggerate any associations identified.
As discussed above, it is also feasible that the attitudes and beliefs of owners might either consciously or unconsciously have influenced responses about the health of their dog, a point that the authors of the previous study emphasised [15]. For example, owners who believe a particular type of diet to be optimal for dog health, might be more likely to perceive their dog to be healthy, whether or not this was objectively true. Whilst such a bias would feasibly affect any diet type, a greater bias might be expected with diets perceived to be ‘unconventional’, again, as acknowledged in the previous study [15]. Although increasing in popularity [1], both vegan diets and raw meat diets are still uncommon choices for owners, and many veterinary professionals either do not recommend them or might even advise against their use. Current evidence suggests that both such diet types are often not formulated appropriately to be complete and balanced for essential nutrients [5,8,9–11]. Further, there are also concerns with feeding raw food given the potential for contamination either with bacteria of possible pathogenic potential or bacteria resistant to antimicrobials, both of which might pose a risk to the health of owners and other in-contact people [11–13]. Given such health concerns, owners who make an active choice to feed either a raw or vegan diet could be more defensive about this diet choice, compared with owners who feed conventional diets, and this attitude which might unconsciously have biased responses about the health of their dog, for example, minimising the significance of any illnesses. In the original study questionnaire, the authors did attempt to minimise the influence of canine diet variables on reported health status, by placing the health-related questions before most other variables [15]. However, it is unclear as to whether this would adequately eliminate pre-existing unconscious biases resulting from diet choice.
One variable that featured consistently in all models was dog age, whereby illness severity worsened as age increased, especially when in dogs that were ≥6 years. Perceptions of health in senior dogs are likely to be associated with quality of life, which is known to be negatively associated with age [76]. However, it might also be because many chronic diseases have associations with increasing age, including osteoarthritis [68], obesity [77], diabetes mellitus [70] and many forms of cancer [78]]. It is noteworthy that such chronic diseases are also negatively associated with quality of life [79–82]. An alternative explanation for the negative association between age and health in the current study is the fact that owners of older dogs perceive their health to be poorer simply because they expect older dogs to be less healthy [83].
Owner perceptions about dog health were also positively associated with them reporting that their dog had received medication in the last 12 months. Once again, the observational study design means that causality cannot be assumed. It is unclear whether such dogs are receiving these medications because they are genuinely less healthy (which is why they were prescribed), whether the knowledge that their dog is receiving medication means the owner perceives their dog to be less healthy or both (Fig 9). Finally, the association might be indirect, if both the outcome and predictor variable were associated with a third variable. In this regard, there was also a strong positive association between owner perception of health and the frequency of veterinary visits, whilst a positive correlation between frequency of veterinary visits and receiving medication was also evident, meaning that all three variables are likely to be inter-related. Further, in some models, another healthcare variable (switching to a therapeutic diet) potentially associated with frequency of veterinary visits was also included. Besides these, several additional associations were identified with XGBoost, but not with logistic regression, the most important of which were owner age and education. Owners that were older and those with a higher level of education perceived their dogs to be healthier than did other owners. Further studies would be required to explore all such associations further and determine their true significance to canine health.
Although the study was large, the population studied was not representative of the typical dog-owning public. In this respect, 33% and 13% of owners reportedly fed their dog raw and vegan diets, respectively, which is greater than expected; for example, in a recent UK survey, the proportion of owners feeding raw and vegan diets were 7% and <1%, respectively [1]. Further, the proportion of the UK human population reportedly consuming a vegan diet is estimated to be between 2 and 3% [2] whilst, in the current study, 23% of owners reported consuming a vegan diet. This raises a concern about how representative this study population as a whole, and is of particular significance given the strong clustering based on owner diet. Although the reasons for such an imbalance is not clear, it might well have resulted from the method of owner recruitment. In this respect, the study was widely advertised via social media but, given concerns that owners feeding dogs unconventional diets might be under-represented, relevant interest groups were actively targeted and invited to participate. This strategy was successful in ensuring that adequate numbers of dog owners feeding unconventional diets were included, ensuring that statistical comparisons were meaningful in both the previous [15] and current studies. However, the recruitment strategy might have inadvertently generated an unrepresentative study population which, in turn, may have influenced the results obtained, as discussed above.
In addition to the unrepresentative study population, multiple associations amongst groups were identified with the potential for confounding amongst variables. For example, feeding a raw diet was more commonly reported by owners from the UK, whilst feeding a vegan diet was more common in owners from Europe. Further, there were positive associations between owner income and education, and between dog age and neuter status. Including healthcare variables in multiple logistic regression analyses might also have hidden the associations seen with the owner-animal metadata, including diet type. In logistic regression analyses, we accounted for these associations by determining multicollinearity using VIF, refining models where this might be a problem. No significant issues of multi-collinearity amongst predictor variables were identified in any of the final models. Whilst this does not necessarily guarantee that variance from each factor in the models was correctly assigned, any errors are likely to be small. The use of XGBoost would further address these concerns given that these techniques do not make prior assumptions about data distribution and can better take account of complexity in interactions amongst variables [69]. Nonetheless, to ensure any owner-animal metadata effects were not unfairly penalised by the inclusion of the healthcare variables, we also examined models containing only the owner-animal metadata. In these models, dog age remained as the variable of most importance, but other variables of lesser importance were identified that had not been seen on the original models (including breed size category, owner age, education level and urban setting), whilst a limited effect of diet remained (i.e., the vegan diet variable was only the 12th most important predictor variable). The significance of any associations between these variables owner perceptions of health are not known. Nevertheless, since predictive accuracy of all models (both multiple logistic regression and XGBoost) utilising just the owner-animal metadata was poor, these analyses should be considered exploratory and any conclusions drawn should be made with caution.
A final approach we took, to ensure that the we fully accounted for any possible effect of the dog diet vegan variable, was to add it back into the best-fit models for both the any health problem and serious illness outcome variables. In both cases, this additional predictor variable did not improve the performance of either model and, interestingly, models with equivalent performance could be created when the owner diet vegan was instead added.
The main limitations to the current study have already been discussed above, including the fact that owner-reported health was the main outcome measure and the fact that the study population was not representative of the general dog-owning population. Further, significant data pre-processing was required to ensure adequate group sizes for statistical analyses. The disadvantage of such an approach was that many smaller groups needed to be excluded, meaning that we might have missed some variables with a potential impact on owner-reported canine health. That said, in preliminary testing, we coded several of our variables in different ways (e.g., dog age as a continuous or ordinal 5-category variable; veterinary visits as a 4- or 5-category ordinal variable), and selected the approach that produced the best model fit. Therefore, whilst some genuine effects from small categories might still have been missed, the final best-fit models we selected were those that most closely fitted the data.
A third limitation was the fact that we did not contact the authors of the original study before conducting our data analyses, to clarify any uncertainties with the dataset and seek guidance on our planned approach. We chose not to do so because the original paper was well written, with clearly presented study methodology and results. Further, since independent replication is critical to the scientific method, it was arguably better not to have direct contact with the primary study authors when conducting statistical analyses.
A fourth limitation was that the outcome variables tested in the current study were derived from only one of the health metrics examined in the original study, namely owner-reported dog health [15]. Other health metrics reported by Knight et al. included asking owners “to report what they believed their veterinarian’s assessment [of their dog’s health] to be” [15] and owner reports of 22 possible health disorders, grouped into broad categories (e.g., allergy, cancer/tumours, hormonal, skin/coat) [15]. Our conclusions might have differed if one or more of these alternative health metrics had been analysed, but we decided against this because, in our opinion, all were inferior measures. Not only were these metrics also based on subjective owner reports, but owners were required to speculate on what their veterinarian’s opinion about the health of their dog would be. Further, there were no clear definitions of the health disorder categories, meaning that these would likely be broad, comprising many different diseases with differing aetiology and severity. Finally, the number of dogs included in many categories was small, being ≤50 in 12 of the 22 [15]. Whilst further work could be considered, where these additional metrics are analysed using the techniques of the current study, it is questionable as to whether anything further would be gained. Instead, it would be preferable to conduct a prospective study utilising objective, veterinary-derived measures of health, as discussed above.
A final limitation was the fact that there might be confounding between healthcare variables (e.g., veterinary visits and received medication) and dietary choices. For example, owners who are sceptical about conventional diets and instead choose to feed either vegan or raw food, might also be less willing either visiting their veterinarian or allowing their dog to receive medication. In support of this, negative associations were identified in the current study between feeding a raw diet and both veterinary visit frequency and the use of medications; further, in the previous study, dogs fed raw diets were less likely to be neutered than dogs fed conventional diets [15]. However, results were inconsistent for dogs fed vegan diets; whilst there was a negative association between feeding a vegan diet and the odds of a dog receiving medications, there was no association with frequency of veterinary visits. Further, in the previous study, dogs fed vegan diets were more likely to be neutered than those fed conventional diets [15]. Given these complexities, further research would be required to investigate associations between different diet choices and healthcare variables.
Conclusions
Variables most strongly associated with owner perceptions of health are dog age and healthcare variables including number of veterinary visits, receiving medication and being switched to a therapeutic diet. These results extend the findings of previously-published research [15] and suggest that, when other variables are accounted for, the association between dog diet choice and owner-perceived canine health is negligible.
Acknowledgements
The authors wish to acknowledge the authors of the original paper [15] for providing making their data freely available to other researchers for further analysis.
Funding statement
This study was not funded by any research grant. The costs of open access publishing were covered by the University of Liverpool Institutional Open Access Fund (https://www.liverpool.ac.uk/open-research/open-access/open-access-funding/).
Competing interests statement
AJG is an employee of the University of Liverpool, but his position is financially-supported by Royal Canin. AJG has also received financial remuneration and gifts for providing educational material, speaking at conferences, and consultancy work, all unrelated to the current study. Royal Canin had no involvement in any aspect of the current work including study design, data analysis, drafting the manuscript or the decision to submit the work for publication.
RBJ is an employee of the University of Liverpool, whose salary is financially-supported by the Higher Education Funding Council for England. RBJ also receives a stipend from United Kingdom Research and Innovation for work as a panellist for the Biotechnology and Biological Sciences Research Council.
None of these disclosures alter our adherence to PLOS ONE policies on sharing data and materials.
Research data for this article
The original research data are accessible at https://rebrand.ly/2020/pet-food-consumers-study. All statistical reports and codes are available in the supplementary information (S1-S5 Files) and online (https://github.com/RichardBJ/CanineHealth22).
Supplementary material
S1 File. Statistical report containing details of all data pre-processing steps to create the dataset for primary decision-makers, as well as initial data visualisation using UMAP.
S2 File. Statistical report containing details of all data pre-processing steps to create the dataset for all owners.
S3 File. Statistical report containing details of all statistical analyses for the Kendall’s rank correlation, and also XGBoost modelling for the any health problem outcome using data from primary decision makers.
S4 File. Statistical report containing details of all statistical analyses for the Kendall’s rank correlation, and also XGBoost modelling for the any health problem outcome using using data from all owners.
S5 File. Statistical report containing details of XGBoost modelling for the significant illness outcome using data from primary decision makers.S6 File. Statistical report containing details of XGBoost modelling for the significant illness outcome using data from all owners.
S7 File. Statistical report containing details of binary logistic regression modelling for the any health problem variable using data from primary decision makers.
S8 File. Statistical report containing details of binary logistic regression modelling for the any health problem outcome using data from all owners.
S9 File. Statistical report containing details of binary logistic regression modelling for the significant illness variable using data from primary decision makers.
S10 File. Statistical report containing details of binary logistic regression modelling for the significant illness variable using data from all owners.
S1 Table. Results of simple (i.e., univariable) binary logistic regression analyses examining associations between owner, animal and veterinary variables and the any health problem binary for all owners.
S2 Table. Results of simple (i.e., univariable) binary logistic regression analyses examining associations between owner, animal and veterinary variables and the significant illness binary for all owners.
S3 Table. Kendall’s tau correlation coefficients for the naïve associations amongst owner-animal metadata variables in Fig 2.
S4 Table. Respective log P-values for the Kendall’s tau correlations amongst owner-animal metadata variables in Fig 2 and S1 Table.
S5 Table. Final multiple binary logistic regression model examining associations between owner, animal and veterinary variables and the presence of any health problem as reported by the owner.
S6 Table. Final multiple binary logistic regression model examining associations between owner, animal and veterinary variables and the presence of a significant or serious illness as reported by the owner.
S1 Fig. Owner-dog metadata visualised by Uniform Manifold Approximation and Projection with Density-Based Spatial Clustering. Owner-pet metadata were pre-processed with all factors as numeric and subject to dimension reduction with the UMAP projection technique. Owner variables included in this visualisation were diet, sex, location, education and income; animal variables included were age, sex, neuter status, diet and breed size category. Healthcare variables were not included. Each individual row of the data (owner-dog combination) contributes one UMAP x and y coordinate, analogous to PC1 and PC2 in a principal component analysis. Four or five distinct clusters were evident. Points are colour-coded by owner gender (a), income (b), location (c), setting (d), education € and decision-maker status (f) as indicated in the legend.
S2 Fig. Owner-dog metadata visualised by Uniform Manifold Approximation and Projection with Density-Based Spatial Clustering. Owner-pet metadata were pre-processed with all factors as numeric and subject to dimension reduction with the UMAP projection technique. Owner variables included in this visualisation were diet, sex, location, education and income; animal variables included were age, sex, neuter status, diet and breed size category. Healthcare variables were not included. Each individual row of the data (owner-dog combination) contributes one UMAP x and y coordinate. Four or five distinct clusters were evident. Points are colour-coded by owner age (a), animal career (b), dog age (c), breed size category (d), dog sex € and neuter status (f) as indicated in the legend.
S3 Fig. Multiple binary logistic regression model, on data from all owners, with any health problem as the outcome variable and including either all owner, animal and including either all healthcare variables (a) or only the owner-animal metadata (b). The dots represent the odds ratio for each variable, whilst the bars represent 99% confidence intervals (99%-CI). Variables where the 99%-CI range does not include 1.0 (vertical dotted line) are depicted in red, whilst those that include 1.0 are depicted in blue. Note the logarithmic scale for the X-axis.
S4 Fig. Multiple binary logistic regression model with significant illness as the outcome variable, on data from owners who were primary carers, either all owner, animal and including either all healthcare variables (a) or only the owner-animal metadata (b). The dots represent the odds ratio for each variable, whilst the bars represent 99% confidence intervals (99%-CI). Variables where the 99%-CI range does not include 1.0 (vertical dotted line) are depicted in red, whilst those that include 1.0 are depicted in blue. Note the logarithmic scale for the X-axis.
S5 Fig. All-variable XGBoost model on the any health problem binary using the all-owner dataset. (a) Receiver operating characteristic curve of a prediction model containing all variables (owner, animal and healthcare). This shows the increasing true positive and false positive rates, with decrease of the threshold probability for prediction of any health problem. Prediction accuracy was good as assessed by ROC analysis (area under curve 0.838, 99%-CI: 0.797-0.879). Acceptance threshold is indicated by the colour bar on the right-hand side and shown at discrete points on the curve. (b) Graph depicting the relative contribution of predictor variables to the all-variable XGBoost model for the any health problem binary. ‘Importance’ represents fractional contribution of each feature to the model, based on the total gain from including each feature; ‘cover’ represents the number of observations related to this feature in the model); ‘frequency’ represents the relative number of times a feature has been used in trees). Variables are organised in order of importance left to right, based on the sum of the 3 metrics, whilst variables that are not shown did not contribute to the final model.
S6 Fig. Reduced (owner-animal metadata) XGBoost model on the any health problem binary using the all-owner dataset. (a) Receiver operating characteristic curve of a reduced prediction model only containing owner and animal variables. This shows the increasing true positive and false positive rates, with decrease of the threshold probability for prediction of any health problem. Prediction accuracy was moderate, as assessed by ROC analysis (AUC 0.682, 99%-CI: 0.628-0.737). Acceptance threshold is indicated by the colour bar on the right-hand side and shown at discrete points on the curve. The fact that the prediction thresholds are all low (< 0.5) shows that this model struggles to predict health issues. (b) Graph depicting the relative contribution of predictor variables to the all-variable XGBoost model for the any health problem binary. ‘Importance’ represents fractional contribution of each feature to the model, based on the total gain from including each feature; ‘cover’ represents the number of observations related to this feature in the model); ‘frequency’ represents the relative number of times a feature has been used in trees). Variables are organised in order of importance left to right, based on the sum of the 3 metrics, whilst variables that are not shown did not contribute to the final model.
S7 Fig. All-variable XGBoost model on the significant illness binary using the all-owner dataset. (a) Receiver operating characteristic curve of a prediction model containing all variables (owner, animal and healthcare). This shows the increasing true positive and false positive rates, with decrease of the threshold probability for prediction of significant illness. Prediction accuracy was good, as assessed by ROC analysis (area under curve 0.884, 99%-CI: 0.796-0.972). Acceptance threshold is indicated by the colour bar on the right-hand side and shown at discrete points on the curve. (b) Graph depicting the relative contribution of predictor variables to the all-variable XGBoost model for the significant illness binary. ‘Importance’ represents fractional contribution of each feature to the model, based on the total gain from including each feature; ‘cover’ represents the number of observations related to this feature in the model); ‘frequency’ represents the relative number of times a feature has been used in trees). Variables are organised in order of importance left to right, based on the sum of the 3 metrics, whilst variables that are not shown did not contribute to the final model.
S8 Fig. Reduced (owner-animal metadata) XGBoost model on the significant illness binary using the all-owner dataset. (a) Receiver operating characteristic curve of a reduced prediction model only containing owner and animal variables. This shows the increasing true positive and false positive rates, with decrease of the threshold probability for prediction of significant illness. Prediction accuracy was moderate, as assessed by ROC analysis (area under curve 0.664, 99%-CI: 0.535-0.693). Acceptance threshold is indicated by the colour bar on the right-hand side, and also shown at discrete points on the curve. (b) Graph depicting the relative contribution of predictor variables to the all-variable XGBoost model for the significant illness binary. ‘Importance’ represents fractional contribution of each feature to the model, based on the total gain from including each feature; ‘cover’ represents the number of observations related to this feature in the model); ‘frequency’ represents the relative number of times a feature has been used in trees). Variables are organised in order of importance left to right, based on the sum of the 3 metrics, whilst variables that are not shown did not contribute to the final model.