Deep learning classification of canine behavior using a single collar-mounted accelerometer: Real-world validation
1Pet Insight Project, Kinship, San Francisco, California, United States of America
2WALTHAM Petcare Science Institute, Melton Mowbray, Leicestershire, United Kingdom
*scottlyle@whistle.comAbstract
Collar-mounted canine activity monitors can use accelerometer data to estimate dog activity levels, step counts, and distance traveled. With recent advances in machine learning and embedded computing, much more nuanced and accurate behavior classification has become possible, giving these affordable consumer devices the potential to improve the efficiency and effectiveness of pet healthcare. Here we describe a novel deep learning algorithm that classifies dog behavior at sub-second resolution using commercial pet activity monitors. We built machine learning training databases from over 5,000 videos of over 2,500 dogs and ran the algorithms in production on over 11 million days of device data. We then surveyed project participants representing 10,550 dogs, they provided us 163,110 event responses to validate real-world detection of eating and drinking behavior. The resultant algorithm displayed a sensitivity and specificity for detecting drinking behavior (0.949 and 0.999, respectively) and eating behavior (0.988, 0.983). We also demonstrated detection of licking, petting, rubbing, scratching, and sniffing. We show that the devices’ position on the collar has no measurable impact on performance. In production, users reported a true positive rate of 95.3% for eating (among 1,514 users), and of 94.9% (among 1,491 users) for drinking. The study demonstrates the accurate detection of important health-related canine behaviors using a collar-mounted accelerometer. We trained and validated our algorithms on a large and realistic training dataset, and we assessed and confirmed accuracy in production via user validation.
Introduction
Much as recent progress in smartwatches has enabled new telehealth applications [1–4], recent progress in internet-connected pet wearables, such as collar-mounted activity monitors, has prompted interest in using these devices to improve the cost and efficacy of veterinary care [5]. Just as with smartwatches in human telehealth, accelerometerbased activity monitors have emerged as an inexpensive, low-power, and informationrich approach to pet health monitoring [6–8].
Accelerometer-based pet activity monitors analyze the moment-to-moment movement measured by a battery-powered accelerometer. They are typically attached to the pet via a collar, though attachment methods may be more elaborate in research settings. Using the device’s accelerometer signal (sometimes in combination with gyroscope, magnetometer, GPS, or other sensor signals), collar-mounted activity monitors can accurately estimate pet activity levels [9–14], step count, and distance traveled [12].
In recent years, advances in machine learning have allowed pet activity monitors to move beyond estimating aggregate activity amounts, to detecting when and for how long a pet performs common activities such as walking, running, lying down, or resting [15–17], These biometric capabilities have progressed to include increasingly specific and varied activities such as drinking, eating, scratching, and head-shaking [8,16,18–21].
The benefits of accurate and quantitative behavior detection in pet health are extensive. Pet activity monitors have been shown to be useful in the detection and diagnosis of pruritis [22,23] and in potential early prediction of obesity [24]. They have also been used in monitoring response to treatments such as chemotherapy [25]. Furthermore, statistical analysis of activity and behavior monitoring on large numbers of pets can be an expedient approach to medical and demographic studies [24], since sample sizes are potentially very large.
Although several studies have demonstrated and measured the accuracy of activity recognition algorithms [8,16,18–21], the datasets used to train and evaluate the algorithms are typically not representative of the broad range of challenging environments in which commercial pet activity monitors must function. For instance, most existing studies use exclusively healthy dogs and are often run in controlled environments that promote well-defined and easily detectable behaviors with a low risk of confounding activities.
Unfortunately, real-world algorithm performance often lags far behind the performance measured in controlled environments [26,27]. For instance, existing studies typically ensure careful installation of the device in a specific position on a properly adjusted collar. In real-world usage, collars vary in tightness and often rotate to arbitrary positions unless the activity monitor device is very heavy. Collar rotation and tightness [28], as well as the use of collar-attached leashes [29], can compromise performance. In our experience, confounding activities like riding in a car or playing with other pets can produce anomalous results if not adequately represented in training datasets. Finally, some studies use multiple accelerometers or harness-mounted devices [30], which limit applicability in many consumer settings.
The work described here was performed as part of the Pet Insight (PI) Project [31], a large pet health study to enable commercial pet activity monitors to better measure and predict changes in a pet’s health by:
- Sourcing training data from project participants and external collaborators to build machine learning training databases and behavior detection models such as those described in this work.
- Combining activity data, electronic medical records, and feedback from more than 69,000 devices distributed to participants over 2-3 years to develop and validate proactive health tools.
- Using the resulting datasets, currently covering over 11 million days in dogs’ lives, to enable insights that support pet wellness and improve veterinary care.
This work presents the results of the PI Project’s efforts to develop and validate these behavior classification models [32]. It includes evaluation of model performance in a real-world context and addresses limitations from controlled research settings such as device fit and orientation.
Materials and methods
Activity Monitor
Data were collected primarily via a lightweight canine activity monitor (Whistle FIT®, Mars Petcare, McLean, VA, Fig 1), which was designed and produced specifically for this study). Smaller amounts of data were collected via the commercially available Whistle 3® and Whistle GO® canine activity monitors. All three devices used the same accelerometer. Unlike the Whistle FIT®, these latter devices are furnished with GPS receivers and cellular radios. However, in all cases, the behavior classification in this study is performed using only the output of the devices’ 3-axis accelerometers.
Accelerometry data collection
All monitoring devices acquired accelerometry data and uploaded it according to their usual operation. That is, the devices acquired 25-50 Hz 3-axis accelerometry data for at least several seconds whenever significant movement was detected. Data were compressed and annotated with timing data using a proprietary algorithm. Data were temporarily stored on-device and then uploaded at regular intervals when the devices were in Wi-Fi range. Uploads were processed, cataloged, and stored in cloud-hosted database services by Whistle servers. The compressed accelerometry data were retrieved on demand from the cloud database services in order to create the training, validation, and testing databases used in this study.
Animal behavior data collection
Animal behavior data was collected (summarized in Table 1 and described further elsewhere in this report) and used to create two datasets used in model training and evaluation:
- Crowd-sourced (crowd) dataset. This dataset contained both (a) long (multi-hour) in-clinic recordings, as well as (b) shorter recordings submitted by project participants. This large and diverse dataset was meant to reflect real-world usage as accurately as possible.
- Eating and drinking (eat/drink) dataset. This dataset consisted of research grade sensor and data using a protocol designed to represent EAT and DRINK behaviors. Other observed behaviors were incidental.
For brevity, we refer to these datasets simply as the crowd and eat/drink datasets.
Eat/drink study protocol
This study was conducted using dogs owned by the WALTHAM Petcare Science Institute and housed in accordance with conditions stipulated under the UK Animals (Scientific Procedures) Act 1986. Briefly, the dogs were pair housed in environmentally enriched kennels designed to provide dogs free access to a temperature-controlled interior and an external pen at ambient temperature. Dogs were provided with sleeping platforms at night. The dogs had access to environmentally enriched paddocks for group socialization and received lead walks and off-lead exercise opportunities during the day. Water was freely available at all times and dogs were fed to maintain an ideal body condition score. The study was approved by the WALTHAM Animal Welfare and Ethical Review Body. One hundred and thirty-eight dogs across 5 different breeds (72 Labrador Retrievers, 18 Beagles, 17 Petit Basset Griffon Vendeens, 14 Norfolk Terriers and 17 Yorkshire Terriers) took part for two consecutive days each. Each dog was recorded once a day during its normal eating and drinking routine using a GoPro camera (GoPro, San Mateo, CA).
In this study, either one (ventral only) or four (ventral, dorsal, left, and right) activity monitors were affixed to a collar. For each observation, the collar was removed from the dog, the correct number of activity monitors were attached, and then shaken sharply in view of the camera to provide a synchronization point that was identifiable in both the video and accelerometer signals (so that any time offset could be removed). The collar was then placed on the dog at a standardized tightness. The dogs were recorded from approximately one minute before feeding until approximately one minute after feeding. In order to increase the diversity of the dataset, collar tightness was varied between a two-finger gap and a four-finger gap, and food bowls were rotated between normal bowls and slow-feeder or puzzle-feeder bowls. For each data recording, researchers noted the date and time, device serial number(s), collar tightness, food amount and type, and various dog demographic data.
Crowd-sourcing protocol
Pet Insight participants were requested to use smartphones to video record their pets performing everyday activities while wearing activity monitors. The participants were told that the activity monitor should be worn on the collar but were not given any other instructions about how the collar or monitor should be worn. Participants were asked to prioritize recording health-related behaviors like scratching or vomiting, but to never induce these events and to never delay treatment in order to record the behaviors. As a participation incentive, for every crowd-sourced video used, the PI project donated one dollar to a pet-related charity.
After recording each video, participants logged into the PI crowd-sourcing website, provided informed consent, uploaded the recorded video, and completed a short questionnaire confirming which pet was recorded and whether certain behaviors were observed. The device automatically uploaded its accelerometry data to Whistle servers.
In-clinic observational protocol
This study was conducted at several Banfield Pet Hospital (BPH) clinics. Its objective was to acquire long-duration (multi-hour) naturalistic recordings to augment the shorter crowd-sourced recordings, which were typically several minutes or less in duration.
Randomly selected BPH clients who chose to participate signed fully informed consent forms. Their dogs were outfitted with Velcro breakaway collars with one attached activity monitor device each. Collar tightness and orientation were not carefully controlled. Video was recorded via a 4-channel closed-circuit 720p digital video security system. Video cameras were ceiling- or wall-mounted and oriented towards the in-clinic kennels so that up to four dogs could be observed at a time. For each recording, researchers noted the date and time, the device serial number, and the dog/patient ID number.
Video labeling
All uploaded videos were transcoded into a common format (H.264-encoded, 720p resolution, and up to 1.6 Mb/s) using Amazon’s managed Elastic Transcoder service, and their audio was stripped for privacy. Video start times were extracted from the video metadata and video filenames. Matching device accelerometry data was downloaded from Whistle’s databases, and automatic quality checks were performed.
Videos were then labeled by trained contractors using the open-source BORIS (Behavioral Observation Research Initiative Software) software application [33]. The resulting event labels were imported and quality-checked using custom Python scripts running on one of the PI project’s cloud-based web servers. Labels were stored alongside video and participant metadata in a PostgreSQL database.
All video labeling contractors were trained using a standardized training protocol, and inter-rater reliability analyses were performed during training to ensure consistent labeling. Videos were labeled according to a project ethogram [8,15,20]. This report describes several of these label categories.
Labelers divided each video into valid and invalid regions. Regions were only valid if the dog was clearly wearing an activity monitor and was fully and clearly visible in the video. Invalid regions were subsequently ignored. In each valid video region, the labeler recorded exactly one posture, and any number (0 or more) of applicable behaviors.
Postures (Table 2) reflect the approximate position and energy expenditure level of the pet, while behaviors (Table 3) characterize the pet’s dominant behavior or activity in a given moment. For instance, during a meal, a dog might exhibit a STAND posture and an EAT behavior. While pausing afterwards, the same dog might exhibit a STAND posture and no behavior. Multiple simultaneous behaviors are rare but possible, such as simultaneous SCRATCH and SHAKE behaviors.
Training data preparation
Although accelerometer data and smartphone video data were both time-stamped using the devices’ network-connected clocks, inaccuracies led to alignment errors of typically several seconds, and sometimes much longer. Short activities such as SHAKE, in particular, require more accurate alignment. We aligned approximately 1,200 videos manually by matching peaks in accelerometer activity to labels for high-intensity behaviors like SHAKE and SCRATCH. We used these manual alignments to develop and validate an automatic alignment algorithm that aligned the remaining videos.
We created each of the two training datasets (crowd and eat/drink) by:
- Selecting appropriate videos from our database.
- Limiting the number of entries per dog to 30 (some dogs are overrepresented in our database).
- Allocating all of each dog’s data into one of 5 disjoint cross-validation folds.
- Downloading each dataset and labeling each time-point with a posture and/or behavior(s).
The specific method of separating data into cross-validation folds (step 3 above) is critical [34]. Classifiers trained on individual dogs have been shown to overperform on those dogs relative to others, even if those classifiers are trained and evaluated using separate experimental observations. Gerencsér et. al. experienced an accuracy reduction from 91% for a single-subject classifier to 70-74% when generalizing to other dogs [35]. Consequently, we were careful to ensure that all of a dog’s videos fall in a single fold, so that data from a single dog is never used to both train and evaluate a classifier.
The overall data acquisition process, from video capture to a completed dataset, is shown in Fig 2.
Deep learning classifier
Our deep learning classifier is based on our FilterNet architecture, which we have published in detail in a previous work [32]. We implemented the model in Python using PyTorch v1.0.1 [36] and the 2020.02 release of the Anaconda Python distribution (64-bit, Python 3.7.5). We trained and evaluated our models on p2.xlarge instances on Amazon Web Services [37] with 4 vCPUs (Intel Xeon E5-2686 v4), 61 GB RAM, and a NVIDIA Tesla k80 GPU with 12 Gb RAM, running Ubuntu 18.04.4.
We used the crowd dataset for cross-validated training and evaluation (Fig 3). Specifically, we trained and evaluated five different models, using a different held-out fold as a test set for each model. We combine the models’ predictions for each of the five test sets for model evaluation, as described below. We also generated behavior classifications for the eat/drink dataset using one of the models trained on the crowd dataset (that is, we did not use the eat/drink dataset for model training). There were no dogs in common between the crowd and eat/drink datasets, so cross-validation was not needed in this step.
Evaluation
For evaluation, we modeled the task as two multi-class classification problems, one for behaviors and one for postures. At each timepoint in each video entry in a dataset we recorded the labeled behavior and posture, and every 320 ms we calculate the most likely predicted behavior and posture. We tallied the labeled and predicted pairs from all five test folds together using the PyCM multiclass confusion matrix library to create separate behavior and posture confusion matrices [38]. We used the PyCM package to calculate metrics derived from the confusion matrices [39].
Because the MIXED posture is used primarily for expediency in labeling, we dropped any timepoints with MIXED labels from the postures confusion matrix, and replaced any MIXED-class posture predictions with the next most likely prediction for that timepoint. We also excluded any timepoints with more than one simultaneous labeled behavior (about 3% of the data) from the behaviors confusion matrix.
Furthermore, following Uijl et. al. [8], we excluded any timepoints within 1 second of a class transition in both classification problems. However, also similar to [8], we treated the SHAKE class differently due to its very short duration. For SHAKE, we only excluded the outer one-third second. In dropping these transition regions, we attemptted to follow established convention for minimizing the effects of misalignment in labeling, and to make our reported results easier to compare to related works.
User validation
Although the crowd dataset is meant to be representative of real-world data, it is subject to biases such as underrepresentation of behaviors that are unlikely to be video recorded, such as riding in cars or staying at home alone. Furthermore, it is impossible to anticipate all of the myriad situations that may serve as confounders.
Consequently, we ran real-world user validation campaigns on the two behaviors that users are most likely to be aware of, EAT behavior and DRINK behavior. We defined events as periods of relatively sustained, specific behaviors detected with high confidence, such as eating events (meals) consisting of several minutes of sustained eating behavior. We adapted our production system, which runs the models described in this work in nearreal-time on all PI project participants, to occasionally send validation emails to participants when an EAT or DRINK event had occurred within the past 15 minutes. Respondents categorized the event detection as correct (“Yes”) or incorrect (“No”) or indicated that they weren’t sure. Users were able to suggest what confounding event may have triggered any false predictions. We excluded any responses that arrived more than 60 minutes after an event’s end, as well as any “Not Sure” responses.
Results
Data collected
After applying the steps described above the crowd dataset contained data from 5,063 videos representing 2,217 subjects, and the eat/drink dataset contained data from 262 videos representing 149 unique dogs.The distribution of weights and ages represented in these datasets is shown in Fig 4, while a breed breakdown is given in Table 4. As expected, the crowd dataset exhibits a far greater diversity of weights, ages, and breeds than the eat/drink dataset, since the eat/drink subjects are sampled from several relatively homogeneous subpopulations.
These datasets also differed in the length and frequency of labeled events, as shown in Table 5. The crowd and eat/drink datasets contain 163.9 and 22.4 hours of video data labeled as VALID, respectively.
The EAT class was highly represented in both the crowd dataset (because participants were specifically requested to submit videos of their dogs at mealtime, since it is an easily filmed and important behavior) and in the eat/drink dataset (due to study design). The eat/drink dataset included only small amounts of incidental LICKSELF, SCRATCH, PETTING, and SHAKE behavior, while the crowd dataset contained many of these events because participants were repeatedly reminded of their importance.
The class distribution of both datasets is highly imbalanced, which presented a challenge for algorithm training. For instance, in the crowd dataset, which we used for training, the EAT class total duration is 117 times greater than that of SHAKE.
The distribution of lengths for each label class was highly skewed, with many short labels and a smaller number of longer labels (Fig 5). Some of this skew was due to label fragmentation, where a long stretch of the labeled activity is interrupted either by the dog temporarily pausing (for instance, lifting up its head to look around several times while drinking or while eating a meal) or by discontinuities in the labeling when the dog leaves the camera’s field of view (since labelers only marked videos as VALID when the dog was fully and clearly visible). The distribution of SHAKE labels was less skewed, likely because it is typically a short behavior and less prone to interruption.
Classification accuracy
Cross-validated classification metrics for the crowd dataset are given in Table 6, and classification metrics obtained from evaluating the eat/drink dataset using a model trained on the crowd dataset are given in Table 7. The tables give metrics both for behavior and posture classes. However, some subsequent sections report only behaviors, because postures are much less carefully labeled and are typically used in an aggregate form where individual misclassifications are less important.
It is important to note that the class balance (class prevalence) of these datasets is not representative of real-world canine behavior. Because the videos are typically taken in stimulating or interesting situations, these datasets exhibit a lower relative prevalence of LIE DOWN and other low-energy postures. Furthermore, the datasets exhibit much higher levels of EAT, DRINK, and possibly other behaviors, due to either study design (in the eat/drink dataset) or because the PI project requested that participants film certain behaviors.
Of the metrics in Table 6 and Table 7, only sensitivity and specificity are independent of class prevalence.
The “behaviors” confusion matrix for the crowd dataset is shown in Fig 6 in nonnormalized and normalized forms. The non-normalized confusion matrix gives raw tallies (that is, the total number of one-third second time points) of predicted and labeled classes, and the normalized confusion matrix gives the percentage of each actual label classified by the algorithms as a given predicted label (so that the percentages in each row sum to 100%). The non-normalized matrix is dominated by correctly predicted NONE and EAT samples, due to their high prevalence and effective classification in this dataset. The normalized matrix suggests the reliable classification of DRINK, EAT, NONE, and SHAKE. The LICKSELF and SCRATCH classes are of moderate reliability, and the LICKOBJECT, PETTING, RUBBING, and SNIFF classes exhibit some systematic misclassification and are of lesser reliability.
Effect of device position on performance
The system’s classification performance, as measured by F1 score, shows no significant dependence on device position (Fig 7). This invariance is a key property that enables real-world performance to approach that of controlled studies.
User validation
Participants responded far better than expected to user validation efforts. Users opened emails, clicked through to the web form, and submitted validation results for 55% of the EAT validation emails and 42% of the DRINK validation emails.
Responses are summarized in Table 8. As described above, we excluded any responses that arrived more than 60 minutes after an event’s end, as well as any “Not Sure” responses. The positive (“Yes”) validation rate was approximately 95% for both event types. As expected, the rate of users responding “Not Sure” was far greater for DRINK (12%) than for EAT (2%). Participant comments confirmed our expectation that users were less aware of DRINK behavior than of EAT behavior. This lack of awareness likely also contributed to the lower DRINK response rate.
As the production system generates candidate EAT and DRINK events, it calculates a confidence score (the mean algorithm confidence over the event’s duration) that varies between 0 and 1.0, and drops any events with a score below a threshold of 0.3. Fig 8 shows how the percentage of “Yes” responses (the true positive rate) varied with this confidence score. For EAT events, the rate grew from 83% for the lowest-confidence bin (0.3-0.4) to 100% (201 out of 201) for the highest-confidence bin (0.9-1.0). Since users do not see the confidence score, this trend suggests that the EAT validation data are relatively reliable. The DRINK data show a less convincing trend, which is consistent with users’ lower awareness of DRINK events.
It is unfortunate that, of the behavior classes measured in this work, only EAT is likely to exhibit the level of user awareness required for validation using this method.
Discussion
Comparison with Previous Work
We compare our dataset and results with several previous works in Table 9, and we tabulate several important qualitative differences between the datasets in Table 10.
In comparing these results, it is important to account for:
- Class distribution. Each dataset exhibits a different distribution of behaviors. In general, classifiers exhibit better F1 scores for common behaviors than for rare behaviors. The classifier sensitivity and specificity are relatively insensitive to this distribution, so we recommend using these metrics for comparing performance across different datasets.
- Dataset collection methods. Classifiers are more accurate when applied to high-quality datasets collected under controlled conditions. Accuracy can drop substantially in naturalistic versus laboratory settings [26,27]. Classifiers benefit from consistent device position, device attachment, and collar tightness, and they also benefit when the labeled behaviors as well as the collection environment are consistent and well-defined.
Previous works have used relatively controlled and high-quality datasets, similar to the eat/drink dataset in this work [8,18,19,21]. The classification performance of the classifier presented here on the EAT and DRINK classes in the eat/drink dataset advances the sensitivity, specificity, and F1 score for these classes. Sensitivity and specificity are independent of class prevalence. The balance between sensitivity and specific is a design choice, so we have calibrated our algorithms to favor specificity in order to minimize false positives.
The classifiers’ performance on SCRATCH in the challenging crowd dataset also advances the state of the art. Comparable detection of LICKOBJECT, LICKSELF, PETTING, RUBBING, and SNIFF has not been previously demonstrated to our knowledge. We note that SCRATCH, LICKSELF, and RUBBING behaviors are highly relevant to dermatological health applications [19], and that PETTING is an important confounder that can be easily misclassified as SCRATCH or LICKSELF in classifiers that are not exposed to this behavior. We have found the classifiers’ detection of SHAKE to be highly accurate (though susceptible to temporal misalignment between device and video data, due to the short event lengths). It is difficult to compare the per-time-sample SHAKE classification metrics here to published per-event metrics due to differing methodologies [8,18].
Challenges
The system excels at certain clearly defined and easily recognizable activities, especially those repeating and universal movement patterns such as drinking (lapping), walking, running, shaking, and most eating behaviors. It also performs well on well-defined instances of scratching and self-licking.
Device positioning and collar tightness do not appear to have a strong effect on system accuracy, meaning that accurate behavior metrics can be acquired via normal activity monitor usage. An important feature of the devices described in this study is their insensitivity (invariance) to collar orientation or position (Fig 7). In real-world settings, and especially with lightweight devices such as the Whistle FIT, the device can be, and often is, rotated away from the conventional ventral (bottom) position at the lowest point of the collar.
The system appears to use the angle of a dog’s neck (that is, whether the dog is looking up or down) as an important behavioral clue. Consequently, activities such as eating or drinking appear to be less accurate when raised dog bowls are used, and activities such as sniffing and scratching, and self-licking can go undetected if performed in unusual positions. Slow-feed food bowls, collars attached to taut leashes, and loose collars with other heavy attachments can also cause misclassifications, but are often classified correctly nonetheless.
Other sets of activities simply present very similar accelerometer data, such as eating wet food, which can be confounded with drinking; or being pet by a human or riding in a moving vehicle, which can be confounded with scratching or self-licking; or even vigorous playing and ‘tug-of-war’, which can be confounded with shaking and other activities. These misclassifications become less common as the models improve, but in some cases confusion may be unavoidable. Some other activities are simply rare or unusual, for instance, drinking from a stream, drinking from a water bottle, or licking food off of a raised plate.
A different type of problem relates to activities that are ambiguous even to human labelers, such as the distinction between eating a small part of a meal versus eating a large treat. Similarly, when a dog repeatedly starts and stops an activity, it is often a matter of the labeler’s judgment whether to use a single long label or multiple short labels. Both of these types of labeling ambiguity can be very deleterious to certain classification metrics, even though it is questionable whether the system’s usefulness or real-world accuracy is affected.
Conclusion
We advanced the sensitivity and specificity for detecting drinking and eating behavior, and we demonstrate detection of licking, petting, rubbing, scratching, and sniffing, which to our knowledge have not been reported in a comparable manner. We demonstrated that system performance is not sensitive to collar position. In production, users reported high rates of true positives, consistent with the metrics measured via cross-validation on the crowd training database. The systems described in this work can further improve via the incorporation of additional training data and through the improvement of the underlying algorithms.
Acknowledgments
We are grateful to Leonid Sudakov and Jeannine Taaffe for their support and vision in enabling the Pet Insight Project; WALTHAM Petcare Science Institute for contributing extensive training data and for numerous helpful discussions; to the many participants in the Pet Insight Project; and to the Whistle, Kinship, and Mars Petcare organizations for supporting the development and publication of this work.