Objective Monitoring of Loneliness Levels using Smart Devices: A Multi-Device Approach for Mental Health Applications
aDonald Bren School of Information and Computer Sciences, University of California, Irvine
bInstitute for Future Health, University of California, Irvine
cDepartment of Psychological Science, University of California, Irvine
dDepartment of Computing, University of Turku, Turku, Finland
eDepartment of Cognitive Science, University of California, Irvine
fSchool of Nursing, University of California, Irvine
*Corresponding author; email: salarjafarlou@gmail.comAbstract
Loneliness is linked to wide ranging physical and mental health problems, including increased rates of mortality. Understanding how loneliness manifests is important for targeted public health treatment and intervention. With advances in mobile sending and wearable technologies, it is possible to collect data on human phenomena in a continuous and uninterrupted way. In doing so, such approaches can be used to monitor physiological and behavioral aspects relevant to an individual’s loneliness. In this study, we proposed a method for continuous detection of loneliness using fully objective data from smart devices and passive mobile sensing. We also investigated whether physiological and behavioral features differed in their importance in predicting loneliness across individuals. Finally, we examined how informative data from each device is for loneliness detection tasks. We assessed subjective feelings of loneliness while monitoring behavioral and physiological patterns in 30 college students over a 2-month period. We used smartphones to monitor behavioral patterns (e.g., location changes, type of notifications, in-coming and out-going calls/text messages) and smart watches and rings to monitor physiology and sleep patterns (e.g., heart-rate, heart-rate variability, sleep duration). We also collected participants’ loneliness feeling scales multiple times a day through a questionnaire app on their phone. Using the data collected from their devices, we trained a random forest machine learning based model to detect loneliness levels. We found support for loneliness prediction using a multi-device and fully-objective approach. Furthermore, behavioral data collected by smartphones generally were the most important features across all participants. The study provides promising results for using objective data to monitor mental health indicators, which could provide a continuous and uninterrupted source of information in mental healthcare applications.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This study did not receive any funding.
Summary of Updates:
1.Introduction
1.3.Current Research
Altogether, the intersection of mental health and technology offers an exciting opportunity to build upon foundational work in the field [33, 39, 42] by developing a more comprehensive monitoring framework capable of assessing the occurrence and predictability of loneliness. In this study, we develop a method of passively and continuously monitoring loneliness using smart wearables and smartphone devices. Using these devices together allows us to assess a wide range of behavioral patterns such as phone usage, communications, and locations as well as physiological patterns such as sleep indices and cardiovascular data. We leverage these devices to propose and evaluate a fully objective loneliness detection method that does not rely on user engagement. We also aim to examine which behavioral and physiological features were most predictive of loneliness for each individual. Finally, our third aim is to determine device efficiency of loneliness monitoring and detection.
2.Method
2.1.Participants
Full-time college students (N=30) at a West coast university between the ages of 18-22 were recruited to participate in an intensive longitudinal study investigating loneliness and mental health. Participants were eligible if they spoke fluent English and used an Android smartphone compatible with the Oura Ring and Samsung Active 2 watch. Students were ineligible to participate if they were married, had children, were returning to school after a three year hiatus, or if they experienced any severe forms of psychopathology (i.e., diagnosed with clinical depression or substance use disorders, psychosis, or any form of suicidal ideation). The exclusion criteria were intended to ensure a relatively homogenous sample of college students and individuals in emerging adulthood. We reasoned that students whose demographics were different (parents, older adults, and individuals returning to school) may have a different experience than emerging adults as they develop their identity and social networks. Students were recruited through faculty outreach where professors may then make announcements to their classes. Social media posts on institutional platforms (e.g., Facebook, Reddit, and Discord) were also used to recruit students. Interested participants reached out via email and were administered a screening survey to assess for moderate to severe depression or suicidal ideation. Individuals who met these criteria were then reached out by the clinical psychologist on our team (JB) for additional support and follow-up. Campus and local wellness resources were provided to all individuals who reached out or completed the screening survey.
2.2.Procedure
All procedures were approved by the researchers’ academic institutional review board (#2019-5153). For the purpose of this investigation, only relevant procedures and our monitoring phase will be described from the larger study. If participants met the inclusion criteria, they were then contacted to schedule their first in-person lab session in which they completed a baseline battery assessment of mental health, emotions, and well-being related measures. The mental health assessment included measures of psychopathology, on the basis of which participants were withdrawn from the study if the research team determined in consultation with JB that continued participation would be unsafe. During the first lab session, Oura Ring, Samsung Gear Sport watch, and smartphone apps related to the study were set up and downloaded for the participants. Participants were then instructed on how to use the devices and were guided through tasks that they would be completing for the next four weeks of their participation. Participants were instructed to wear their devices at all times throughout the study period, except under certain conditions (e.g., while charging devices or while performing intensive activities that risked damage to the watch). For the smartphone apps, participants downloaded AWARE that monitored their phone usage and a study-designed app, mSavorUs, which would prompt individuals to complete a brief survey five times per day, using an interval-based ecological momentary assessment (EMA) design. The survey included questions about affect, interactions with others, and feelings of loneliness, social isolation, and connectedness. Apps relating to the wearable devices were also downloaded onto participants’ phones for proper device functionality. Participants completed this monitoring phase for about 8-weeks and were compensated for their involvement in all components of the study.
2.3.Data Collection
To support our longitudinal study design, we built a platform capable of collecting and storing data efficiently. We designed a research dashboard that allowed the researchers to have constant access to the data being collected in order to monitor participants’ progress. The dashboard visually displayed the collected data and a summary of each participant’s daily activities. The information was then used to track and assess possible connectivity issues with the devices throughout the study duration. The collected data fell under three categories: objective physiological data, objective behavioral data, and subjective self-report questionnaires. Participants’ objective physiological data was collected using the Oura ring and Samsung watch. The Oura ring was used to assess information regarding the participant’s sleep patterns, and the Samsung watch was employed to conduct continuous physiological assessments throughout the day. The objective behavioral data was collected using the AWARE phone application. AWARE provides us with mobile sensory data and participants’ phone usage details, which can then be used to derive behavioral features. Subjective self-report data was collected using the mSavorUs phone application that was developed for the purpose of collecting subjective experiential data and delivering interventions in this study.
2.3.1.Objective Physiological Data using Samsung Watch and Oura Ring
The Oura Ring offers a range of metrics regarding the user’s physical activity, and sleep patterns [7]. However, Photoplethysmography (PPG)-based wearable devices, including the Oura Ring, can encounter noise during signal collection, particularly when utilized in a home-based monitoring setting. For our purposes, we used data from Oura that included the user’s sleep, physiology and physical activity. The final list of selected data included heart-rate related information during sleep, the estimated time of start of sleep after getting into bed, as well as the number of minutes with high, medium and low level of physical activity during the day (refer to Appendix Table A.3 for a detailed list of features of Oura ring used in this study). Samsung Active 2 watch operated on the Tizen open-source operating system [43], allowing for the development of custom data collection applications. This open-source platform enabled us to develop an application specifically designed for the watch to collect the data in a manner customized to the purposes of our study. The watch app activated every two hours to collect twelve minutes of PPG signals, used to extract heart rate and heart rate variability. The collected data was then synchronized with our server through wifi or bluetooth.
2.3.2.Objective Behavioral Data using the AWARE app and Subjective Experiential Data using mSavorUs
For the objective behavioral data, we used the AWARE app available for Android. This app’s sensors were configured to collect participants screen usage (e.g., lock, unlock time), application usage (e.g., notification, time of opening any app), their keyboard strokes pattern, battery level, communications (e.g., message/calls sending and receiving time) and GPS location [44]. Finally, for subjective experience, we designed and developed the mSavorUs application as a portion of a larger intervention-based study. This application [45] included assessments of loneliness and social isolation using mobile app interfaces and push notifications.
2.4.Data Analytic Plan
2.4.1.Feature Extraction
We developed Heart Rate (HR) and Heart Rate Variability (HRV) extraction methods from the raw PPG signals collected by the Samsung watch. The raw PPG signals are highly susceptible to environmental noise motion artifacts. To illustrate the problem, we plot two 60-second PPG signals. Figure 1 (a) is a clean PPG signal, showing heart beat oscillations. However, Figure 1(b) is a noisy PPG, distorted by the subject’s hand movements. Such distorted signals will result in unreliable HR and HRV features extraction. Therefore, we developed a PPG processing pipeline to address this problem.
The pipeline includes 3 main stages: signal quality assessment (SQA), signal reconstruction, and PPG peak detection. The SQA classifies PPG signals as “clean” or “noisy” by extracting five features from the signal, including interquartile range, standard deviation of the power spectral density, range of energy of heart cycles, average Euclidean distances, and average correlation between a template and heart cycles [46]. After we performed SQA, short-term “noisy” segments (less than 15 seconds) were reconstructed using a trained generative adversarial network (GAN) model [47]. The GAN model was trained to reconstruct noisy PPG using the information both in the distorted part and its proceeding clean signals. Then, a trained Dilated Convolution Neural Network (DCNN) was employed to detect the systolic peaks [48] and inter-beat intervals (IBI). Finally, HR and HRV-related features were extracted from the IBI signals. We access the Oura ring data through an application programming interface (API) that provides the processed features of sleep and activity. No further feature extraction or post-processing steps were done on the ring data (please refer to Appendix Tables A.3, A.4 for the full list and description of the features extracted from ring and watch).
We developed methods to extract behavioral parameters of the participants’ daily living through the longitudinal behavioral data collected by the smartphones. For calls, we used the duration sums and number of calls in each category (e.g., outgoing, incoming and voicemail). For location data we first located the participant’s house as the most frequently visited place during the nights of study. Then given a location time window, we extracted variance of latitude, variance of speed, mean of speed, number of places, home duration, outdoor duration, mean of outdoor duration (>=2 places), standard deviation of outdoor duration (>= 2 places), longest duration type other than home (>= 2 places), and total travel distance(>= 2 places). For messages and notifications, we counted the number of messages in each category. Table 1 demonstrates all of the categories and three samples for each of them (please refer to the Appendix Table A.3 for a detailed list of all features extracted from the phone.)
Some of the features from the devices were collected in a single assessment per day (i.e., sleep related features) whereas others were assessed with higher resolution, multiple times throughout the day (i.e., location, phone lock-screen, and PPG recordings; for a list of all features, see Tables A1-3 in the appendix). For data that was sampled multiple times per day, we tested different time windows (e.g., from 4 to 48 hours) for aggregating the values in order to pair with the subjective responses (i.e., self-reported loneliness used for labeling). Figure 2 shows how different window lengths for each modality have been used to compile data records. After extracting mentioned feature values given a time window, we selected the optimum time window for each feature based on the correlation with the target questionnaire response values.
2.5.Missing data
For intensive longitudinal studies assessing experiences in real-time, many factors could interrupt continuous data collection and result in missing data. These factors may result from participants forgetting to wear their devices or charge them overnight, or may result due to technical issues (e.g., server congestion or permission issues over the phone). Using the monitoring tools in the data collection server, we were able to track most of these issues over the course of study and solve them in the shortest time possible. After all these arrangements, having a certain degree of missing data is inevitable and must be taken care of in the analysis stage. In our analysis, we tested two data imputation methods to handle missing data. We replaced the missing values with the average of: A) two preceding and succeeding valid samples, or B) all valid values [49] of that participant. The latter method yielded a higher correlation with the target questionnaire response values for the majority of the features. Therefore, for consistency, we used this method to handle missing data for all of the features.
2.6.Classification
We defined the binary classification labels (true/false values) based on the participant’s self-reported loneliness rating (scaled from 0-100, 0 = not at all, 100 = extremely) as below or above the median to have a balanced number of labels. We developed a random forest [50] method to predict these classification labels using the 84 selected features. Random Forests is a machine learning algorithm that combines multiple decision trees to make more accurate predictions. Leveraging multiple decision trees enables this algorithm to process high dimensional data efficiently. Before training predictive models, we Z-normalized the features using each participant’s data in the training set to reduce interpersonal bias in feature space. For model evaluation, we used a one-subject-out policy for predictive modeling evaluation, where we used the most recent 50% of each participant’s data as a test set and the remaining data of that participant with all the other participants for the training model. We calculated the accuracy, F1-score, precision, recall, and mean squared error for all the test entries combined. We defined true positive and false positive predictions if our model detected the loneliness class respectively correct and incorrect (according to EMA labels). Similarly, true negative and false negative predictions were instances where the model did not detect the loneliness class while they were correct and incorrect, respectively. Given these terms we can define aforementioned performance metrics as:
2.7.Feature Importance
A limitation of using classic model evaluation metrics is that they lack insight into the dynamics of prediction. To address this limitation, the SHapley Additive exPlanations (SHAP) method was used to investigate the contribution of features to predictions given a model [51]. SHAP offers explainability and insight into the contribution of each feature used by a given model in the prediction stage. Given a trained model, this method generates numerical estimates (called SHAP values) that represent how much a feature value affects the output value of the model and what direction (toward either of the classification labels) this effect is. Variations of this method have been proposed for different machine learning models. Some of them consider the trained model as black-box and extract local explanations using efficient sampling in the feature space [52]. Other variations of SHAP analysis methods are proposed for specific machine learning models. In this work, we leveraged path dependent feature perturbation algorithms [53] developed for tree-based models. This method splits the feature space while recursively following the decision path for a given datapoint. This approach enabled us to enhance the accuracy and interpretability of our models, making them more effective tools for decision-making and analysis.
3.Results
3.1.Aim 1: Loneliness detection overall performance
We first present the performance of our proposed loneliness detection method enabled by using the multiple ubiquitous sensing devices AWARE, Oura ring, and Samsung watch. As previously mentioned, we conducted a two-month monitoring study with 30 participants who completed five self-report assessments per day, totaling to about 7,300 data points. The trained model was evaluated by comparing the estimated loneliness values with the corresponding ground truth values (i.e., collected by subjective questionnaires). This procedure was repeated for every participant; in other words, 30 different personal models were built. In the following, we report the aggregated results obtained from all the trained models. Our loneliness detection models obtained an accuracy of 82% Table 2, precision score of 81%, and value of 0.83 Area Under Curve (AUC). Higher AUC generally represents better classification capability (ranging from 0.5 showing random classification to 1 that is for a perfect classification). Figure 3 shows the confusion matrix of the loneliness detection models extracted from about 3,600 tested samples. The obtained true positives (detecting loneliness correctly) and true negatives (estimating no loneliness correctly) were considerably lower compared to the false positives and false negatives.
3.2.Aim 2: Explainability and Feature Importance Analysis
As discussed, we obtained loneliness detection accuracy of 82% using all the modalities from three devices (smartwatch, ring, and phone). In this section, we explore further into the analysis to investigate the effect of the features on the performance of loneliness detection using explainable machine learning techniques. To this end, we carried out SHapley Additive exPlanations [51] (SHAP) analysis to gain deeper insight into the detection models. SHAP values serve as a measure showing the significance of each feature in the model’s ability to perform the detection. As previously mentioned, we trained a personal model for each participant. To assess the impact of each feature on the loneliness detection among participants, we calculated the mean SHAP values across the test samples.
Figure 4 depicts the absolute SHAP values of the 20 features for each participant, with dark blue and light green hues showing highest and lowest values, respectively. Overall, the 6 most influential features from our prediction model were related to behavioral markers extracted from smartphones (i.e., the AWARE platform), while the HRV features – collected from the smart watch – had less impact. However, the influence of each feature varied across participants. For example, the number of notifications from lifestyle and communication applications had a significant impact on Participants 15 and 20, respectively, whereas these features did not exert as similar of an impact on the remaining participants. This inter-individual difference indicates the importance of personalization in loneliness detection methods.
The findings shown in Figure 4 further offer a broad understating of the impact of individual features on the loneliness detection. However, due to averaging the absolute values, this representation obscures the distribution and direction of the effects of each feature on each participant. Specifically, a positive SHAP value denotes that the corresponding feature shows a positive impact on the loneliness class detection, while a negative value contributes to the non-loneliness class detection. Thus, Figure 5a,5b shows the positive and negative SHAP values of two randomly selected participants, revealing the variation in the importance of the features. Loneliness class is represented on the right side of the y axis in this Figure, while non-loneliness is represented on the left side. The red and blue colors illustrate negative associations and positive associations respectively for a given feature. For example, physiological features were the most influential for Participant 8 (Figure 5a), whereas Participant 10’s most important features (Figure 5b) were behavioral features. It also shows that on average, Participant 8’s lower values of HRV metrics (represented by the blue color) are associated with higher feelings of loneliness (left side of the vertical line). These figures also indicate how the order and impact of features vary between the participants for loneliness detection. For example, for Participant 8, the most influential features are extracted from the smartwatch. However, this effect is less evident for Participant 10 (Figure 5b). For instance, decreases in the HRV CVSD feature (presented in blue color) results in loneliness for Participant 10. This is while higher values of this feature (red) results in non-loneliness class.
3.3.Aim 3: Loneliness Detection using Different Sets of Devices
We investigated the effects of different modalities on the performance of our loneliness detection method. To this end, in the training phase, we built four random forest models, three of which were trained with data collected from a single device (i.e., smart ring, smart watch, or smartphone), and one of which was trained using data from all three devices. Subsequently, the models were evaluated using test data from the same device(s). This assessment strategy allowed us to select the optimal set of devices based on the desired criteria, including performance, costs, or user burden. The performance of the models is indicated in Table 2. The smart ring (Oura) is small, lightweight, and easy-to-use, with a battery life of approximately one week following a full charge. The device provides sleep quality, physical activity, and nocturnal HR and RMSSD parameters with a rather high accuracy [54, 55, 35]. Although Oura obtained the highest recall value compared to all other cases, it showed the poorest accuracy, F1-score, and precision for loneliness detection. The second-best accuracy of a single device was obtained using the smart watch (Samsung), which enabled us to acquire HR and multiple HRV measurements at all times. In comparison to Oura, the watch obtained better precision but worse recall. The AWARE framework, which only uses smartphone logging and sensing features, obtained the highest accuracy and precision but the lowest recall compared to the Oura ring and Samsung watch. AWARE runs as a background application on the user’s smartphone. This passive sensing eliminates the need for user input, in contrast to wearable devices that require continuous wear. It should be noted however, that the AWARE data collection is limited to the behavioral and contextual parameters. For instance, our previous study [56] indicated that COVID-19 lockdown had an adverse impact on such passive smartphone-based data collection, as user mobility decreased (e.g., fewer location changes). Our results showed that multimodality has a positive impact on the overall performance of loneliness detection, providing an improved trade-off between precision and recall. Specifically, the three devices used in the study together obtained higher precision values compared to the Oura ring alone, and slightly higher recall values compared to Samsung Watch and AWARE alone.
4.Discussion
Advances in technology offer exciting future direction for mental health prevention and treatment. With the availability of passive sensing and assessments of people’s experiences, behaviors, and mental states in real-time and with greater ecological validity, researchers and medical professionals can have a better understanding of individualized experiences of mental health. Loneliness is an important predictor of physical health and mortality [57]. Furthermore, loneliness is associated with several sleep and physiological factors [58]. The U.S. has previously announced a loneliness epidemic [59] and several recent studies are focusing on loneliness across different populations [60]. To address loneliness requires a multimodal approach. Our study adopted a comprehensive approach using multiple devices to passively and continuously monitor people’s experiences in order to predict loneliness and assess how loneliness might look across individuals. Our findings in this study showed that loneliness detection using multimodal assessment with commercially available devices could yield an accuracy of 82%. We compare our proposed ML-based loneliness detection method with current state-of-the-art methods for fully and objectively (i.e., only using sensors from wearable and mobile devices) identifying loneliness. To the best of our knowledge, there exists only one article in the literature which attempts to objectively detect instantaneous loneliness (i.e., not daily or weekly classification) while having an acceptable sample size (N ≥ 10). It should be noted that in order to utilize a predictive model in just-in-time adaptive interventions (JITAI) [61], it is critical to deploy models with the ability of instantaneous (fine-grained) predictions. Wu et. al. [32] conducted a three-week data collection of 129 individuals using smartphones. The EMA was collected randomly up to four times a day in which participants were asked to select 4 options of loneliness levels. Using a random forest classifier, they obtained the average AUC of 0.73-0.74 for loneliness detection, which is 10% lower than our method’s AUC. Although their results showed how contextual data collected from smartphone sensors could be used in loneliness detection, they did not leverage physiological measurement and multi-modal assessment in their study. As presented in the following, physiological indicators are rich sources of information for loneliness detection models when coupled with contextual markers.
With respect to feature importance for predicting loneliness, we found that across devices, assessments of people’s phone behaviors using the AWARE app had the highest average values with loneliness across all participants. More specifically, people’s use of their phones either through notifications or engagement with phone apps and calls were highly associated with loneliness. In this respect, our findings do support the existing literature that has primarily used mobile app assessments of loneliness using people’s geolocation and phone interactions with others [33, 27, 32]. HRV parameters collected through the Samsung smartwatch were second in their prediction weights on loneliness.
Surprisingly, sleep measures as assessed through the ring had less weight as predictors of loneliness. This finding is somewhat inconsistent with existing literature. Past studies that have used subjective indices of sleep suggest a link between reported sleep disturbances and loneliness [40, 41]. Furthermore, studies that have used objective indicators of sleep (such as polysomnography, sleep watches, and neural scans) are consistent with subjective indices, where sleep quality and deprivation was also found to be associated with social withdrawal and loneliness [62]. One potential limitation of our work that may help to explain our finding is that our model included several features from assessments collected during the day, whereas the Oura ring was the only device that captured bedtime features. While behavioral and some of the activity data are being generated throughout the day, sleep data has values per day. This difference in the frequency of sleep data could suggest further exploration on specific data fusion techniques.
With respect to the nuances across behavioral features and cardiovascular features, one possible explanation for why behavioral features weighed more heavily is the growing literature examining the importance of context [63, 64, 65]. For example, the links across emotions, physiological assessments, and health vary widely across different racial/ethnic groups [66, 67]. Assessing these objective markers lack meaning if not measured with the context of the individual. Loneliness is highly related to a person’s perceived social interactions or lack thereof. If an individual feels slight changes in their interactions with others, whether it be in their use of phone apps or virtual engagement with others, it might have changes in their loneliness. Therefore, in having a more comprehensive approach to understanding how loneliness occurs, it still is needed to ground it in the context of social interactions and relationships.
Our aim was to obtain the best performance for objective loneliness detection using all available devices. In other words, in addition to assessing the performance of loneliness detection methods, it is vital to evaluate how each device performs individually. Through this process, researchers can gain insights on how to improve other metrics, such as feasibility and usability within the context of remote health monitoring studies. This information can assist researchers in selecting the optimal set of devices according to their specific requirements. Our findings indicate that multimodality yields a beneficial effect on the overall performance of loneliness detection, offering an enhanced balance between precision and recall. Additionally, multimodality offers competitive data fusion [68] in loneliness detection, which can enhance service availability. This enhancement is achieved through the utilization of three independent battery-powered devices, which ensures that loneliness detection remains operational even if one or two of the devices become unavailable due to factors such as battery depletion, technical problems, or decreased user mobility. Moreover, multimodality provides complementary fusion of sensing modalities [68], which can improve the explainability of the analysis. Each device is not directly dependent on other devices, but the data can be aggregated to provide a more complete depiction of loneliness under observation. The Oura ring collects physical activity and nocturnal health parameters, including sleep quality and HR. The Samsung watch provides nighttime and daytime physiological parameters, including HR and HRV. AWARE enables us to acquire behavioral information (e.g., social interactions) and contextual information (e.g., location). In other words, each device assesses different phenomena, and assessing all of these at the same time, in real-time, can provide a more comprehensive view of how loneliness occurs for each individual. By including each device in analyses, a potentially more accurate examination of associations between loneliness and diverse health and well-being parameters can be tested. Therefore, it can provide holistic actionable insights for health providers. Overall, multimodality is a promising approach for improving the performance, availability, and explainability of loneliness detection methods, however, it increases the cost and user burden.
5.Conclusion
In this paper we proposed a method for detecting loneliness in a multimodal fully objective setup using smart phones and commercially available wearable devices. We tested our method in a 2-month long study and showed that we can detect loneliness in a continuous way. Our results showed that smartphones and activity-related information during the day are among the most important features for most of the participants. Leveraging multi-modal assessments and passive sensing to continuously and objectively monitor human experiences can help reduce participant burden and support on-going efforts to design personalized treatment and programming to support well-being. Assessing human experiences in real-time can allow us to take a more preventative approach for health treatment rather than a reactive approach. The work conducted in this study was restricted to a demographic of college students and limited to a 2 month period, which therefore suggests further investigation for confirming the generalizability. We are planning to further investigate device-specific feature extraction techniques and personalized modeling for our proposed method.
Data Availability
All data produced in the present study are available upon reasonable request to the authors.