Development of a Simplified Smell Test to Identify Patients with Typical Parkinson’s as Informed by Multiple Cohorts, Machine Learning and External Validation
1Neuroscience Program, Ottawa Hospital Research Institute, Ottawa, ON., Canada
2Methodological, Implementation Research Program, Ottawa Hospital Research Institute, Ottawa, ON., Canada
3the Methods Centre, Ottawa Hospital Research Institute, Ottawa, ON., Canada
4University of Ottawa Brain and Mind Research Institute, Ottawa, ON., Canada
5Department of Cellular and Molecular Medicine, University of Ottawa, Ottawa, ON, Canada
6Department of Medicine, University of Ottawa, Ottawa, ON, Canada
7School of Epidemiology and Public Health, Faculty of Medicine, University of Ottawa, Ottawa, ON, Canada
8Division of Neurology, Department of Medicine, The Ottawa Hospital, Ottawa, ON., Canada
9Paracelsus-Elena-Klinik, Kassel, Germany
10Department of Neurology, University Medical Center Goettingen, Germany
11Department of Neuroscience, Carleton University, Ottawa, ON, Canada
12Memory Program, Bruyère Research Institute, Ottawa, ON, Canada
13Deutsches Zentrum für Neurodegenerative Erkrankungen (DZNE), Goettingen, Germany
14Aligning Science Across Parkinson’s (ASAP) Collaborative Research Network, Chevy Chase, MD 20815
*Corresponding authors Correspondence to: Dr. Juan Li Ottawa Hospital Research Institute juli@ohri.ca; Dr. Brit Mollenhauer Paracelsus-Elena Klinik brit.mollenhauer@med.uni-goettingen.de; Dr. Michael Schlossmacher Ottawa Hospital Research Institute mschlossmacher@toh.caABSTRACT
Background
Reduced olfaction is a common feature of patients with typical Parkinson disease (PD). We sought to develop and validate a simplified smell test as a screening tool to help identify PD patients and explore its differentiation from other forms of parkinsonism.
Methods
We used the Sniffin’ Sticks Identification Test (SST-ID) and the University of Pennsylvania Smell Identification Test (UPSIT), together with data from three case-control studies, to compare olfaction in 301 patients with PD or dementia with Lewy bodies (DLB) to 36 subjects with multiple system atrophy (MSA), 32 individuals with progressive supranuclear palsy (PSP) and 281 neurologically healthy controls. Individual SST-ID and UPSIT scents were ranked by area under the receiver operating characteristic curve (AUC) values for group classification, with 10-fold cross-validation. Additional rankings were generated by leveraging results from eight published studies, collectively including 5,853 unique participants. Lead combinations were further validated using (semi-)independent datasets. An abbreviated list of scents was generated based on those shared by SST-ID and UPSIT.
Findings
We made the following five observations: (i) PD and DLB patients generally had worse olfaction than healthy controls, as published, with scores for MSA and PSP patients ranking as intermediate. (ii) SST-ID and UPSIT scents showed distinct discriminative performances, with the top odorants (licorice, banana, clove, rose, mint, pineapple and cinnamon) confirmed by external evidence. (iii) A subset of only seven scents demonstrated a similar performance to that of the complete 16-scent SST-ID and 40-scent UPSIT kits, in both discovery and validation steps. Seven scents distinguished PD/DLB subjects from healthy controls with an AUC of 0.87 (95%CI 0.85-0.9) and PD/DLB from PSP/MSA patients with an AUC of 0.73 (95%CI 0.65-0.8) within the three cohorts (n=650). (iv) Increased age was associated with a decline in olfaction. (v) Males generally scored lower than females, although this finding was not significant across all cohorts.
Interpretation
Screening of subjects for typical Parkinson’s-associated hyposmia can be carried out with a simplified scent identification test that relies on as few as seven specific odorants. There, the discrimination of PD/DLB subjects vs. age-matched controls is more accurate than that of PD/DLB vs. PSP/MSA patients.
Funding
This work was supported by: Parkinson Research Consortium; uOttawa Brain & Mind Research Institute; and the Aligning Science Across Parkinson’s Collaborative Research Network.
Research in context
Evidence before this study
Chronic hyposmia is a common feature of Parkinson disease (PD) and dementia with Lewy bodies (DLB), which often precedes motor impairment and cognitive dysfunction by several years; it is also frequently associated with α-synuclein aggregate formation in the bulb. The presence of hyposmia increases an individual’s likelihood of having -what has recently been proposed as- a neuronal synucleinopathy disease, by >24-fold. Despite the strong association of PD with reduced olfaction, little is understood about it clinically, such as whether it is affected by sex and age, and whether hyposmia of PD is associated with the same scent identification difficulty seen in other conditions that present with parkinsonism. Moreover, due to its time-consuming nature and traditional administration by healthcare workers, extensive olfactory testing is not routinely performed during neurological assessments in movement disorder clinics.
Added value of this study
We analyzed the performance of both the Sniffin’ Sticks Test kit and UPSIT battery to discriminate between healthy controls, patients with PD/DLB and those with MSA or PSP. Comparison to and juxtaposition with eight other published studies allowed for the generation of a markedly abbreviated smell identification test that unified both tests, as described. Group classification performance by each scent and its distractors was further analyzed using machine learning and advanced Item Response Theory methods. Relations between each scent tested, sex and age were analyzed for the first time. Our findings suggest concrete steps to be implemented that would allow for simplified, routine olfaction testing in the future.
Implications of all the available evidence
Olfaction testing has emerged as an important neurological assessment part when examining subjects with Parkinson’s and those at risk of it. A simple, validated smell test containing fewer scents than current options could facilitate rapid testing of olfaction in clinic settings and at home, without supervision by healthcare workers. The usefulness of such a non-invasive test in population health screening efforts could be further enhanced when coupled to a self-administered survey that includes questions related to other risk factors associated with PD. As such, large-scale community screening and applications to routine practice in family doctors’ offices as well as in specialty clinics could be made operationally feasible and cost-effective.
Article notes
Competing Interest Statement
The authors have declared no competing interest.
Funding Statement
This work was supported by funding from Parkinson Canada (to T.M., D. M., M.G.S; 2018; to J.L.; 2019-2021), Michael J. Fox Foundation for Parkinson's Research (to T.M., D. M., M.G.S), Department of Medicine (T.M., T.R., D.M., M.G.S.), The Ottawa Hospital Foundation (Borealis Foundation to J.L.) and the Uttra & Subhash Bhargava Family (M.G.S.), the Paracelsus-Elena-Klinik Kassel, Parkinson Fonds Deutschland, and the Deutsche Parkinson Vereinigung (B.M.; C.T.). The study was also funded by the joint efforts of The Michael J. Fox Foundation for Parkinson's Research (MJFF) and the Aligning Science Across Parkinson's (ASAP) initiative. MJFF administers the grant [Grant ID: ASAP-020625] on behalf of ASAP and itself. The funders had no role in the design and execution of the study; the collection, management, analysis, and interpretation of the data; the preparation, review, or approval of the manuscript; and the decision to submit the manuscript for publication.
INTRODUCTION
Hyposmia is a common non-motor sign of Parkinson disease (PD) and dementia. The reported prevalence of olfaction loss in PD ranges from 45% to >90% based on populations selected, testing methods, and threshold criteria.[1] Chronic hyposmia is also described as predictive, with reduced olfaction preceding PD diagnosis by 4-20 years.[1]-[3] Olfactory testing may also help in the differentiation of parkinsonian syndromes.[4] Several screening tools and predictive models for the incidence of PD have included subjective or objective assessments of olfaction.[5]-[9]
Two commonly used smell tests for evaluating olfactory function include the University of Pennsylvania Smell Identification Test (UPSIT),[10] a battery of 40 scratch-and-sniff questions that are self-administered, and the Sniffin’ Sticks Test (SST) battery,[11],[12] with three subtests (for Identification; Discrimination; Threshold) comprising 16 scents each, of which the administration usually involves a trained supervisor. The SST-Identification (SST-ID) and UPSIT are comparable in that they both assess one’s ability to identify a range of scents. Within SST, SST-ID is more commonly used in PD cohort studies and known to have better diagnostic performance than the SST-Discrimination subtest.[13]
Smell test kits were initially developed to assess olfaction in the general population but have been increasingly used in research settings that study disorders of the brain. Using different cohorts and methods, some studies have ranked odorants in the UPSIT[14]-[17] and SST-ID kits[13],[18]-[20] by their diagnostic performances, and reported that certain subsets of scents appeared to have equal or better performance than the entire complement of 40- or 16-scents-based tests. Other studies[21]-[24] have examined subsets of UPSIT but without any details on scent ranking. Moreover, external validation was frequently missing in these analyses, proposed scent combinations were found to be cohort-specific without agreement across different studies,[16],[25] and the role of distractors (versus the correct scent offered) in such multi-choice settings were understudied. Furthermore, analyses of UPSIT and SST-ID kits were always conducted separately despite the similarities between the two tests. Finally, olfaction scores in patients with other, atypical forms of parkinsonism were not assessed in PD-centric studies.
In the current work, we aimed to assess olfaction performances in commonly encountered forms of parkinsonism, to assess the key features of both UPSIT and SST-ID odorants, to explain any observed differences between them, and to develop a simplified smell test to unify both kits using proper internal and external validation steps. To this end, eight published scent rankings,[13]-[20] which collectively examined 5,853 participants, were incorporated into our study to make the proposed abbreviated test generalizable and to avoid overfitting. We also added Item Response Theory-based analyses when examining the behaviour of participants’ responses to multiple choices provided for each scent; this, to enhance sensitivity and specificity of future test versions. Lastly, we analyzed the effects of age and sex on performance in scent identification.
METHODS
The study was conducted in adherence with STARD[26] guideline, see Supplemental Table 1.
Source of data and participants
We used de-identified data from three observational, retrospective, case-control studies: the “De Novo Parkinson disease study” (DeNoPa);[27] Ottawa (PREDIGT) Trial; and Prognostic Biomarkers in Parkinson’s Disease Study (PROBE).[28] Detailed description and inclusion/exclusion criteria of these three cohorts are provided in Supplemental Methods; their demographic and diagnostic characteristics are summarized in Table 1. Data of the cross-sectional Ottawa Trial study, baseline data of the longitudinal PROBE study, and three visits of the longitudinal DeNoPa study (baseline; 48-month; and 72-month follow-up visits) were used. Patients with PD or dementia with Lewy bodies (DLB) (DeNoPa: n=129; Ottawa Trial: n=70; PROBE: n=102), subjects with multiple system atrophy (MSA) or progressive supranuclear palsy (PSP) (DeNoPa: n= 9; Ottawa Trial: n=6; PROBE: n=53), and neurologically healthy controls (HC) (DeNoPa: n=109; Ottawa Trial: n=118; PROBE: n=54) were included. Diagnostic criteria were applied according to international guidelines; [27]-[30] diagnoses were revised, as necessary, during longitudinal follow up visits or by independent raters who were blind to olfaction test results. Most study participants with PD in the three trials were classified as Hoehn and Yahr stage II-III. No participant overlap existed between the three studies.
Analyses of these deidentified cohort data were approved by Investigational Review Boards at Paracelsus-Elena-Klinik (Kassel, Germany) in Frankfurt, Hesse (FF 89/2008), at The Ottawa Hospital (Ottawa, Canada; 20180010-01H) as well as all PROBE Study-affiliated sites, with the participants’ consent.
Study assessments and statistical analyses
Sniffin’ Sticks test (SST)
Used in DeNoPa, the SST comprises a supervised test for one’s sense of smell using pen-like odor dispensing devices, as administered in a clinic setting.[11],[12] The test has three subtests: Identification (SST-ID), Discrimination (SST-DS), and Threshold (SST-TH), each with 16 odorants. In SST-ID, subjects are presented a stick and choose the scent from four options (one correct, three distractors). SST-DS is performed using triplets of odorants that are of similar intensity and hedonic tone, where subjects are required to identify which stick of the triplet has a different scent from the other two. SST-TH is performed using triplets of sticks where only one is filled with odorant at a certain dilution whereas the other two are filled with odor-free solvent. SST-TH determines at what dilution subjects can consistently identify the odorant-filled stick. The entire SST (in German) was completed by all DeNoPa participants at their baseline visit, and the SST-ID subtest was re-administered at 48-month and 72-month follow-up visits.
The University of Pennsylvania Smell Identification Test (UPSIT)
UPSIT was used in the Ottawa Trial and PROBE; this self-administered kit contains 40 scratch-and-sniff questions; each question has one correct answer and three incorrect distractors.[10]
Data preparation
Observations that had no valid response to SST-ID or UPSIT were removed from the analysis. Observations with incomplete responses were imputed with 0s, indicating incorrect responses. A dichotomous response-based transformation (0 = incorrect response; 1 = correct response) was used to calculate the sum scores and assess discriminative performances for each scent. The exact indices of chosen options were used for Item Response Theory analysis (see below).
Demographic and diagnostic characteristics
Demographic and diagnostic characteristics of the study cohort were summarised using n (%) or median (the interquartile range (IQR)). The reported p-values represented the significance from corresponding Fisher’s exact test or Kruskal-Wallis rank sum test, with false discovery rate correction for multiple testing; p-values smaller than 0.05 were considered significant.
Comparing different smell tests
Score distributions of corresponding smell tests in each subject group were illustrated using Cummings estimation plots.[31] The raw UPSIT and SST-ID scores were also normalized into percentiles based on age and sex, and hyposmia was defined by SST-ID percentile ≤ 10%[32] or UPSIT percentile ≤ 15%.[10] Discrimination performances of these subtests were compared using area under the receiver operating characteristic (ROC) curve (AUC) values with bootstrap estimated 95% confidence interval (CI),[33] in order to distinguish diagnostic groups. Table 2 also reports optimal thresholds and their associated sensitivity and specificity that correspond to the maximum Youden indices.[34]
Machine learning workflow of developing and validating the abbreviated smell test
Figure 1 illustrates the machine learning workflow of ranking UPSIT/SST-ID scents, developing and validating respective simplified tests and the unified abbreviated test. Data of the Ottawa Trial and baseline data of DeNoPa were used as discovery cohorts, and baseline data of PROBE and follow-up data of DeNoPa were used for (semi-)independent validation.
- Ranking the individual scents: Using the corresponding discovery cohorts with 10-fold cross validation (see Supplemental Methods), individual scents in SST-ID and UPSIT were ranked separately based on their AUC values in differentiating PD/DLB patients from healthy controls. To control over-fitting, the SST-ID and UPSIT scent rankings from our study was compared with eight external rankings,[13]-[20] four for each test, and two final lists were generated by averaging internal and external rankings. Eleven scents are shared by both smell tests, and an additional “Shared” ranking was constructed using their respective positions in the averaged SST-ID and UPSIT rankings.
- Developing and validating the best-performing, simplified tests: For each scent ranking, beginning with the highest-ranked odorant, subsets were constructed by adding one scent at a time in descending ranking order. A total of 95 and 220 distinct SST-ID and UPSIT subsets with various numbers of scents were compared, using their AUC values in distinguishing PD/DLB from HC, to develop the best-performing simplified tests, including one that unified both smell tests. The resulted abbreviated smell tests were also validated using (semi-)independent datasets.
Exploring observed differences in scent performance
The percentage of correct scent identification within each subject group was calculated. These percentages were further compared to examine the relationship of scent identification versus sex and age (using spline smoothing, see Supplemental Methods). Furthermore, Item Characteristic Curves (ICCs) for each scent within PD/DLB and HC groups from baseline DeNoPa and Ottawa Trial visits were used to analyze scent performance and the influence of distractors (see Supplemental Methods for more details).
Statistical analyses were performed using ‘R’ (version 4.3.1). Cummings estimation plots were generated using ‘dabestr’,[36] and ICC curves were generated using ‘TestGardener’;[37] all other plots were generated using ‘ggplot2’.[38] The package ’pROC’[33] was used for ROC and AUC, and ‘fda’ was used for spline smoothing.[39]
RESULTS
Performances of individual scents differ in discriminating PD/DLB from HC
We next sought to determine whether subsets of scents were more informative to discriminate between PD/DLB and controls compared to the complete 16-scent (SST-ID) or 40-scent (UPSIT) tests. Figure 3(a) shows the distribution of AUC values for each SST-ID scent across 10 folds using baseline data from the DeNoPa cohort. Clusters of scents identified included banana and mint as the two most discriminative scents (individual AUC values, >0.725), followed by anise, coffee, licorice, fish and rose in the second-most discriminative cluster. Compared with SST-ID scents, clustering was less obvious for the UPSIT scents (Figure 3(c)), whose AUC values ranged between 0.5 to 0.77. For the Ottawa Trial cohort, the top-7 UPSIT scents for identifying PD/DLB patients included rose, wintergreen, root beer, licorice, dill pickle, mint, and grass.
The observed differences in each scent’s discriminative performance were further examined by visualizing the percentages of correct scent identification within each diagnostic group (Figure 3(b), (d)) and by the percentage differences between HC and PD/DLB groups (Supplemental Figure 2). Regardless of the study cohort and smell test used, PD/DLB patients showed significantly lower percentages of correctly identifying each scent than control subjects. Scents that were easy to identify in the HC group but difficult for the PD/DLB group (i.e., larger percentage differences as in Supplemental Figure 2) had larger single-scent AUC values. Scents had poorer discriminative performances when both HC and PD/DLB groups found them easy (e.g. SST-ID: orange, UPSIT: leather) or difficult (e.g. SST-ID: apple, UPSIT: lemon).
Therefore, two rankings for scents from the SST-ID and UPSIT kits were constructed. To overcome the inherent risk of overfitting due to the modestly sized cohorts used, external evidence was introduced. Figure 4 compared the scent rankings from this study with previously published studies, and two “Average” rankings were derived. For the SST-ID kit, different studies -despite their differences in cohort design and methods applied (Supplemental Table 1)-showed consensus in that anise, licorice, mint, banana, coffee, fish and rose were the most discriminative scents in distinguishing PD/DLB subjects from HCs (Figure 4(a)). For the UPSIT battery, however, its related studies showed less agreement on their scent rankings (Figure 4(b)), which could be partially explained by results shown in Figure 3(c); there, many UPSIT scents showing similar performances and revealed fewer clusters than did SST-ID-based scents. Nonetheless, the top-7 UPSIT scents in the final “Average” ranking were coconut, clove, wintergreen, banana, licorice, grass and cherry. With respect to a possibly common, simplified list, we noted that there are eleven scents shared by SST-ID and UPSIT, and thus, an additional “Shared” ranking was generated to construct a unified, abbreviated smell test (Table 3).
Item Characteristic Curves further reveal details for each scent and the influence of distractors
The findings above revealed subsets of scents that were relatively discriminative (from a PD perspective), which could suggest a disease-specific and/or scent processing-related change that is linked to participant performance. However, for multiple-choice tests like SST-ID and UPSIT kits, the selection of distractors paired with each scent could also influence odorant performance. Item Characteristic Curves (ICCs) can help address this, especially when the scent is shared by different tests.
In the current context, mint and licorice were two well-performing scents, and their ICCs (Figure 5(1)-(4)) showed similar characteristics in that HCs generally correctly identified them, while PD/DLB patients had more difficulty in choosing the right option. However, there were also some noteworthy differences: when scoring on the scent for ‘mint’, PD/DLB patients could rule out ‘chive’ and ‘onion’ in SST-ID and ‘fruit punch’ in UPSIT, indicating that they detected some scent, but it was not declarative enough to choose ‘mint’. However, for ‘licorice’, particularly in the UPSIT kit, there was strong evidence of random guessing whereby patients couldn’t detect any scent to help favor or eliminate an option. Here, ICCs of scoring by HCs also eliminated the possibility of the corresponding pen (SST-ID) or encapsulated sticker (UPSIT) being defective.
The scent for ‘banana’ ranked 1/16 in DeNoPa but only 34/40 in the Ottawa Trial (Figure 3); however, these inconsistent performances were not due to differences in distractors. The option ‘cherry’ distracted many PD/DLB patients and HCs in the Ottawa Trial, but not in the DeNoPa study (Figure 5(5)-(6)). Here, cohort- or odorant (e.g., its concentration or composition for the artificial scent)-related differences might be more plausible explanations.
‘Orange’ and ‘lemon’ were both ranked low in the two tests but for different reasons (Figure 5(7)-(10)): ‘orange’ in SST-ID was too easy, even for hyposmic PD/DLB patients. ‘Orange’ in UPSIT, however, had different distractors that were active within PD/DLB (‘bubble gum’) and HC (‘turpentine’). For ‘lemon’, the distractors of ‘grapefruit’ in SST-ID and ‘motor oil’ in UPSIT confused both patients and healthy persons. Such ICC results within the normosmic control group might be evidence of a flawed odorant choice or an explanation that is rooted in chemical manufacturing of the scent. In Supplemental Figures 3-5 we listed the ICCs for all other scents.
Development and validation of abbreviated smell tests
Based on scent rankings in Figure 4 and Table 3, 95 SST-ID subsets and 220 UPSIT subsets were compared by their AUC values in distinguishing PD/DLB patients from HCs. Figure 6 shows the results within each internal and external validation set. When using an increasing number of highly rank-ordered scents, we observed that the corresponding AUC values for odorant subsets increased steeply for the first four; surprisingly, any improvement in performance thereafter was marginal. Compared with other published rankings, the “Average” rankings as well as their subsets appeared to be more discriminative with robust performance in all the validation sets. Considering a potential trade-off between subset performance and number of scents administered, the SST-ID subset with seven scents and UPSIT subset with ten scents, based on their corresponding “Average” rankings, emerged as the best-performing test batteries.
With respect to potentially developing a unified smell test that could be applied to large study populations (those examined here were based in North America and Europe of predominantly White ethnicity), the subset of seven scents from the “Shared” ranking (Table 3) with the highest performance in all validation sets comprised licorice, banana, clove, rose, mint, pineapple and cinnamon. Mean (standard deviation) scores and AUC values of this subset within each cohort are reported in Table 4. Note, due to inherent cohort differences, the AUC value in the diagnostic classification of PD/DLB versus HC when combining all three cohorts was 0.87 (95%CI 0.85-0.9), i.e., slightly lower than those of each separate cohort. In the combined cohort, the cut-off value for distinguishing PD/DLB versus HC that corresponded to the maximum Youden index was 4.5 (score range 0-7), the resulted sensitivity and specificity were 0.76 and 0.85, respectively.
Performances of scents in discriminating PD/DLB from MSA/PSP
The same workflow was applied to the PROBE cohort to generate a subset of scents specialized in distinguishing between PD/DLB versus MSA/PSP patients. There, a subset with ten scents (clove, dill pickle, cinnamon, soap, rose, pizza, root beer, turpentine, gasoline, licorice) achieved an AUC value 0.78 in the validation set within PROBE, a promising improvement when compared with the entire 40-scent UPSIT test (AUC = 0.68; Supplemental Figure 6). Independent validation is needed to retest the usefulness of this subset given the small number of MSA/PSP subjects enrolled in the other two cohorts studied herein.
Assessment of age and sex on scent identification
We also investigated the influence of age and sex on scent identification. Supplemental Table 2 shows the coefficients of the linear regression for the relationship between smell test score with age, sex, and diagnostic groups within each cohort. Not surprisingly, progression in age significantly lowered olfaction across all groups, and males generally had a worse sense of smell than their female counterparts, although the latter was not significant across the three cohorts.
When focusing just on the eleven scents shared by both tests, relationships between scent identification and sex were further evaluated by comparing the percentages of correct scent identification across groups (Figure 7(a)). In line with the regression results, females showed higher percentages of correct identification than males for most of the scents, except for cinnamon, turpentine, and leather. Next, we compared the probability of identifying each scent correctly across ages between the PD/DLB and HC groups (Figure 7(b)). As expected, older participants showed decreasing percentages for identifying specific scents correctly. The fitted lines for PD/DLB and HC groups were usually in the same direction and of similar slopes, with some exceptions, but these were not consistent across all three cohorts.
DISCUSSION
To our knowledge, this is the most comprehensive study to date describing olfactory dysfunction in late-onset, typical PD and two other neurological disorders presenting with parkinsonism using both the multimodal SST battery and UPSIT kit. The principal take home message from this study is that when probing for hyposmia in PD, the following points matter: PD/DLB patients had worse olfaction than healthy subjects, and scores of MSA/PSP patients were intermediate; there was no observed difference in olfaction between MSA and PSP patients; scent identification testing is sufficient, and threshold as well as discrimination testing could be omitted when screening populations for PD using the SST kit; fewer scents can reduce examination time and test taking fatigue without sacrificing diagnostic accuracy; the selection of specific scents should be informed by their discriminative performance in specific group classification efforts; random guessing could lower diagnostic accuracy; and from a test design perspective, choices provided as distractors influence scent performance. Importantly, we found that a simplified smell test, with specific scents, is sufficient to identify PD/DLB-linked hyposmia. Such a test, which is now being piloted by us, holds the potential to facilitate olfactory testing in the clinical setting, used for at-home testing and population-based screening methods.
In developing and validating a simplified smell test for this purpose, we used a machine learning approach and found that only seven scents (licorice, banana, clove, rose, mint, pineapple, and cinnamon) were required to approximate the diagnostic performance of administering the 16-scent SST-ID or 40-scent UPSIT batteries and the value of adding more scents was negligible. Moreover, a test kit for rapid screening with nearly similar performance could be constructed using as few as four scents.
We demonstrated the impact distractors have on detecting specific scents using IRT analysis. We uncovered uncertainty in eliciting a choice for some scents, even for HCs with intact olfaction. This could be explained by the difficulty of biological scent discrimination or the to-be-improved selection of artificial odorants. By extension, our analyses revealed the opportunity to remove ill-performing scents, e.g., orange and lemon, from currently used kits.
Unexpectedly, we found a high level of guessing among PD patients for licorice, indicating patients’ difficulty in detecting this scent. SST-ID and UPSIT batteries are multiple choice-based tests, in which participants are instructed to always choose even when they cannot smell anything; such random guessing will introduce errors into data sets. Advanced IRT methods can treat missing responses as an additional option; administrators of tests would then prefer the participants to leave any uncertain questions unanswered rather than forcing a guess. However, for the future administration of olfaction tests, or for designing a new one, we would suggest adding an extra choice, such as “I cannot identify the scent” to reduce random guessing. Based on our experience in administrating smell tests, the extra option would also help improve participant experience and eliminate frustration, especially for patients with severe hyposmia.
An easy-to-administer, inexpensive, sensitive and non-invasive smell test (with 4-7 scents) could have important clinical usefulness, particularly when coupled with a short, self-administered questionnaire capturing demographic information and known risk factors of developing PD.[9] Such a questionnaire could also explore other factors leading to hyposmia unrelated to neurodegeneration, e.g., previous nasal injuries, microbial infections, seasonal allergies, and chronic exposure to air pollution, to augment specificity for PD. Upon validation, such a kit could be used as the initial step of large-scale community screening, or in routine clinic practice of a movement disorders-oriented clinic, or for early detection within a family medicine office. When it comes to screening efforts for PD, more invasive and expensive tests, e.g., the α-synuclein seeding amplification assay from cerebrospinal fluid (CSF), or the administration of a dopamine transporter scan, could be administrated as an additional step to increase screening accuracy further, such as when considering enrolling PD subjects into specific, disease-modifying clinical trials.[40]-[44]
Despite the findings regarding scent ranking and subset analyses, it remains unclear whether a specific PD olfaction deficit exists, rather than a global reduction in olfaction, and what the underlying mechanisms would be. We and others recently found that olfactory deficits were significantly associated with positivity on the α-synuclein seeding amplification assay in CSF samples, suggesting that patients may have an underlying disease linked to the dysregulation of SNCA expression and/or protein processing.[45],[46]
Mechanistically, it remains unknown as to how chronic hyposmia arises in PD/DLB (and REM Sleep Behaviour Disorder) as well as some MSA/PSP subjects, at what age it begins, the cause and its underlying circuit-based and molecular mechanisms, and whether olfactory deficits are shared for specific scents among persons with typical PD versus those with dementia. Large scale population screening, including with a simplified testing battery derived from SST-ID and UPSIT kits, could begin to answer some of these questions.
STRENGTH
In this work, olfaction performance using SST and UPSIT batteries was studied in detail. Both internal and external validation efforts were used to test/retest performance and avoid overfitting. Scent ranking derived from this study was also compared with eight previously published studies. By incorporating results from external studies, the unified abbreviated olfaction test kit of using just seven scents emerged as very generalizable. The observed performance differences for each scent in group classification was explained using complementary techniques and further studied considering sex and age. Our study also explored scent identification performance at the level of choices provided, highlighting the importance of distractors, which allows for future improvement in the design of test kits.
LIMITATIONS AND FUTURE WORK
Performance of the unified abbreviated smell test was estimated by simulation: responses of the related scents were partitioned from the SST-ID and UPSIT data sets, and the original distractors from SST-ID and UPSIT kits were used. Real performance of the abbreviated test, when manufactured as a stand-alone product with distractors determined for each scent, should be assessed in the same cohorts, in new studies, and in routinely run clinic settings. Further, more data are needed to validate the scent ranking and the associated subset developed here for distinguishing PD/DLB versus MSA/PSP patients.
The three cohorts used are highly homogeneous cohorts with most participants being white. Although scent rankings and the selection of a simplified smell test have been rigorously developed and validated with external information incorporated, future calibration and cultural adaption efforts will be necessary when applying them to other populations, especially those that differ from Western Europeans.
Case-control studies have an inherent potential for selection bias in their recruitment. Especially because of the age- and sex-matched design, age- and sex-effects were likely underestimated. Population studies, such as in the initial community screening effort undertaken by PARS planners and the ‘PPMI hyposmia’ effort, could provide complementary data sets; however, they too have potential setbacks: as the majority of participants will have a normal sense of smell, the score distribution could be highly skewed; a low percentage of PD patients will make the resulting dataset imbalanced; when smell test data are reduced to a single sum score, sub-analyses will be difficult to complete, which will limit interrogations of data between established cohorts.
For screening purpose, a one-time administered smell test may not be informative enough to assess a subject’s sense of smell completely, because other factors, such as temporary olfaction reduction/loss due to infection, seasonal allergies, occupational exposure and/or drinking, eating, smoking before taking the test, could skew results. Retesting at appropriate time intervals will be required for even higher accuracy.
Supporting information
ACKNOWLEDGEMENTS
The authors acknowledge the commitment of study participants in the three cohorts and are grateful to all clinical research coordinators at all the study sites. Responsible supervisors for the:
DeNoPa Study: Paracelsus-Elena-Klinik: Brit Mollenhauer.
Ottawa Trial Study: The Ottawa Hospital: Michael Schlossmacher, Élisabeth Bruyère Hospital: Andrew Frank.
PROBE Study: PROBE Steering Committee: Voyager Therapeutics: Bernard Ravina; Brigham and Women’s Hospital: Clemens Scherzer, University of Ottawa: Michael Schlossmacher, Avid Radiopharmaceuticals: Andrew Siderowf, University of Rochester: David Oakes, Arthur Watts; Institute for Neurodegenerative Disorders: Kenneth Marek; Georgetown University: Ira Shoulson.
We thank Nathalie Lengacher for help in graphic design; Nadine Mauri, Nancy MacDonald, and Yoobin Lee for help in data management/input. This work was supported by funding from Parkinson Canada (to T.M., D. M., M.G.S; 2018; to J.L.; 2019-2021), Michael J. Fox Foundation for Parkinson’s Research (to T.M., D. M., M.G.S), Department of Medicine (T.M., T.R., D.M., M.G.S.), The Ottawa Hospital Foundation (Borealis Foundation to J.L.) and the Uttra & Subhash Bhargava Family (M.G.S.), the Paracelsus-Elena-Klinik Kassel, Parkinson Fonds Deutschland, and the Deutsche Parkinson Vereinigung (B.M.; C.T.). The study was also funded by the joint efforts of The Michael J. Fox Foundation for Parkinson’s Research (MJFF) and the Aligning Science Across Parkinson’s (ASAP) initiative. MJFF administers the grant [Grant ID: ASAP-020625] on behalf of ASAP and itself.
The funders had no role in the design and execution of the study; the collection, management, analysis, and interpretation of the data; the preparation, review, or approval of the manuscript; and the decision to submit the manuscript for publication. We are grateful for the ongoing support and feedback from people with lived experiences, such as through the board of the Parkinson’s Research Consortium Ottawa and members of Partners Investing in Parkinson’s Research, and to Drs. P. Wells and D. Lewis for their ongoing encouragement.
CONFLICT OF INTEREST STATEMENT
In 2024, MGS co-founded NeuroScent Inc to develop a home-based testing platform for hyposmia.
CODE AVAILABILITY
The code for data analyses and figures is publicly accessible in GitHub (link to be updated).
DATA AVAILABILITY
The datasets used in this study can be accessed via zenodo (link to be updated) upon request.
SUPPLEMENTAL METHODS
Detailed description of the study cohorts
De Novo Parkinson disease study (DeNoPa): The DeNoPa cohort[27] is an ongoing, single-center study based in Kassel, Germany. The DeNoPa cohort is an observational, longitudinal study of patients with a newly established diagnosis of PD (UK Brain Bank Criteria[29]), who were naïve to L-DOPA therapy at baseline, and of age- and sex- and education-matched, neurologically healthy controls (HC). Details of inclusion/exclusion criteria have been described elsewhere.[27] Diagnostic accuracy was ensured by ongoing follow-up visits every two years (as of 2023, 10-year follow up visits were underway). Data used were received on May 16th, 2023.
Ottawa (PREDIGT) Trial: The Ottawa Trial is a pilot study to evaluate the performance of a 2-step screening tool that combines the PREDIGT questionnaire[8],[9] and the UPSIT test to distinguish patients with PD/DLB from age-matched neurologically healthy controls and patients with various other neurological diseases. Enrolment and assessment of this cross-sectional, case-control study was completed in March 2024. A manuscript that describes this cohort is in preparation. Diagnostic accuracy was ensured by independent chart review by three subspecialty-trained neurologists (JS, MGS, and AF) according to UK Brain Bank Criteria and MDS Criteria.[30]
Prognostic Biomarkers in Parkinson Disease (PROBE): PROBE[28] is a longitudinal, case-control study to test biomarkers in PD subjects and controls to determine their feasibility and potential utility as markers of risk and prognosis for PD. Details of inclusion/exclusion criteria have been described elsewhere.[28] Participants were enrolled from August, 2007 to December, 2008. Diagnosis of PD, probable MSA, and probable PSP met UK Brain Bank criteria, Consensus Criteria, and NINDS-PSP Criteria, respectively.
Supplemental methods
10-fold cross validation (CV)
For each smell test, the discovery dataset was randomly partitioned into 10 parts, where the case-control ratio was maintained in each part. In each fold, 9/10 parts were used as the development set and the remaining 1 part was used for internal validation. This procedure was repeated for 10 times, and results were showed either in distribution or average across 10 folds.
Relationship between scent identification and age
Within each cohort, participants’ ages were segregated into four bins with similar sample size. For each scent, the proportion of correctly identification was calculated for each bin, and spline smoothing was then used to represent the relationship between the proportions and age.
Item characteristic curves (ICCs)
The version of ICCs used in this study differed from traditional parametric ICCs in two aspects: 1) the x-axis was the score percentage rank in [0,100], not the latent trait on the whole real line; and 2) ICCs represented spline smoothing lines that fit response data, rather than being fitted to any pre-defined parametric model.[35]
SUPPLEMENTAL FIGURES
Supplemental Figure 1: Distribution of Sniffin’ Sticks Threshold (SST-TH) and Discrimination (SST-DS) scores for each subject group in the DeNoPa Study at baseline and their ROC curves for group classification. The Cummings estimation plots (a, b) were used to illustrate and compare smell test score distributions in each group: (a) SST-TH, (b) SST-DS. Each data point in the upper panels represents the score of one participant, and colors represent different groups and diagnosis as shown in the legend. The vertical lines in the upper panels represent the conventional mean ± standard deviation error bars. The lower panels show the mean group difference (the effect size) and its 95% confidence interval (CI) estimated by bias-corrected and accelerated bootstrap, using HC as the reference group. Panel (c) shows the ROC curves of each smell test (indicated by color) to distinguish PD/DLB versus HC (left) and PD versus OND (right). HC = healthy control. PD = Parkinson disease. DLB = dementia with Lewy bodies. MSA = multiple system atrophy. PSP = progressive supranuclear palsy. ROC = receiver operating characteristic. AUC = area under the ROC curve.
Supplemental Figure 2: Percentage differences of correct scent identification between HC and PD/DLB groups (% HC - % PD/DLB) in the DeNoPa (a) and Ottawa Trial (b) cohorts. The scents are ordered in descending orders of their mean single-scent AUC value (see Figure 3 (a) and (c)); the color of each scent changes gradually from the most discriminative to the least discriminative odorant, as indicated by the legend.
Supplemental Figure 3: Influence of distractors in the multiple-choice smell tests: the remaining 6 scents shared by UPSIT and SST-ID. Panels with odd numbers show the Item Characteristic Curves (ICCs) of six SST-ID scents, and panels with even numbers show the ICCs of the corresponding UPSIT scents. In each figure, the left panels are the ICCs using data of the PD/DLB patients, and the right panels are corresponding to healthy controls. The x-axis is transformed score indices (percentage rank of the SST-ID score) within the corresponding group. The y-axis is the probability of choosing each option at a particular score index. The correct option of each item is highlighted using the thicker blue curves. Numbers in the color legends are the option indices. The horizontal dashed lines represent 50% probability. The vertical dashed lines represent five quantiles (5%, 25%, 50%, 75%, and 95%).
Supplemental Figure 4: Item Characteristic Curves (ICCs) of the remaining SST-ID scents using baseline data from the DeNoPa Cohort. In each panel, the left panels are the ICCs using data of the PD/DLB patients, and the right panels are corresponding to healthy controls. The x-axis is transformed score indices (percentage rank of the SST-ID score) within PD and HC group, respectively. The y-axis is the probability of choosing each option at a particular score index. The correct option of each item is highlighted using the thicker blue curves. Numbers in the color legends are the option indices. The horizontal dashed lines represent 50% probability. The vertical dashed lines represent five quantiles (5%, 25%, 50%, 75%, and 95%).
Supplemental Figure 5: Item Characteristic Curves (ICCs) of the remaining UPSIT scents using the Ottawa Trial cohort. In each panel, the left panels are the ICCs using data of the PD/DLB patients, and the right panels are corresponding to healthy controls. The x-axis is transformed score indices (percentage rank of the SST-ID score) within PD and HC group, respectively. The y-axis is the probability of choosing each option at a particular score index. The correct option of each item is highlighted using the thicker blue curves. Numbers in the color legends are the option indices. The horizontal dashed lines represent 50% probability. The vertical dashed lines represent five quantiles (5%, 25%, 50%, 75%, and 95%).
Supplemental Figure 6: Range of performances for UPSIT scents in differentiating PD/DLB from MSA/PSP subjects in the PROBE cohort and AUC values for numerical subsets of odorants in group classification. Panel (a) illustrates distribution of AUC values of each scent across 10-fold cross-validation using violin plots, with 25%, 50%, and 75% quantile lines. The scents are ordered in descending order of their mean single-scent AUC value. The color of each scent changes gradually from the most to the least discriminative odorant, as indicated by the legend. Panel (b) shows the percentage of subjects correctly identifying each scent within the MSA/PSP and PD/DLB groups, and panel (c) shows the percentage differences between the two groups. Scents in panels (b) and (c) follow the same rank order as in panel (a). In panel (d), the x-axis is the number of scents included for each subset, with individual points representing average AUC values for the validation set across ten CV folds gathered in PROBE. The black horizontal, dashed line indicates the corresponding AUC values of the whole test (= 40 scents). The red horizontal, dashed line indicates AUC = 0.8 as a predetermined reference line.
SUPPLEMENTAL TABLES