Auditory perceptual history is communicated through alpha oscillations
1School of Psychology, University of Sydney, Brennan MacCallum Building A18, Manning Rd, Camperdown, NSW 2006, Australia
2School of Medical Sciences, University of Sydney, Anderson Stuart Building F13 Eastern Avenue, NSW 2006, Australia
3Department of Translational Research on New Technologies in Medicine and Surgery, University of Pisa, via San Zeno 31, 56123 Pisa, Italy
4Scientific Institute Stella Maris, Viale del Tirreno, 331, 56018 Calambrone, Pisa, Italy
5Department of Neuroscience, Psychology, Pharmacology, and Child Health, University of Florence, Via di San Salvi 12, 50139 Florence, Italy
*Corresponding authors: tam.ho@sydney.edu.au, dave@in.cnr.itAbstract
Sensory expectations from the accumulation of information over time exert strong predictive biases on forthcoming perceptual decisions. These anticipatory mechanisms help to maintain a coherent percept in a noisy environment. Here we present novel behavioural evidence that past sensory experience biases perceptual decisions rhythmically through alpha oscillations. Participants identified the ear of origin of a brief sinusoidal tone masked by dichotic white noise, and response bias oscillated over time at ∼9 Hz. Importantly, the oscillations occurred only for trials preceded by a target to the same ear and lasted for at least two trials. These findings suggest that each stimulus elicits an oscillating memory trace, specific to the ear of origin, which subsequently biases perceptual decisions. This trace is phase-reset by the noise onset of the next trial, and remains within the circuitry of the ear in which it was elicited, modulating the sensory representations in that ear.
Introduction
It has been long known that perception depends heavily on expectations and perceptual experience. Helmholtz (1867) introduced the concept of unconscious inference, suggesting that perception is at least partly inferential, or generative. Gregory (1997) described perception as a series of hypotheses to be verified against sensory data, using many compelling illusions to back his view. Both suggested that perception is a proactive, predictive process, where the brain makes best guesses about the world to test against sensory data, updating the guesses as needed.
Recent techniques lend themselves to quantitative study of predictive perception. One example is serial dependence: under many conditions, the appearance of images in a sequence depends strongly on the stimulus presented just prior to the current one. Judgements of orientation (Fischer and Whitney, 2014), numerosity (Cicchini et al., 2014), motion (Alais et al., 2017), facial identity or gender (Liberman et al., 2014; Taubert et al., 2016), beauty and even perceived body size (Alexi et al., 2018) are strongly biased towards the previous image. Strong serial biases are also observed in audition, so far only for pitch discrimination (Arzounian et al., 2017; Chambers and Pressnitzer, 2014). To our knowledge, no study has investigated whether other aspects of auditory processing, such as sound localisation, also exhibit serial dependence.
Sequential effects can last for seconds, even minutes (Chopin and Mamassian, 2012), suggesting that perception does not rely solely on the instantaneous stimulation, but on predictions conditioned by events over a long time-course. Counter-intuitively, the biases introduced by serial dependence can lead to more efficient perception (Cicchini et al., 2018): while the biases constitute perceptual errors, these errors are offset by a reduction in response variance and are therefore beneficial overall. As the world tends to remain constant over the short term (Dong and Atick, 1995; Voss, 1975), past events are good predictors of the future. On the assumption that nothing has physically changed, the previous stimulus acts as a predictor, or a prior, to be combined with the current sensory data to enhance the signal. And on average, that predictor is useful in reducing error.
However, not all serial effects are positive. Classical adaptation aftereffects are negative (Thompson and Burr, 2009): prolonged inspection of downward motion causes stationary stimuli displayed to the same location to appear to move upwards (Addams, 1834). Similar negative aftereffects have been observed in almost every dimension, including orientation (Gibson and Radner, 1937), colour (Neitz et al., 2002), faces (Leopold et al., 2001), and numerosity (Burr and Ross, 2008). They also exist in audition and cover a similarly wide range of aspects, from auditory motion (Grantham and Wightman, 1979), sound localisation (Kashino and Nishida, 1998) and tone intensity (Reinhardt-Rutland and Anstis, 1982) to gender of human voices (Schweinberger et al., 2008) and even vocal affect (Bestelmeyer et al., 2010). While the classical paradigm usually involves long periods of adaptation, sizeable negative aftereffects can arise after very brief presentations in vision as in audition (Aagten-Murphy and Burr, 2016; Alais et al., 2017, 2015).
What determines whether assimilation or contrastive effects prevail, and how do the two opposing mechanisms interact? One factor is stimulus strength: strong, salient, high-contrast, long-duration stimuli tend to lead to negative aftereffects, while brief, less salient low-contrast stimuli result in positive aftereffects (Kanai and Verstraten, 2005; Pantle et al., 2000; Yoshimoto et al., 2014). Taubert et al. (2016) reported strong positive serial dependence for judgments of the gender of faces (a stable attribute) and negative aftereffects for judgments of emotion (a labile attribute, where change is important) on the same stimuli. However, negative and positive effects may also reflect individual preferences. Abrahamyan et al. (2016) found that some individuals, termed “switchers”, tend to respond opposite to their previous choice, while others prefer to “remain” with whatever response they gave in the preceding trial. Finally, Chopin and Mamassian (2012) made a surprising observation within the same experiment: stimuli close in time were negatively correlated with the current percept, while those more remote positively. Importantly, their modelling results suggest that both negative and positive dependencies could be explained by a single predictive adaptation mechanism. However, how such a predictive adaptation mechanism could be implemented neurally is unclear.
Generally, the neuronal mechanisms underlying serial dependence are largely unknown. It is assumed that the prior is generated at mid-high levels of analysis, and fed back to early sensory areas, which in turn modify the prior (Friston, 2005; Lee and Mumford, 2003; Summerfield and de Lange, 2014; Summerfield and Koechlin, 2008; Yuille and Kersten, 2006), but we do not know how this information is propagated, nor do we have a grasp on the underlying neural mechanisms. One possibility is that recursive propagation and updating of the prior is related to low-frequency neural oscillations (Friston et al., 2015; Sherman et al., 2016; VanRullen, 2017).
Although controversial, evidence is gathering that synchronous rhythmic neural activity may serve to bind stimulus components into a unitary percept (Buzsáki and Draguhn, 2004; Engel et al., 2001; Gray et al., 1989). Besides the neural evidence in animals and humans, recent psychophysical evidence shows that neural oscillations in mammalian brain modulate perceptual performance, and that these oscillations can be synchronized either by abrupt perceptual stimuli (Fiebelkorn et al., 2013, 2011; Landau and Fries, 2012), or by motor-action (hand or eye movements) (Benedetto et al., 2018, 2016; Benedetto and Morrone, 2017; Tomassini et al., 2017, 2015). Using signal detection theory (Green and Swets, 1966; Macmillan and Creelman, 2004), we have recently shown within the same experimental setup that oscillations can occur in both sensitivity (accuracy) and criterion (response bias) at different frequencies (theta for sensitivity, alpha for criterion), suggesting separate mechanisms underlying those two perceptual properties (Ho et al., 2017). Oscillations in bias could plausibly reflect the reverberation of recursive error propagation within a generative perception framework.
Alpha oscillations in criteria are consistent with EEG evidence of an association between criteria shifts and modulations of alpha power (Craddock et al., 2017; Haegens et al., 2014; Iemi et al., 2017; Iemi and Busch, 2018; Limbach and Corballis, 2016) and phase (Sherman et al., 2016). Decreases in alpha power consistently coincide with more liberal decision criteria in vision (Iemi et al., 2017; Iemi and Busch, 2018; Limbach and Corballis, 2016) as well as touch (Craddock et al., 2017; Haegens et al., 2014). Notably, the alpha modulation can be observed several hundred milliseconds before stimulus onset (Craddock et al., 2017; Limbach and Corballis, 2016), suggesting that criteria shifts are driven, at least in part, by alpha oscillations. This notion is further supported by the observation that the phase of prestimulus alpha can predict criteria changes (Sherman et al., 2016). Furthermore, the relationship between prestimulus alpha and criteria change was strongly influenced by target expectancy, reinforcing the view that alpha oscillations are involved in sensory anticipation or prediction (de Lange et al., 2013; Klimesch et al., 1998; Mayer et al., 2016; Premereur et al., 2012; Rohenkohl and Nobre, 2011; Sanders et al., 2014).
VanRullen and Macdonald (2012) have proposed an oscillatory mechanism by which past perceptual history is stored in memory. By cross-correlating the EEG response with the corresponding visual stimulus, they showed that random, non-periodic luminance changes elicited a “perceptual echo”, a reverberatory response at ∼10 Hz that lasted for at least 1 s. Recent findings have confirmed that this long-lasting oscillation serves to maintain sensory representations over time (Chang et al., 2017; Huang et al., 2018), which could influence subsequent perception.
Given these findings, we hypothesised that the alpha oscillations observed in criteria (Ho et al., 2017) may underlie a predictive mechanism that biases perceptual decisions, giving rise to sequential effects. To test this idea, we asked participants to indicate the ear of origin of a brief monaural sinusoidal tone masked by uncorrelated streams of dichotic white noise in each ear (illustrated in Fig. 1A). The tone was presented at random intervals after the noise onset which served to reset the phase of ongoing neural oscillations. From our previous work (Ho et al., 2017), we expected decision criterion to oscillate at alpha frequencies (8–12 Hz). If these oscillations are instrumental in communicating perceptual expectations, the rhythmic fluctuations in criteria should be contingent on the previous stimuli.
Results
Stimulus history affects response bias but not sensitivity
Participants performed a two-alternative forced-choice (2AFC) task which required them to identify the ear of origin of a weak, brief tone (10 ms) embedded in white noise (2 s in duration). The target tone was delivered randomly with equal probability in either ear with a stimulus-onset asynchrony (SOA) randomly drawn from an interval of 0.2–1.2 s post noise onset. Using a staircase procedure, we kept the target’s intensity around individuals’ thresholds (75 % accuracy).
We first examined whether observer sensitivity and response bias, as measured respectively by d-prime (d’) and decision criterion (c) (Eq. 1&2, see also Fig. 1B), were contingent on the preceding stimuli. As shown in Fig. 2A, stimulus history did not influence observer sensitivity. Whether the last target (1-back) and second-last target (2-back) occurred in the left or right ear had no effect on detecting the current target: the difference in sensitivity contingent on the preceding left (blue bars) and right stimuli (red bars) was not statistically significant after multiple comparison correction (with FDR = 5 %), p = 0.07 and p = 0.56 for 1-and 2-back, respectively. However, Fig. 2B shows that criterion was affected by stimulus history. Curiously, there was no significant 1-back effect (when FDR adjusted), p = 0.86, but a strong 2-back effect (p = 0.0002) of observer response biases towards the previously presented stimulus, either left or right. The effects remained significant for stimuli three trials back, p = 0.008, and four trials back, p = 0.009, but was non-significant five trials back, p = 0.85.
We then examined the dependency of the previous response (Fig. 2C), which revealed a negative influence of one trial back, marginally significant after FDR correction (p = 0.053). This response-dependent negative serial dependence could result from response switching, mentioned in the Introduction. However, two trials back, which should not be affected by response switching, showed the same positive aftereffect found for the stimulus-based analysis, p = 0.0001 (FDR corrected). Trials further in the past (not shown) had no significant effect, all p’s > 0.05. Even though not statistically robust, this repulsive effect could account for the stimulus-based null effect in the 1-back (Fig. 2B), discussed further in the Discussion section.
To gain a better understanding of the sequential effects in response bias, we examined the effects in individual subjects. Figure 2D plots the individual biases (differences in responses preceded by left from those preceded by right) for 1-back trials against 2-back trials. All subjects except one (subject 5) showed a positive serial effect for trials that were 2-back, while there was much greater subject variability in the 1-back condition, with 5 participants showing a positive and the remaining 9 showing a negative effect. This pattern of results is consistent with previous reports of individual differences in switching responses from one trial to the next (Abrahamyan et al., 2016).
That both negative and positive effects are observed only for criteria is particularly interesting, as we have previously shown that response biases oscillate at low alpha frequencies, ∼8 Hz, and there is evidence that alpha oscillations are involved in predictive processes and modulation of decision criteria (de Lange et al., 2013; Sherman et al., 2016). In a preliminary analysis of the temporal dynamics of the 1-back effect, we sorted the individual response data by target SOA (from noise onset) and grouped them into 12 bins of 83 ms, computed the difference in left and right bias contingent on the preceding trial and averaged across all participants. As can be seen from Fig. 2E, the effect of stimulus history on response bias oscillates smoothly over time, warranting further investigation.
Oscillation of response bias but not sensitivity
We applied two different methods to examine response bias and sensitivity for rhythmic fluctuations. The first is the same curve fitting approach we used in our previous study (Ho et al., 2017). For this analysis, we pooled the individual trials across all participants, sorted the trials by target SOA and grouped them into hundred 10-ms bins, from 0.2 to 1.2 s post noise onset. For each bin, we computed d’ (Eq. 1) and c (Eq. 2) and fitted sine curves of different frequencies to the temporal sequences of d’ and c.
We first considered the data as a whole (without dividing the trials contingent on the preceding stimuli) and searched for significant modulations within the range of 4-12 Hz in 0.1-Hz steps. This frequency range encompasses the theta and alpha band. Figures 3 A&B show the binned aggregate data as a function of target SOA from noise onset, expressed as both sensitivity (Fig. 3A) and criterion (Fig. 3B). The shaded yellow and green areas enveloping the lines represent ±1 standard error (N = 14) of the group mean (SEM), computed by bootstrapping the individual trials 2,000 times and applying the same curve fitting analysis to the surrogate data. It is evident that while sensitivity shows no clear oscillation, the criterion does oscillate around the frequency of 9.4 Hz, as indicated by the grey curve in Fig. 3B. For comparison, we fitted a sinusoidal with the same frequency, 9.4 Hz (grey curve in Fig. 3A) to the sensitivity data.
We evaluated the goodness of each harmonic fit (4-12 Hz) using R2 combined with a permutation procedure, which involved shuffling the individual responses 2,000 times and applying the same curve fitting analysis to the surrogate data. From the surrogate data, we constructed a distribution of maximal R2 (across all tested frequencies) for sensitivity (Fig. 3D) and criterion (Fig. 3E), against which we compared the R2 of the original data. The results are summarised in Fig. 3C. For sensitivity (orange line), no frequency produced a good fit, none near the corrected 95% confidence threshold of R2 = 0.12 (dotted black line). However, the criterion data (green line) showed significant modulations between 9.2 and 9.6 Hz, with a strong peak at 9.4 Hz (R2 = 0.15). The significance test (illustrated in Fig. 3D&E) confirmed that this frequency, 9.4 Hz, was significant for criterion, p = 0.009 (Fig. 3D), but not sensitivity, p = 0.9 (Fig. 3D). The phase of the 9.4 Hz oscillation at trial onset relative to the noise burst onset was 179° ± 16° SD (by bootstrap).
The results shown in Fig. 3 were based on an aggregate data analysis (i.e., by pooling all trials across subjects). In principle, the criterion oscillation in Fig. 3B could derive from a few subjects with very strong oscillatory effects. To verify that this is not the case, we examined the individual data and evaluated their coherence as a group using the linear regression analysis in Eq. 5 (illustrated in Fig. 1C). As detailed in the Methods, this approach does not require data binning. We applied the regression analysis to the individual accuracy (correct or incorrect) and response data (‘left’ or ‘right’).
The results, summarised in Fig. 4, corroborate those of the aggregate data analysis. Figure 4A plots the amplitude spectrum for accuracy (an approximation of sensitivity) and Fig. 4B that for response bias (an approximation of criterion), which were computed using the individual estimates of the fixed-effect regression parameters, β1 and β2, and their vectorial mean (Eq. 6). The shaded yellow and green areas enveloping the lines represent ±1 SEM. Response bias (Fig. 4B) shows a strong peak around the same frequency, 9.4 Hz, as the R2 results for criterion in Fig. 3C. Using a similar permutation procedure as in the aggregate analysis, we evaluated the significance of each frequency between 4-12 Hz in 0.1-Hz steps. The corrected p-values are plotted in Fig. 4D&E for accuracy and response bias, respectively. The results for response bias shows that the oscillation at 9.4 Hz is significant (p = 0.027) after correction for multiple comparisons with maximal statistics (illustrated in Fig. 4F). The amplitude at 9.4 Hz is A = 0.02 ± 0.01 SEM. While there are several peaks in the amplitude spectrum for accuracy (Fig. 4A), none was significant after multiple comparison correction (Fig. 4D).
Figure 4F illustrates how the p-values were computed and corrected. The green dots (2,000 in total) clustering in a ring around the origin (0, 0) represent the joined distribution of maximal B1 and B2 (group averages of the individual β1 and β2) obtained by permutation. To construct this distribution, we determined the maximal vector of each randomisation, irrespective of the frequency. The red dot shows the group mean estimated from the original data at the peak frequency, 9.4 Hz. The red circle has a radius equal to the distance of the red dot (B1, B2) to the origin (0, 0). The p-value reflects the proportion of permuted data exceeding the original data (the green points that fall outside the circle).
Finally, Fig. 4C shows the individual phase and amplitude vectors for response bias at 9.4 Hz. The length of the vectors reflects the amplitude of the 9.4 Hz oscillation and the vector direction indicate individual phases at noise onset. The vectors are tightly clustered around a phase angle, θ, of 172° ± 16° SEM (Eq. 7). This is consistent with the phase angle we obtained from the curve fitting analysis with the aggregate data (see above).
Oscillation in response bias is driven by stimulus history
Having established the existence of rhythmic fluctuations in bias and criterion in both the aggregate and individual data, we investigated how the oscillations relate to serial effects driven by expectations from stimuli of previous trials using two different analysis. In the first analysis, we separated the trials based on whether the previous stimulus had been presented to the same ear (congruent) or a different ear (incongruent), and tested both sets of data for oscillations in both sensitivity and bias, using the same curve fitting and linear regression analyses as above. As predicted, the oscillation observed in criterion (Figs. 3&4) was driven by the preceding stimulus history, occurring only when the previous stimulus was congruent with the current stimulus.
Figure 5 shows the results of the analysis of the aggregate data. As before, we fitted the temporal sequences of sensitivity (d’) and criterion (c) for congruent (Fig. 5A) and incongruent trials (Fig. 5C) with sinusoids ranging in frequency between 4-12 Hz in 0.1 Hz steps. Congruent trials (dark green line), but not incongruent trials (light green line), showed a good fit at 9.4 Hz (thick grey line). The goodness of fit at all tested frequencies are plotted in Fig. 5D. The largest R2 was obtained around 9.4 Hz for congruent trials, with R2 = 0.15, while for incongruent trials, the goodness of fit did not approach significance at any frequency. As in the previous analysis, we created a distribution of maximal R2 (Fig. 5D) from the 2,000 surrogate datasets obtained by permutation and determined the 95 percentile for both congruent (Fig. 5D) and incongruent (Fig. 5E) trials. Any frequency with R2 exceeding this threshold (R2 = 0.12, dotted black line in Fig. 5B), was considered significant. Only the peak (green line, congruent trials) around 9.4 Hz survived the multiple comparison correction, with p = 0.01. The phase of this 9.4-Hz oscillation for congruent trials was, at noise onset, 180° ± 17° SD (by bootstrap). For incongruent trials, the phase at 9.4 Hz was 179° ± 34° SD.
As before, we examined the individual subject data using the linear model in Eq. 5, separately for congruent and incongruent trials. The amplitude spectra in Fig. 6A and 6D are based on the individual estimates of β1 and β2 averaged across the 14 participants (see Eq. 7). As for the aggregate data, the congruent trials (dark green line) yielded a large peak around 9.4 Hz with A = 0.03 ± 0.009 SEM (Fig. 6A). At this frequency, the amplitude is reduced for the incongruent amplitude spectrum with A = 0.02 ± 0.01 SEM (Fig. 6D). Inspection of the individual vectors shows a tight cluster around a mean phase angle (at noise onset) of 164° ± 13° SEM for congruent trials (Fig. 6B), while incongruent trials had a slightly greater phase dispersion, with a mean phase angle (at noise onset) of 177° ± 18° SEM (Fig. 6E).
The results of the 2D permutation test plotted in Fig. 6G corroborate the observations in Figs. 5A&D. The only frequencies to survive the strict multiple comparison correction were around 9.4 Hz (dark green line, congruent trials) with p = 0.045. In contrast, incongruent trials showed no significant frequencies (light green line). The 2D permutation distributions for congruent and incongruent trials at 9.4 Hz are plotted in Fig. 6C and 6F, respectively. The thick red dot shows the amplitude and phase of the group mean with respect to the permutation distribution. Only in the congruent case does the group mean exceed the distribution of maximal vectors significantly (Fig. 6C).
We also analysed the sensitivity for congruent and incongruent trials, but failed to find significant oscillations in either data set (data not shown). The lack of modulation in sensitivity in the present study is consistent with our previous findings that sensitivity oscillates in antiphase between the left and right ears (Ho et al., 2017): because they are out of phase, they should cancel each other under the conditions of this study, where results for left and right ears need to be combined for the sensitivity analysis. We checked this hypothesis by a post-hoc regression analysis of the accuracy of congruent trials separated by ear of origin. We first examined the phase coherence across subjects at all frequencies of interest (4-12 Hz) for each ear and congruence condition separately. Observers showed significant phase consistencies for both ears for the congruent trials at only one region in the frequency spectrum, from 9.1 to 9.7 Hz (consistent with the grey shaded region of significance in Fig. 6B). At no frequency was there strong phase coherence in the incongruent trials. Using the Watson-Williams test (circular analogue to a two-sample t-test; Berens, 2009), we confirmed that the group phase distributions for left- and right-ear accuracy in the congruent trials were significantly different (p < 0.05, Bonferroni corrected) between 9.1 and 9.7 Hz, peaking at around 9.4 Hz. Figure 6H shows the individual vectors at 9.4 Hz for congruent trials containing a left target with a mean direction (thick red line) of ∼189° ± 12° SEM. For congruent trials containing a right target, the individual vectors show a mean direction of ∼304° ± 14° SEM at 9.4 Hz (Fig. 6I). Although the relationship between the left and right accuracy on congruent trials is not exactly antiphase (∼115°), they do show a trend in this direction. This suggests that the absence of sensitivity oscillations (Figs. 2&3) in the present study is likely due to a cancellation of left and right ear oscillations.
As the serial dependence analysis showed strong 2-back effects (Fig. 2B), we hypothesised that the 9-Hz oscillation could be long-lasting, spanning at least two trials, which is also suggested by recent findings (Chang et al., 2017). To investigate whether events two trials back show measurable oscillations near 9.4 Hz, we separated the 1-back incongruent data (which show no significant oscillations in Fig. 5C) further on the basis of whether the stimulus two trials back was congruent or incongruent. Given that the 1-back incongruent data showed no oscillation in previous analyses, any memory trace should be relatively weak and confined to a very narrow frequency window 0.5 Hz wide, 9.1-9.6 Hz, the frequency range where the congruent 1-back data showed a significant oscillation (after correction for multiple comparison; grey shaded box in Fig. 6B). We used the same analysis as before, but instead of testing every frequency from 4-12 Hz, we allowed the curve-fitting algorithm to search for the best frequency within our limited window of interest (9.1-9.6 Hz) for both the original and shuffled data. The results, shown in Fig. 7 confirm a 9-Hz oscillation in the 2-back congruent data which was best fitted by a sinusoidal with frequency of ∼9.2 Hz, R2 = 0.08. The permutation test confirmed that this fit is significant (Fig. 7B), p = 0.037 when compared with the best fits of shuffled data within the frequency of interest, 9.1-9.6 Hz. There was no modulation near this frequency in the 2-back incongruent data (Figs 2E&F): R2= 0.02, with p = 0.7 (Fig. 7D). No other frequency within the range of 4-12 Hz approached significance, after multiple comparison correction.
To examine the individual data and group coherence at 9.2 Hz, we applied the same regression analysis as before (Eq. 5). Figure 7C shows the individual vectors at 9.2 Hz based on the 2-back congruent trials and Fig. 7G the individual vectors based on the incongruent trials. The individual phases in the congruent condition cluster around a similar phase, 190° ± 18°, similar to that of the 1-back congruent data (Fig. 6B). The mean phase angle in the incongruent condition is 295° ± 15° and bears no relation to either the mean phase in the 1-back congruent or incongruent condition (Figs. 6B&E). Figures 7D&H show the results of the 2D permutation test at 9.2 Hz, which are consistent with the results of the aggregate data analysis. Very few points (dark cyan dots in Fig. 7D) from the permutation distribution (p = 0.013) exceed the group mean vector (thick red dot) in the congruent condition, while many more points exceed the mean vector in the incongruent condition, p = 0.36 (Fig. 7H). Taken together, the results suggest that the 9-Hz oscillation lasts at least two trials. Although the reverberation is weaker 2-back than 1-back, as to be expected, it may be sufficient to induce the long-lasting serial dependence observed on average (Fig. 2).
Discussion
To maintain a stable and coherent percept in a world that is naturally noisy and ambiguous, observers take advantage of past information to anticipate forthcoming sensory input. While there is a good deal of behavioural evidence in favour of this predictive account of perception, little is known about the underlying neural mechanisms. Here, we present evidence suggesting that predictive perception is implemented rhythmically through alpha-band oscillations, along the lines of the “perceptual echo” suggested by VanRullen and Macdonald (2012).
Our current study shows that identifying the ear of origin of a weak tone was strongly biased by previous stimuli, two or even three trials before the current trial. Although the immediately previous trial had no average serial effect, we observed a strong rhythmic fluctuation in bias (measured by criterion) at ∼9.4 Hz, which was critically dependent on stimulus history: strong oscillations occurred only when the previous target had been presented to the same ear as the current one, and weaker oscillations could be elicited by stimuli in the same ear two trials previously. The results reinforce our previous study (Ho et al., 2017) showing oscillations in criterion (for a different task) at similar frequencies (around 7.5 and 8.7 Hz), and extend these in an important way by showing that the oscillations are contingent on past stimulus history.
Why do oscillations occur only when the previous trial was presented to the same ear as the current one? One possible mechanism, illustrated in Fig. 8, is based on the “perceptual echo”, as suggested by VanRullen and Macdonald (2012; see also Bowen, 1989). An auditory signal presented to one ear should elicit a neural reverberatory response that oscillates in the alpha range, and affects subsequent perceptual decisions. Our data further suggest that the perceptual echo is confined to the ear in which it is evoked. We assume that the phase of this oscillation becomes aligned to the onset of the noise-burst of the following trial (dotted vertical black lines in Fig. 8), and that the alignment phase is opposite for each ear (180° for the left and 0° for the right ear). This echo could then bias the response to subsequent signals presented to the same ear, either by modulating the neural gain of the signals (Rahnev et al., 2011), or by causing a shift in the decision boundary.
To illustrate, Fig. 8 shows putative internal representations of noise and signal. Presentation of a stimulus may set up an internal “echo” at about 9 Hz, confined to that ear. The onset of the noise burst resets the phase of the echo, to about 180° in the left ear (Fig. 8A) and 0° in the right ear (Fig. 8B). If we assume that the echo modulates gain, then the amplitude of the response to stimuli presented to the same ear will depend on time of presentation, more amplified if presented at a peak than at a trough. This modulation will be reflected in response bias, but not average sensitivity, as the echoes in the two ears are assumed to be out of phase (discussed below). The new stimulus will in turn elicit a new echo in that ear (red trace), which reverberates to the next trial.
How could cyclic increases in sensitivity or gain result in changes in criterion but not in d’? As the sequence of stimuli was completely random, past information was not informative about the current trial, and therefore could not increase d’, either on average or in a cyclical manner; but it can change the trial-by-trial judgment of which ear carried the weak signal. It is important to point out that criterion changes do not necessarily reflect decision processes (such as shifts in decision boundaries), but can also reflect perceptual changes (Peters et al., 2016; Witt et al., 2015). In the current experimental design, calculations of d’ can only be made after combination of hits and false alarms to stimuli presented to the two ears. If the modulation of activity in the two ears is in counterphase, they will cancel each other out when the output is combined for d’, but add together for criterion. Evidence from our previous (Ho et al., 2017) and present study (of Figs. 6H&I) suggests that this does occur: when the congruent data are separated for ear of origin of the signal, the oscillations in the left ear tend to be out of phase with those of the right (∼115° on average in this study).
Why did our experiments reveal no positive serial dependence on the immediately previous trial (1-back), as is normally observed in studies of serial dependence? There are (at least) two possible, non-mutually exclusive explanations. The first is a consequence of the forced-choice paradigm used here (while most studies of serial dependence use a reproduction technique). Forced-choice paradigms can lead to sequential response biases independent of the stimuli, such as alternation of responses (Abrahamyan et al., 2016), especially after a long run of similar responses (probably related to the gambler’s fallacy (Burns and Corpus, 2004; Tversky and Kahneman, 1971). As the responses were strongly correlated with stimuli (75% correct), a systematic alternation would tend to cancel out positive serial dependence based on the previous stimulus (but have no effect on average for trials two-back). Not all observers alternate: some show the opposite tendency, “sticking” with the current response. Figure 2C is consistent with this idea: while only one observer showed negative serial dependence for 2-back trials, the variability in 1-back trials was much higher: five out of 14 observers showed positive 1-back biases, showing a tendency for “sticking”, while the others could be switchers. Further evidence comes from the response-based analysis, which showed a strong negative serial dependency, meaning that responses tended to be different from the previous one (independent of the stimuli). As stimuli and responses were highly correlated, this response-dependent bias will also impact the dependence on previous stimuli, potentially cancelling any positive serial dependence that may have been present.
Another possible reason for the lack of positive 1-back effects is that the stimuli may have caused both positive serial dependence and negative adaptation aftereffects. As mentioned in the introduction, both types of effects have been reported in sequential judgments, sometimes within the same experiment (Abrahamyan et al., 2016; Chopin and Mamassian, 2012; Taubert et al., 2016). Negative aftereffects tend to be shorter lived than the positive dependencies, and should therefore affect only 1-back, not 2-back trials (Chopin and Mamassian, 2012). This seems to be less plausible given the strong dependence on response (while aftereffects are stimulus-driven), but we cannot rule out the possibility.
Whatever the reason for the lack of positive serial dependence in the averaged results, our study suggests that oscillations may be a more sensitive signature of memory-based perceptual effects than simply looking at average results. Many competing effects could reduce or annul average serial dependence effects, without affecting rhythmic, time-dependent oscillations.
As both vision and audition have similar serial dependencies and oscillations of decision criteria at alpha rhythm, the mechanism described in Fig. 8 may generalise to other sensory modalities. However, considering the anatomical and organisational differences between the visual and auditory system, it is possible that this mechanism had to be adapted to the specific architecture of each sense. For instance, oscillatory effects that are robust in vision are not readily observed in audition (VanRullen et al., 2014; Zoefel et al., 2015; Zoefel and Heil, 2013; Zoefel and VanRullen, 2017). However, by using dichotic rather than diotic stimulation (as all previous studies had done) we were able to show that similar to visual sensitivity, auditory sensitivity also oscillates but in antiphase for the two ears (Ho et al., 2017). As a result, when we summed the ears (as would be the case in diotic stimulations), the sensitivity oscillations cancelled each other out. Our findings therefore show that the ears need to be stimulated separately in order to observe certain oscillatory effects.
In summary, we showed that when participants were asked to determine the ear of origin of a brief sinusoidal target masked by white noise, their decision criteria oscillated at an alpha rhythm. Further analyses revealed that the alpha oscillation in criteria was driven entirely by trials in which the target occurred in the same ear as that of the immediately preceding trial. When the previous and current targets occurred in opposite ears, we found no observable oscillations of criteria. To account for these findings, we proposed that every target elicits a long-lasting reverberation at alpha rhythm that is ear-specific and continues to the next trial, where it synchronises with other ongoing brain oscillations possibly related to both perceptual and decisional processes. Because this ‘echo’ is ear specific, it will only bias perceptual decisions (towards or away from the ear in which the echo was elicited), if the following target occurs in the same ear. The bias is not absolute but fluctuates over time, and may become stronger as more evidence of the same stimuli occurs successively in the sequence. Although this model is very basic, we believe it could be elaborated to explain more complex serial dependence effects and provide a deeper understanding of how expectation or prediction guides perception.
Materials and Methods
Participants
Eighteen healthy participants took part in the experiment. All reported normal hearing. Four participants were male and two left-handed. After inspecting participants’ auditory thresholds and (log-transformed) reaction times, we excluded four participants (one male) from the data analysis for the following reasons. Three exhibited large differences in mean auditory threshold between the left and right ear (1 standard deviation from the group mean) and one displayed an atypical distribution of very long reaction times (2.5 standard deviations from the group mean). The mean age of the remaining 14 participants was 21.14 ± 4.22. All participants provided written, informed consent. The study was approved by the Human Research Ethics Committees of the University of Sydney. No power analysis was used to decide the number of subjects or the number of trials per subject. We based our sample size estimations on our previous study (Ho et al., 2017), which showed oscillations in auditory perceptual performance before, and other studies on similar behavioural rhythms in vision (Fiebelkorn et al., 2013; Landau and Fries, 2012).
Experimental procedure
The experimental design was kept largely the same to our earlier study (Ho et al., 2017). However, instead of the bilateral tone discrimination task (2×2 design) used in our previous experiment, we asked participants to determine the ear of origin of a similarly brief sinusoidal tone of 10-ms duration and 1,000 Hz in frequency masked by dichotic white noise.
As in the previous experiment, participants sat in a dark room and listened to auditory stimuli via in-ear tube-phones (ER-2, Etymotic Research, Elk Grove, Illinois). They wore earmuffs on top of the tube-phones (3M Peltor 30 dBA) to isolate external noise. The broadband white-noise masks were 2 s long and randomly generated each trial. To ensure that the noise masks were both clearly lateralised and uncorrelated, the left- and right-ear maskers were in antiphase (each a time-reversed duplicate of the other). As illustrated in Fig. 1A, the noise segments were presented to both ears with simultaneous onset. The target tone was delivered randomly with equal probability in either ear during the 2-s noise masker with a stimulus-onset asynchrony (SOA) randomly drawn from an interval of 0.2–1.2 s post noise onset. Participants responded via button press on a response box (ResponsePixx, Vpixx Technologies, Saint-Bruno, Quebec). All participants used their thumbs to respond were instructed to respond as soon as possible while the noise masker was still present. For each ear, the target’s intensity was kept around individuals’ thresholds (75 % accuracy) using an accelerated stochastic approximation staircase procedure (Faes et al., 2007; García-Pérez, 2011). The next trial started after a silent inter-trial interval (ITI) of random duration between 1.2–2.2 s.
Participants completed 2,800 trials (40 blocks of 70 trials) in total. Blocks lasted ∼5 minutes and participants were allowed to rest after every block. Prior to the experiment, they completed a practice block of 20 trials with feedback but no feedback was provided during the experiment. Stimuli were presented using the software PsychToolbox (Brainard, 1997) in conjunction with DataPixx (Vpixx Technologies, Saint-Bruno, Quebec) in MATLAB (Mathworks, Natick, Massachusetts).
Signal detection theory
Prior to the data analysis, trials in which the response occurred before the target onset or after the noise offset were excluded. Additionally, we eliminated trials with reaction time (RT) exceeding the 99 % confidence intervals of individuals’ RT and with target intensity beyond the 95 % confidence intervals of individuals’ thresholds.
We computed d’ and c using Eq. 1 and 2 from signal detection theory (SDT) (Green and Swets, 1966; Macmillan and Creelman, 2004). As illustrated in Fig. 1B, the calculation of the hit rate (Hright) was based on the hits from the right target condition and the false alarm rate (FAleft) based on the false alarms from the left target condition. d’ is given by the difference between z-transformed hit and false alarm rates:
Half the sum of the z-transformed hit and false alarm rates gives c, the bias towards one of the responses (left ear when positive and right ear when negative).
Behavioural oscillations
Aggregate data analysis
To look for oscillations in sensitivity and response bias, we applied two different methods. The first is the same curve fitting approach we used in our previous study (Ho et al., 2017). For this analysis, we pooled across all 14 participants, sorted the trials by target SOA and grouped the data into hundred 10-ms bins, from 0.2 to 1.2 s post noise onset. The mean number of trials per bin was 151 ± 24 for the left-target condition and 152 ± 25 for the right-target condition. For each bin, we computed d’ and c as above (Eq. 1&2). Using the standard MATLAB function fit included in the Curve Fitting toolbox, we fitted the following Fourier series model (Eq. 3) to the resulting sequence of sensitivity and criterion (see Fig. 3A&B): where t is time (t = 0.2, 0.3… 1.2 in seconds), ω is the angular frequency (ω = 2πf) we want to test, a0 is a constant and a1 and b1 are the sine and cosine coefficients respectively. A and Φ represent amplitude and phase of the sinusoidal fit. We used a non-linear least-squares method (a standard implementation in MATLAB) whereby the summed squares of the residuals were minimised through successive iterations (400 iterations in total). As in our previous study (Ho et al., 2017), sensitivity displayed a decreasing non-linear trend over time which we removed before curve fitting using the same polynomial fit (polyfit in MATLAB). Detrending was not applied to the criterion curve. For each frequency between 4 and 12 Hz (in 0.1 Hz steps), we fitted the best sinusoid, allowing amplitude and phase to vary (two degrees of freedom). This yielded a measure of goodness of fit, R2, as a function of frequency (Fig. 3C). To test the significance of every fit, we applied a permutation procedure (Ernst, 2004) whereby we shuffled the responses of each individual trial over all SOAs to generate 2,000 surrogate datasets which we submitted to the same binning and curve fitting procedure as the original data. To correct for multiple comparisons, we determined the maximal R2 for every surrogate dataset irrespective of frequency. This resulted in a distribution of 2,000 maximal R2 (Figs. 3D&E) against which we compared each fit to the original dataset. Any frequency that exceeded the 95 percentile of the maximal-R2 distribution (dotted line in Fig. 3C) was considered significant. We also estimated the variability in the aggregate data by applying the bootstrap method which involves the random selection of the same number of trials (with replacement) as in the original data and submitting the surrogate datasets to the same binning and curve fitting procedure as above.
Individual and group analysis
The curve fitting method described above requires a sufficient number of trials per time point to accurately estimate the oscillations underlying d’ and c. For that reason, we had to pool across all subjects and bin the aggregated data. In order to examine the individual data for oscillations and evaluate their coherence across subjects, we used a different approach which does not require data binning but allows for an estimation of participants’ sensitivity and response bias based on single trials (for similar approaches, see (Benedetto et al., 2018; Tomassini et al., 2017). As illustrated in Fig. 1C, the response yi (i = 1, 2… n, where n is the total number of trials) to a target presented at time ti (i.e., the interval from noise onset to target onset in seconds) can be modelled as the linear combination of harmonics at each tested (angular) frequency (compare with Eq. 3):
where is the predicted responses and β0, β1 and β2 are fixed-effect regression parameters that can be estimated using the linear least-squares method implemented in MATLAB as the fitlm function from the Statistics and Machine Learning toolbox. To retain some consistency between the individual data and aggregate data analyses, we analysed sensitivity and response bias separately, that is, yi reflected either individual accuracy or response bias. Therefore, the model we used is a special case of the general linear model (GLM) as it is restricted to one dependent variable (also called multiple linear regression). This model estimates the regression parameters adequately when the sampling rate is uniform across the time series. As this condition may not always be met, we applied the full model which includes a third independent regressor containing information about the stimulus:
where S is the stimulus at time t and takes the value –1 or +1 for left and right target respectively. We examined sensitivity based on participants’ correct and incorrect responses in which case yi = 1 for correct and yi = 0 for incorrect. To analyse response bias, yi took the value 1 when participants made a ‘right’ response and 0 when they made a ‘left’ response.1 To measure the group’s coherence in terms of both phase and amplitude at each frequency, we averaged the sine and cosine regression parameters across all participants, i.e.,
and
, where n = 15 and i = 1, 2… n, and obtained an amplitude spectrum (Figs. 4A&B) by taking the vectorial average of the individual estimates given by the square root of their sum squared for every frequency:
We computed the mean phase θ by taking the arctangent of the ratio of the averaged sine to cosine regression parameter:
The significance of the model fit in Eq. 5 was evaluated using a two-dimensional (2D) permutation test. Specifically, we tested the null hypothesis that the individual response data contain no temporal structure and thus their time stamps should be interchangeable. To this end, we shuffled the SOAs (i.e., the interval between noise onset and target presentation), keeping the relationship between the response (‘left’ response or ‘right’ response) and stimulus (left target or right target) constant (Kennedy and in Statistics-Simulation, 1996; Winkler et al., 2014). The permutation was carried out at the individual subject level 2,000 times per dataset. Each surrogate dataset was fitted using the same model described in Eq. 5 and the resulting β1 and β2 were averaged across subjects for every frequency tested. This yielded a joined distribution of 2,000 surrogate means for each frequency from 4-12 Hz in 0.1-Hz steps. Similar to the multiple comparison correction we applied in the aggregate data analysis, we determined the maximal vector of each joined distribution, irrespective of the frequency. This resulted in a joined distribution of 2,000 maximal vectors, against which we compared the original mean (see Fig. 4F). This comparison entails determining the proportion of surrogate amplitude group means equal or exceeding the original group mean.