Abstract
Auditory perception is the set of processes by which the brain turns the pressure wave arriving at the two ears into an organized experience of distinct sound sources located in space. Its central problem is inverse and ill-posed: the eardrum receives a single summed waveform, yet the listener hears a violin, a voice, and a passing car as separate things. This article follows the problem from transduction in the cochlea, through the perceptual dimensions the system extracts, pitch and spatial location, to auditory scene analysis, the parsing of the mixture into streams, and the temporal-coherence account of how features are bound. It then turns to auditory objects and the cortical what and where pathways, and closes with perceptual restoration and the questions the framework leaves open. Three interactive demonstrations model the missing fundamental, binaural localization, and stream segregation.
Keywords: auditory perception, pitch, sound localization, auditory scene analysis, temporal coherence
Hearing is the sense that never closes. Sound arrives from every direction at once, passes through obstacles that stop light, and reaches the ears as a single fluctuating pressure at each eardrum, the summed contribution of every source in the environment mixed into one waveform. From this the auditory system must recover the number of sources, what each is, where each lies, and what each means, continuously and in real time. This article follows that achievement in order, from the mechanics that convert sound into a neural code, through the perceptual attributes the brain computes from that code, to the inverse problem of organizing a mixture into the separate objects a listener actually hears.
- Auditory perception solves an inverse problem: recovering many distinct sound sources from the single summed waveform at each ear.
- The cochlea performs a mechanical frequency analysis, mapping frequency to place along the basilar membrane and preserving fine timing in the pattern of neural firing.
- Pitch is derived from both the place of excitation and the temporal pattern of the waveform, which is why a tone with its fundamental removed still sounds at the fundamental pitch.
- Sound is localized by comparing the two ears: interaural time differences dominate at low frequencies and interaural level differences at high frequencies.
- Auditory scene analysis groups sound into streams over time and frequency; temporal coherence among a source's features is the proposed binding cue, and attention shapes what is heard.
What Auditory Perception Is
Auditory perception is the set of processes that take the fluctuating air pressure at the two eardrums and deliver an organized experience of sound sources: their identity, their location, and their meaning. The defining difficulty is that the input is a single mixture. Where the eye receives a spatial image in which different objects fall on different receptors, the ear receives one summed pressure wave at each side, into which every simultaneous source has been added together. The listener nonetheless hears a scene of separate things, and the work of audition is to invert that summation, recovering the sources that produced the wave from the wave alone.
That recovery is organized around the notion of an auditory object: a perceptual entity that corresponds to a sound source and persists as a coherent thing across time even as its acoustic properties change (Griffiths & Warren, 2004). Building such objects from a mixture is the problem the psychologist Albert Bregman named auditory scene analysis, by analogy with the way the visual system organizes an image into surfaces and objects (Bregman, 1990). The rest of this article follows the problem from the periphery inward: first the transduction that converts sound to a neural code, then the attributes, pitch and location, that the system computes, then the grouping processes that assemble those attributes into objects, and finally the cortical systems that carry the computation and the questions that remain open.
From Sound to a Neural Code
Auditory perception begins with a mechanical frequency analysis in the inner ear. Sound entering the cochlea sets up a travelling wave along the basilar membrane, a structure whose stiffness and mass vary along its length so that each place responds maximally to a particular frequency, high frequencies near the base and low frequencies at the apex. This produces a tonotopic map: frequency is converted into a place of maximal excitation, and that orderly mapping of frequency onto position is preserved all the way up the auditory pathway into the cortex (Moore, 2012).
Place alone does not carry all the information. The transduction itself is performed by the hair cells of the organ of Corti, whose stereocilia open ion channels when deflected, converting mechanical motion into a receptor potential with extraordinary speed and sensitivity (Hudspeth, 1989). Because hair cells can follow the fine structure of the waveform up to a few kilohertz, auditory nerve fibres tend to fire at a consistent phase of the stimulating cycle, a property called phase locking. The auditory system therefore holds two codes for frequency at once: a place code, given by which fibres are most active, and a temporal code, given by the timing of their spikes. The dual code is the foundation for the two attributes considered next, pitch and spatial location, each of which draws on place and timing in a different proportion.
Pitch Perception
Pitch is the perceptual attribute that lets sounds be ordered from low to high and that carries melody in music and intonation in speech. For a periodic sound it corresponds to the repetition rate of the waveform, the fundamental frequency, and the classic puzzle of pitch is that this percept survives the removal of the very frequency it names. When the fundamental is filtered out of a complex tone, leaving only its higher harmonics, the sound continues to be heard at the pitch of the missing fundamental, because that fundamental is the common spacing of the remaining harmonics and the period the whole pattern still repeats at (Oxenham, 2012). Pitch is thus not read off the energy at one frequency but computed from the pattern across many.
Two families of mechanism have been proposed, corresponding to the two codes the periphery supplies. Place theories derive pitch from the pattern of excitation along the basilar membrane, resolving the lower harmonics at distinct places and inferring the fundamental from their spacing; temporal theories derive it from the periodicity in the phase-locked firing of the nerve, which repeats at the fundamental rate whether or not energy is present there (Oxenham, 2012). Neither alone accounts for all the data: the most salient and musically useful pitches come from the low, resolved harmonics that place coding handles best, yet a pitch persists for high-harmonic and unresolved stimuli that only timing can explain (Moore, 2012). The first demonstration builds a harmonic tone and lets the reader remove harmonics, showing that the perceived pitch tracks the spacing of whatever harmonics remain rather than the presence of the fundamental.
Demonstration 1
The missing fundamental
fundamentalhigher harmonicperceived pitch
Sound Localization
A mixture recovered as objects is of little use unless those objects can be placed in space, and here the auditory system exploits the fact that it has two ears separated by a head. A sound off to one side reaches the near ear slightly before the far ear and slightly louder, because the head casts an acoustic shadow. These two cues divide the frequency range between them. At low frequencies, where the wavelength is long compared with the head, the interaural time difference, the lead of one ear over the other, is the dominant and unambiguous cue; at high frequencies, where the head casts an effective shadow, the interaural level difference dominates, an arrangement long summarized as the duplex theory of localization (Middlebrooks & Green, 1991).
Extracting a microsecond time difference requires machinery of remarkable precision. The influential proposal is a set of coincidence-detecting neurons fed by axons of differing length from the two ears, so that a given neuron fires maximally when the delay along its two input lines exactly cancels the interaural time difference, converting a difference in arrival time into a place code for azimuth (Jeffress, 1948). This delay-line model organized decades of work on the binaural system, and while its exact form is debated, its central idea, that the brain localizes by comparing timing across the two ears through coincidence detection, remains foundational. The second demonstration lets the reader move a source around the head and reads out the interaural time difference the two cues imply.
Demonstration 2
Locating a sound with two ears
Auditory Scene Analysis
The deepest problem in hearing is not extracting an attribute from a clean sound but recovering many sounds from a mixture, the problem that selective listening at a crowded party first made vivid, where a listener can follow one talker among many and yet notice their own name spoken across the room (Cherry, 1953). Bregman framed this as auditory scene analysis and distinguished two kinds of grouping. Simultaneous grouping decides which frequency components present at the same moment belong to the same source, using cues such as common onset, harmonic relationship, and common changes over time. Sequential grouping decides which sounds arriving one after another belong to the same source over time, building the perceptual streams that carry a melody or a voice (Bregman, 1990). Table 1 lists the principal grouping cues and what each contributes.
| Cue | Grouping type | What it exploits |
|---|---|---|
| Common onset | Simultaneous | Components that start together tend to arise from one source |
| Harmonicity | Simultaneous | Frequencies that are integer multiples of one fundamental fuse into one tone |
| Temporal coherence | Simultaneous | Channels whose activity rises and falls together belong to one source |
| Frequency proximity | Sequential | Successive tones near in frequency form one stream; distant tones split |
| Temporal continuity | Sequential | Smoothly continuing sounds are followed as a single ongoing source |
Sequential streaming has been studied most closely with sequences of alternating tones. When two tones of different frequency alternate, they are heard either as a single connected stream that jumps up and down, or, when the frequency separation is large or the sequence fast, as two separate streams each of steady pitch, one high and one low, that the listener cannot easily recombine. This build-up of stream segregation has a neural counterpart: recordings from primary auditory cortex show that responses to the two tones become progressively better separated with repetition and with frequency distance, in a way that tracks the perceptual split (Micheyl et al., 2005). The third demonstration realizes this alternating-tone paradigm, letting the reader vary the frequency separation and the rate and watch the point at which one stream splits into two.
Demonstration 3
One stream or two?
integratedhigh streamlow stream
Temporal Coherence and Attention
What binds the many frequency components of one source together and separates them from another? An influential answer is temporal coherence: components whose neural responses rise and fall together over time are bound into a single stream, while components that fluctuate independently are assigned to different sources (Shamma et al., 2011). On this account the auditory system continually measures the correlation between the activity in different frequency channels, and coherent channels are grouped whatever their absolute frequency, which explains how a talker's harmonics, spread across the spectrum, cohere into one voice because they all modulate in step with the glottal pulse and the syllabic rhythm.
Crucially, this binding is not purely automatic. Temporal coherence is proposed to operate together with attention, so that attending to one feature of a source, its pitch or its location, pulls the coherent remainder of that source into the perceptual foreground and pushes competing sources back (Shamma et al., 2011). This makes auditory grouping partly object-based: attention selects an auditory object as a whole rather than a raw frequency band, and the features bound to that object are selected with it (Shinn-Cunningham, 2008). The account thus links the low-level grouping cues of scene analysis to the high-level control of attention, treating them as two aspects of a single process that forms and selects auditory objects.
Auditory Objects and the What and Where Pathways
The auditory object, the perceptual thing that grouping produces, is what higher processing acts upon. An auditory object is a representation of a sound source that binds its features, pitch, timbre, and location, into a unit that can be attended, remembered, and recognized as the same source across changes in its acoustics (Bizley & Cohen, 2013). Defining objecthood for sound is harder than for vision, because a sound has no fixed spatial boundary and exists only over time, but the concept does real work: it explains why a listener can track a single instrument through an orchestra, and why the features of an unattended source are perceived as belonging together even when ignored (Griffiths & Warren, 2004).
Beyond the primary areas, the processing of these objects appears to divide along two pathways, echoing the organization of vision. Physiological work in the primate identifies an anterior stream specialized for the identity of a sound, its what, and a posterior stream specialized for its spatial location, its where, projecting to partly distinct regions of the frontal lobe (Rauschecker & Tian, 2000). The division is not absolute, and the two streams interact, but it captures a real functional split: recognizing which source a sound is and computing where it lies are partly separable operations carried by partly separate cortex (Bizley & Cohen, 2013). Figure 1 traces the whole path, from the cochlea's dual code through grouping into an auditory object to the two cortical streams.
Figure 1
From the Cochlea to the What and Where Pathways
Note. Schematic of the auditory pathway (after Rauschecker & Tian, 2000, and Bizley & Cohen, 2013). The cochlea supplies a place-and-timing code that scene analysis groups into an auditory object; anterior and posterior cortical streams then compute the object's identity and location in parallel. The diagram is illustrative, not drawn from data.
The Auditory Cortex
The cortical destination of the auditory pathway is the superior temporal plane, where a core of primary areas, tonotopically organized like the cochlea that feeds them, is surrounded by belt and parabelt regions that respond to progressively more complex sounds (Rauschecker & Tian, 2000). Ascending the hierarchy, tuning shifts from pure tones toward the spectrotemporal patterns that characterize natural sounds such as speech and music, so that later areas represent behaviourally meaningful sound features rather than raw frequency.
A striking property of this cortex is its functional asymmetry. The two hemispheres appear to trade off temporal against spectral resolution: the left auditory cortex resolves rapid temporal change well, suiting it to the fast transitions of speech, while the right resolves fine frequency detail well, suiting it to the pitch relations of melody (Zatorre et al., 2002). This trade-off is a natural consequence of a fundamental limit, since a system cannot simultaneously achieve arbitrarily fine resolution in both time and frequency, and it offers an elegant account of why speech and music, though both carried by the same ascending pathway, engage the two hemispheres to different degrees (Zatorre et al., 2002). The cortex, on this view, is not a single general-purpose analyzer but a set of regions tuned to the statistical structure of the sounds that matter.
Perceptual Restoration and Sound Texture
Because perception is inference from an incomplete and noisy signal, the auditory system routinely fills in what the acoustics leave out. The clearest case is phonemic restoration: when a speech sound is excised and replaced by a cough or a burst of noise, listeners report hearing the missing sound clearly and cannot correctly locate the interruption, because the system uses lexical and contextual knowledge to reconstruct the masked segment (Warren, 1970). The percept is not of a gap but of a complete word, evidence that what is heard is the system's best interpretation of the source rather than a transcript of the signal at the ear.
Inference also operates on the many everyday sounds that have no clear pitch or event structure at all, such as rain, wind, or a crackling fire. These sound textures are perceived through their time-averaged statistics rather than their moment-to-moment detail: synthesizing a novel sound that merely matches the statistics of the auditory periphery is enough to reproduce the perceived identity of a texture, showing that the system represents such sounds by summary statistics computed over time (McDermott & Simoncelli, 2011). Restoration and texture perception make the same point from two directions: auditory perception is a constructive, inferential process that represents the probable source, not a passive registration of the waveform.
What the Framework Does and Does Not Explain
The consensus picture, a peripheral frequency analysis feeding grouping processes that build attended auditory objects along what and where pathways, is well supported, but several of its central terms remain incompletely specified. The first is the auditory object itself. Unlike a visual object it has no agreed definition or boundary, and different researchers operationalize it differently, so the concept that anchors much of the field is doing more intuitive than technical work and its neural signature is still contested (Bizley & Cohen, 2013; Griffiths & Warren, 2004). Until objecthood is pinned down, claims about how objects are formed and selected inherit that vagueness.
A second open question concerns the mechanism of grouping. Temporal coherence is an attractive and quantitative binding cue, but how it is computed neurally, how it interacts with the classical cues of harmonicity and common onset, and how much of streaming is driven by peripheral channelling rather than central correlation are not settled (Shamma et al., 2011; Micheyl et al., 2005). What is not in dispute is the shape of the problem and the broad architecture that solves it: sound is analyzed into frequency and timing at the periphery, grouped into streams and objects by cues that include temporal coherence, selected by attention, and processed for identity and location by partly separate cortical systems (Bregman, 1990; Rauschecker & Tian, 2000). The debates concern the definition of the object and the mechanism of binding, not whether the scene is analyzed at all.
Worked Example
Consider the three demonstrations with their default constants. The pitch demonstration builds a harmonic complex on a fundamental of 200 hertz, so the components lie at 200, 400, 600, 800, and 1000 hertz. The perceived pitch equals the repetition rate of the whole pattern, which is the greatest common divisor of the present components. With all five present the divisor is 200 hertz. Removing the 200-hertz fundamental leaves 400, 600, 800, and 1000, whose greatest common divisor is still 200, so the perceived pitch is unchanged at 200 hertz, the missing fundamental. Only when the remaining components no longer share 200 as a divisor, for instance if just 800 and 1200 remained, giving a divisor of 400, does the perceived pitch shift. The demonstration computes the divisor of whatever harmonics the reader leaves in and confirms it stays at 200 across the removal of the fundamental.
The localization demonstration models the interaural time difference with the Woodworth formula, in which the delay equals the head radius divided by the speed of sound, times the quantity azimuth in radians plus the sine of the azimuth. Taking the head radius as 0.0875 metres and the speed of sound as 343 metres per second gives a leading factor of about 255 microseconds per radian. Straight ahead, at zero degrees, the difference is zero. At thirty degrees, which is 0.524 radians, the formula gives 255 times the quantity 0.524 plus 0.500, or about 261 microseconds. At ninety degrees, fully to one side, it gives 255 times the quantity 1.571 plus 1.000, or about 656 microseconds, the maximum. Moving the source in the demonstration traces this curve, and the readout crosses from a time-difference regime to a level-difference regime as the modelled frequency rises, illustrating the duplex theory. The streaming demonstration is explicitly illustrative rather than measured: it estimates a probability of hearing two streams as a logistic function of the frequency separation in semitones minus a threshold that falls as the presentation rate rises, so that a large separation at a fast rate segregates into two streams while a small separation at a slow rate coheres into one, reproducing the qualitative boundary found by experiment without asserting particular measured values.
Discussion
Auditory perception is best understood as the solution to an inverse problem posed under an unusually severe constraint: the entire scene arrives as one pressure wave at each ear, and the brain must recover from it the sources that summed to produce it. The system's answer begins in the mechanics of the cochlea, which splits the wave into frequency channels and preserves its fine timing, supplying the place and temporal codes from which pitch and spatial location are computed (Hudspeth, 1989; Oxenham, 2012; Middlebrooks & Green, 1991). On top of this peripheral analysis the system runs the grouping processes of auditory scene analysis, using onset, harmonicity, and above all temporal coherence to bind a source's scattered components into a stream and to segregate competing sources, under the guidance of an attention that operates on whole objects (Bregman, 1990; Shamma et al., 2011; Shinn-Cunningham, 2008).
The objects that grouping produces are then processed for identity and location along partly separate cortical pathways, in a hierarchy that moves from tonotopic frequency maps to representations of behaviourally meaningful sound, with the two hemispheres trading temporal against spectral resolution to serve speech and music (Rauschecker & Tian, 2000; Zatorre et al., 2002; Bizley & Cohen, 2013). That the whole process is inferential rather than passive is shown most clearly when the signal is incomplete or unstructured, in the restoration of masked speech and the representation of sound textures by their summary statistics (Warren, 1970; McDermott & Simoncelli, 2011). The enduring questions, what exactly an auditory object is and how temporal coherence is computed, are questions about the mechanism and vocabulary of a system whose overall design is not in doubt.
Glossary
- Auditory object.
- A perceptual representation of a sound source that binds its features into a unit persisting across time, the entity that attention selects and memory stores.
- Auditory scene analysis.
- The set of processes that organize a mixture of sounds into perceptual streams corresponding to distinct sources, comprising simultaneous and sequential grouping.
- Basilar membrane.
- The structure within the cochlea whose graded stiffness makes each place respond best to a particular frequency, producing the tonotopic map.
- Duplex theory.
- The account of sound localization in which interaural time differences dominate at low frequencies and interaural level differences at high frequencies.
- Hair cell.
- A sensory receptor of the cochlea whose stereocilia open ion channels when deflected, converting mechanical motion of the basilar membrane into a neural signal.
- Interaural level difference.
- The difference in sound intensity between the two ears, caused by the acoustic shadow of the head and used to localize high-frequency sounds.
- Interaural time difference.
- The difference in arrival time of a sound at the two ears, the dominant cue for localizing low-frequency sounds in the horizontal plane.
- Missing fundamental.
- The phenomenon in which a complex tone whose fundamental frequency has been removed is still heard at the pitch of that fundamental, inferred from the spacing of its harmonics.
- Phase locking.
- The tendency of auditory nerve fibres to fire at a consistent phase of the sound waveform, providing a temporal code for frequency up to a few kilohertz.
- Phonemic restoration.
- The illusory perception of a speech sound that has been physically replaced by noise, in which context and lexical knowledge fill in the missing segment.
- Pitch.
- The perceptual attribute that orders sounds from low to high, corresponding for periodic sounds to the repetition rate of the waveform, the fundamental frequency.
- Stream segregation.
- The perceptual splitting of a sequence of sounds into two or more separate streams, promoted by large frequency separation and fast presentation rate.
- Temporal coherence.
- The proposed binding cue by which frequency components whose neural responses fluctuate together over time are grouped into a single auditory object.
- Tonotopy.
- The orderly mapping of sound frequency onto position, established on the basilar membrane and preserved through the auditory pathway into the cortex.
- What and where pathways.
- Partly separate cortical streams for processing the identity of a sound (anterior) and its spatial location (posterior), projecting to distinct frontal regions.
Key Researchers
Albert S. Bregman. Canadian cognitive psychologist at McGill University; he defined and organized the field of auditory scene analysis, describing how the auditory system parses a mixture of sounds into perceptual streams belonging to distinct sources. Faculty Page - Google Scholar - Wikipedia
Brian C. J. Moore. British auditory psychophysicist at the University of Cambridge; his work on pitch, frequency selectivity, and masking, and on modelling cochlear hearing loss, shaped the modern psychology of hearing. Faculty Page - Google Scholar - Wikipedia) - ORCID
Andrew J. Oxenham. Auditory perception researcher at the University of Minnesota; he studies the perceptual and neural coding of pitch, showing how the system extracts pitch from harmonic structure and temporal fine structure. Faculty Page - Google Scholar - ORCID
Josh H. McDermott. Auditory neuroscientist at the Massachusetts Institute of Technology; he pioneered the study of sound-texture perception through statistical synthesis and builds computational models of human audition. Faculty Page - Google Scholar - ORCID
Shihab A. Shamma. Auditory neuroscientist at the University of Maryland; he developed the temporal-coherence theory of auditory scene analysis and studies cortical spectrotemporal processing. Faculty Page - Google Scholar - ORCID
Robert J. Zatorre. Cognitive neuroscientist at the Montreal Neurological Institute, McGill University; he established the functional asymmetry of auditory cortex, trading temporal resolution for speech against spectral resolution for melody. Faculty Page - Google Scholar - ORCID
Frequently Asked Questions
What is auditory perception?
It is the set of processes by which the brain turns the summed pressure wave at the two ears into an organized experience of separate sound sources, recovering their identity, location, and meaning from a single mixture (Bregman, 1990; Griffiths & Warren, 2004).
How does the ear separate different pitches?
The cochlea performs a mechanical frequency analysis, with each place on the basilar membrane responding best to a particular frequency, so that frequency is mapped onto position in an orderly tonotopic code preserved up to the cortex (Moore, 2012; Hudspeth, 1989).
Why can a tone with no fundamental still have a pitch?
Because pitch reflects the repetition rate of the whole waveform, which equals the common spacing of the harmonics; removing the fundamental leaves that spacing unchanged, so the sound is still heard at the missing fundamental (Oxenham, 2012).
How do we tell where a sound is coming from?
By comparing the two ears: a sound reaches the near ear sooner and louder, and the system uses the interaural time difference at low frequencies and the interaural level difference at high frequencies to compute direction (Middlebrooks & Green, 1991; Jeffress, 1948).
How do we follow one voice in a noisy room?
The auditory system groups the sound into streams by auditory scene analysis and binds each source's components by their temporal coherence, while attention selects one whole object and pushes competing sources into the background (Cherry, 1953; Shamma et al., 2011).
What is an auditory object?
It is a perceptual representation of a sound source that binds its pitch, timbre, and location into a unit persisting across changes in the acoustics, the thing that can be attended and recognized as the same source over time (Bizley & Cohen, 2013).
Are the identity and location of a sound processed separately?
Largely yes; physiological work identifies an anterior what stream specialized for the identity of a sound and a posterior where stream specialized for its location, projecting to partly distinct frontal regions (Rauschecker & Tian, 2000).
Does the brain fill in sounds that are not there?
Yes; in phonemic restoration a speech sound replaced by noise is heard as present, and sound textures are represented by their time-averaged statistics, both showing that perception infers the probable source rather than transcribing the signal (Warren, 1970; McDermott & Simoncelli, 2011).
References
Bizley, J. K., & Cohen, Y. E. (2013). The what, where and how of auditory-object perception. Nature Reviews Neuroscience, 14(10), 693-707. https://doi.org/10.1038/nrn3565
Bregman, A. S. (1990). Auditory scene analysis: The perceptual organization of sound. MIT Press. https://doi.org/10.7551/mitpress/1486.001.0001
Cherry, E. C. (1953). Some experiments on the recognition of speech, with one and with two ears. The Journal of the Acoustical Society of America, 25(5), 975-979. https://doi.org/10.1121/1.1907229
Griffiths, T. D., & Warren, J. D. (2004). What is an auditory object? Nature Reviews Neuroscience, 5(11), 887-892. https://doi.org/10.1038/nrn1538
Hudspeth, A. J. (1989). How the ear's works work. Nature, 341(6241), 397-404. https://doi.org/10.1038/341397a0
Jeffress, L. A. (1948). A place theory of sound localization. Journal of Comparative and Physiological Psychology, 41(1), 35-39. https://doi.org/10.1037/h0061495
McDermott, J. H., & Simoncelli, E. P. (2011). Sound texture perception via statistics of the auditory periphery: Evidence from sound synthesis. Neuron, 71(5), 926-940. https://doi.org/10.1016/j.neuron.2011.06.032
Micheyl, C., Tian, B., Carlyon, R. P., & Rauschecker, J. P. (2005). Perceptual organization of tone sequences in the auditory cortex of awake macaques. Neuron, 48(1), 139-148. https://doi.org/10.1016/j.neuron.2005.08.039
Middlebrooks, J. C., & Green, D. M. (1991). Sound localization by human listeners. Annual Review of Psychology, 42, 135-159. https://doi.org/10.1146/annurev.ps.42.020191.001031
Moore, B. C. J. (2012). An introduction to the psychology of hearing (6th ed.). Brill.
Oxenham, A. J. (2012). Pitch perception. The Journal of Neuroscience, 32(39), 13335-13338. https://doi.org/10.1523/JNEUROSCI.3815-12.2012
Rauschecker, J. P., & Tian, B. (2000). Mechanisms and streams for processing of "what" and "where" in auditory cortex. Proceedings of the National Academy of Sciences, 97(22), 11800-11806. https://doi.org/10.1073/pnas.97.22.11800
Shamma, S. A., Elhilali, M., & Micheyl, C. (2011). Temporal coherence and attention in auditory scene analysis. Trends in Neurosciences, 34(3), 114-123. https://doi.org/10.1016/j.tins.2010.11.002
Shinn-Cunningham, B. G. (2008). Object-based auditory and visual attention. Trends in Cognitive Sciences, 12(5), 182-186. https://doi.org/10.1016/j.tics.2008.02.003
Warren, R. M. (1970). Perceptual restoration of missing speech sounds. Science, 167(3917), 392-393. https://doi.org/10.1126/science.167.3917.392
Zatorre, R. J., Belin, P., & Penhune, V. B. (2002). Structure and function of auditory cortex: Music and speech. Trends in Cognitive Sciences, 6(1), 37-46. https://doi.org/10.1016/S1364-6613(00)01816-7