How We Hear in 3D – The Neuroscience Behind Stereo Perception

1024 1024 Michelangelo

“Phase is not the sound — it is the space between sounds.” Taken as a metaphor, this captures something fundamental about human hearing: the brain does not process the signals arriving at our two ears independently. It continuously compares them, using extraordinarily small differences in time, level and spectral balance to reconstruct the position of sound around us.

In high-end audio, we often concentrate on equipment, formats, resolution and frequency response. Yet beneath every recording and every playback system lies a more fundamental question: how does the human auditory system transform two incoming signals into a convincing three-dimensional world?

Stereo reproduction works not because two loudspeakers recreate the original sound field perfectly, but because they provide the auditory system with enough carefully organised information to create a plausible spatial scene. Understanding that process changes how we think about microphone placement, phase relationships, room acoustics and the meaning of realism in recording.

Hearing Is a Reconstruction

Sound arriving at the ears does not contain a ready-made map of the space around us. The auditory system must infer the direction, distance and environment of a sound source from a collection of acoustic clues.

For horizontal localisation, the most important binaural cues are differences in arrival time and sound level between the two ears. For elevation and front-to-back discrimination, the auditory system also relies heavily on direction-dependent spectral filtering created by the listener’s head, torso and outer ears.

These mechanisms operate together. They should not be understood as three entirely separate systems, nor as rigid frequency zones with precise boundaries. Natural sounds are usually broadband, reflections complicate the incoming signals, and the brain integrates multiple cues over time.

The Principal Cues to Spatial Hearing

1. Interaural Time Differences

When a sound source is positioned to one side of the listener, the sound normally reaches the nearer ear slightly before reaching the farther ear. This difference is known as the interaural time difference, or ITD.

The delays involved are extremely small—often measured in microseconds—but the auditory system is remarkably sensitive to them. For low-frequency sounds, neural activity can follow the temporal structure of the waveform closely enough for the brain to compare the timing received at the two ears.

ITD sensitivity to the fine structure of pure tones becomes progressively less effective as frequency rises, with human sensitivity deteriorating sharply around the region of approximately 1.4 to 1.5 kHz. This should be treated as a broad transition rather than a universal dividing line. High-frequency sounds can still carry timing information through changes in their amplitude envelope.

2. Interaural Level Differences

A sound arriving from one side will also tend to be louder at the nearer ear. At shorter wavelengths, the head obstructs part of the sound travelling towards the farther ear, producing an acoustic shadow. The resulting difference in level is known as the interaural level difference, or ILD.

Level differences generally become more pronounced at higher frequencies because shorter wavelengths are more strongly affected by the head. However, ITD and ILD do not simply exchange responsibility at one exact frequency. For complex sounds, the auditory system can combine timing and level information across several frequency regions.

The relative importance of these cues also changes with the sound itself, its distance, the surrounding reflections and the listener’s hearing.

3. Spectral Cues and the Head-Related Transfer Function

Time and level differences are especially useful for identifying whether a sound is located towards the left or right. They are less able, by themselves, to resolve whether a source is above, below, in front of or behind the listener.

For this, the complex shape of the outer ear becomes essential. The folds of the pinna, together with the head and upper body, alter the spectrum of incoming sound in a direction-dependent way. Some frequencies are reinforced, while others are attenuated or notched.

The complete acoustic transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because every person’s anatomy is different, HRTFs are individual. The brain gradually learns the spectral patterns associated with particular directions. Experiments in which the shape of the outer ear was temporarily altered have shown that localisation initially becomes less accurate, particularly for elevation and front-to-back judgements, but can improve again as listeners adapt to the modified cues.

Where Phase Fits In

The word phase is used in several related but distinct ways in audio, which can easily create confusion.

At low frequencies, a difference in arrival time between the ears can also be described as an interaural phase difference for a periodic waveform. In this context, phase difference is one of the ways the auditory system obtains spatial information.

Neurons within the auditory brainstem, including those associated with the medial superior olive, are specialised for processing extremely small interaural timing differences. Their responses contribute to the neural representation of horizontal sound direction.

However, it would be too simple to conclude that phase is used only to locate a sound and never contributes to what that sound is. The timing structure of a waveform—often described as its temporal fine structure—also contributes to aspects of pitch perception, auditory masking and the separation of sounds in complex listening environments.

The more useful distinction for recording engineers is therefore not between “phase” and “sound identity,” but between the different roles that temporal and phase relationships can play.

Within a stereo recording, relationships between the two channels may affect:

  • the apparent position of a phantom image;
  • the perceived width and stability of the soundstage;
  • the impression of depth and surrounding ambience;
  • the result when the recording is reproduced in mono;
  • frequency-response changes caused by constructive and destructive interference.

Phase is therefore neither an isolated technical curiosity nor a universal explanation for every spatial quality. It is one component within a larger system of timing, level, spectrum, reflection and playback interaction.

Why This Matters for Stereo Recording

Stereo microphone techniques create spatial information in different ways.

A spaced pair such as AB introduces arrival-time differences between the microphones and may also produce level differences. A near-coincident configuration such as ORTF deliberately combines microphone spacing with directional level differences.

Coincident techniques such as XY and Mid-Side minimise the arrival-time difference between microphones and create direction mainly through differences in level. The Blumlein pair, using two coincident figure-of-eight microphones, also derives its directional information from the polar patterns and polarity relationships of the two channels while capturing substantial information from the surrounding acoustic environment.

None of these methods is automatically natural or unnatural in every situation. Each encodes the original acoustic event differently, and each interacts differently with loudspeakers, headphones, room reflections and listener position.

This is particularly important because conventional stereo loudspeaker reproduction does not send the left channel exclusively to the left ear or the right channel exclusively to the right ear. Each loudspeaker reaches both ears, introducing additional timing, level and spectral interactions. The listening room then adds its own reflections.

The task of the recording engineer is therefore not merely to create a wide image. It is to create interchannel relationships that remain meaningful when reproduced through the intended playback system.

From Spatial Effect to Spatial Credibility

A stereo recording can sound spectacularly wide while still producing unstable localisation, exaggerated scale or an uncertain centre image. Conversely, a narrower presentation may feel more convincing because its timing, level and reverberant cues form a more internally consistent spatial picture.

This suggests a more useful question than simply asking whether a recording sounds spacious:

Do the spatial cues reproduced by the system support one another strongly enough for the brain to construct a stable and believable acoustic scene?

Phase coherence is part of that question, but it should not be treated as a single measurement that determines realism on its own. Microphone polar pattern, spacing, angle, source distance, direct-to-reverberant ratio, loudspeaker placement and room acoustics all contribute to the final perception.

What Comes Next

In the next article, I will examine how AB, ORTF, XY, Mid-Side and Blumlein recording techniques encode spatial information differently, and why a technique that produces impressive width does not necessarily produce the most credible depth or localisation.

I will also explore the relationship between coincident microphone techniques, loudspeaker reproduction, room acoustics and the preservation of stable interchannel relationships.

The central principle is simple: recording technology should not be considered separately from human perception. The microphone arrangement, recording space and playback format are parts of one perceptual chain.

References and Further Reading

  1. Brughera, A., Dunai, L. and Hartmann, W. M. “Human interaural time difference thresholds for sine tones: the high-frequency limit.” Journal of the Acoustical Society of America, 2013.
    View publication
  2. Wightman, F. L. and Kistler, D. J. “The dominant role of low-frequency interaural time differences in sound localization.”
    Journal of the Acoustical Society of America, 1992.
    View publication
  3. Salminen, N. H., Tiitinen, H., Yrttiaho, S. and May, P. J. C. “The neural code for interaural time difference in human auditory cortex.”
    Journal of the Acoustical Society of America, 2010.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning sound localization with new ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Moore, B. C. J. “The role of temporal fine structure processing in pitch perception, masking, and speech perception for normal-hearing and hearing-impaired people.”
    Journal of the Association for Research in Otolaryngology, 2008.
    View publication
  6. Eargle, J. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View publication

An earlier version of this article was published on LinkedIn. This Direct Sound Records Journal edition has been revised, expanded and technically updated, with additional context and scientific references.

Author

Michelangelo

All stories by: Michelangelo