Posts Tagged :

spatial hearing

Space and Sound

The Space Between Sounds: How the Brain Reconstructs the Soundstage

1024 755 Michelangelo

In musical realism, we do not listen only to frequencies, timbres and dynamics. We listen to relationships: differences in arrival time, level, reflection and spatial organisation. When those relationships remain credible, a recording can become more than a collection of sounds. It can become an inhabitable acoustic space.

When listening to a high-fidelity system, we often say that a recording “sounds real.” But what does that actually mean?

Frequency extension matters. So do controlled bass, natural midrange, transient response, low distortion and harmonic detail. Yet realism also depends on something less obvious and considerably more fragile: whether the recording and playback system allow us to perceive a believable relationship between sound sources and the space around them.

A voice is not simply a voice. It is a presence located at a particular distance and position.

A piano is not merely a collection of strings, hammers and resonances. It is a physical body occupying a volume of air, projecting energy into a room and interacting with surrounding surfaces.

A quartet is not simply divided into left, centre and right. It is an acoustic event in which musicians, distances, reflections and silence form one spatial relationship.

The human auditory system does not receive this scene passively. It interprets, compares and reconstructs it. In a sense, the brain triangulates.

The Soundstage Is Not an Effect

In high-end audio, words such as soundstage, depth, focus, air and presence are used constantly. They are useful descriptions, but they can become vague unless we remember that they have a perceptual foundation.

The soundstage is not a picture physically stored inside a recording. Nor is it simply drawn between two loudspeakers.

It is a perceptual reconstruction created from the information reaching the listener’s ears.

The brain interprets several cues simultaneously:

  • tiny differences in the time at which sound reaches each ear;
  • differences in level between the ears;
  • direction-dependent changes in the spectrum;
  • the relationship between direct and reflected sound;
  • the evolution of reflections and reverberation over time;
  • and prior knowledge of familiar voices, instruments and environments.

When these cues support one another, a sound source can appear stable, physical and separate from the loudspeakers.

When they conflict, a recording may still sound impressive, detailed or extremely wide, but the scene can feel unstable or artificial.

The soundstage is therefore not merely an effect.

It is information interpreted by the listener.

How the Brain Reads Acoustic Space

For horizontal localisation, the auditory system relies heavily on differences in timing and level between the two ears.

Interaural Time Differences

When a sound arrives from one side, it normally reaches the nearer ear slightly before the farther ear. These interaural time differences, or ITDs, can be extremely small—sometimes measured in tens of microseconds—yet the auditory system is remarkably sensitive to them.

Timing cues are particularly important at lower frequencies, where neural activity can represent the temporal fine structure of the waveform with sufficient precision. Structures within the auditory brainstem, especially the medial superior olive, contribute to the analysis of these binaural timing relationships.

ITD sensitivity does not end at one perfectly defined frequency, but sensitivity to the fine structure of pure tones deteriorates substantially through the region around 1 to 1.5 kHz. Higher-frequency sounds can still convey timing information through changes in their amplitude envelopes.

Interaural Level Differences

At shorter wavelengths, the head creates a more substantial acoustic shadow. A sound arriving from the left will generally produce a higher level at the left ear than at the right.

This interaural level difference, or ILD, becomes an increasingly useful directional cue as frequency rises. Neural circuits involving the lateral superior olive contribute to its early processing.

In natural listening, timing and level differences do not operate as two isolated systems with a rigid boundary between them. Broadband sounds contain multiple cues, and the brain combines them according to frequency, source position, environment and reliability.

Spectral Cues and the Shape of the Listener

Timing and level differences provide strong information about left and right, but they cannot always distinguish whether a sound is above, below, in front of or behind the listener.

For those dimensions, the shape of the outer ears, head and torso becomes essential.

The folds of the pinnae filter incoming sound differently according to direction. Certain frequencies are reinforced, while others are attenuated or notched. The complete transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because human anatomy varies, each person’s HRTF is individual. Over time, the brain learns the correspondence between these spectral patterns and positions in space.

Experiments in which the outer ears were temporarily reshaped have shown that vertical localisation initially becomes much less accurate. With experience, listeners can learn to interpret the altered cues—evidence that spatial hearing is not merely mechanical, but calibrated and adaptive.

What we perceive is therefore never only the sound emitted by the source. It is the result of a relationship between source, environment, listener and brain.

Where Phase Enters the Picture

In audio engineering, phase is often discussed as a problem.

Signals may be described as “out of phase” when they weaken the centre image, reduce low-frequency energy, produce cancellations or behave unpredictably in mono.

These concerns are real, but they represent only part of the subject.

Phase is not simply an error waiting to be corrected. Phase and timing relationships can also carry important information about position, width, depth and the interaction between sound sources.

For a periodic low-frequency signal, a difference in arrival time between the ears can also be expressed as an interaural phase difference. In this sense, phase contributes directly to spatial localisation.

Within a stereo recording, interchannel timing and phase relationships can influence:

  • the position and stability of phantom images;
  • the apparent width of the presentation;
  • the sense of distance and depth;
  • the relationship between direct sound and ambience;
  • mono compatibility;
  • and frequency-dependent reinforcement or cancellation when signals combine.

However, it would be misleading to say that phase is used only for location and never contributes to the perceived character of a sound. Temporal fine structure is also involved in pitch, masking and the separation of simultaneous sources. In recording and reproduction, phase relationships can additionally alter the spectrum whenever correlated signals combine acoustically or electrically.

The useful distinction is therefore not between phase and sound, but between the different roles that temporal relationships perform.

Phase is not the sound itself, but it can help organise the space in which that sound is perceived.

Coherence Does Not Mean Perfection

The expression phase coherence is frequently used as though it described one measurable quality that a recording either possesses or lacks.

Reality is more complicated.

Every acoustic environment contains delays. Reflections arrive after the direct sound and from different directions. Instruments radiate differently according to frequency. Microphones have frequency-dependent polar patterns. Loudspeakers and rooms introduce further interactions.

A natural acoustic event is not phase-identical at every point in space.

Spatial coherence should therefore not mean eliminating every difference or delay. It means preserving relationships that remain compatible enough for the auditory system to interpret them as belonging to one plausible event.

When this happens, the voice can stabilise at the centre without appearing glued to either speaker. Instruments occupy a readable volume rather than appearing as thin lateral points. The room does not feel like reverberation placed behind the music; it surrounds and continues the performance.

The silence between instruments stops being empty.

It becomes air.

Why Some Recordings Sound Large but Not Real

A wide soundstage is not necessarily a realistic soundstage.

Modern production provides an enormous range of tools for creating size: multiple microphones, pan controls, delay, artificial reverberation, stereo widening, decorrelation and Mid-Side processing.

These tools can be artistically valuable. They can also create a guitar broader than its physical source, a voice floating beyond the loudspeakers or a reverberant field that could never have existed around the original performers.

There is nothing inherently wrong with that. Recording is also an art of construction.

But width and credibility are not synonymous.

The auditory system evaluates more than the apparent size of the scene. It also evaluates whether the timing, spectral, directional and environmental information is mutually plausible.

A very wide image may be initially impressive. Yet if the centre lacks stability, the reverberation does not belong to the sources or the spatial relationships change unnaturally with frequency, the illusion becomes less convincing.

It is similar to viewing a photograph with intense colour and extraordinary sharpness but incorrect perspective. The image attracts attention immediately, yet something feels wrong.

A recording can contain remarkable detail and separation while remaining spatially two-dimensional.

Beautiful, perhaps.

But not alive.

Every Microphone Technique Makes a Decision

For acoustic music, the choice of microphone technique is never neutral.

AB, ORTF, XY, Mid-Side and Blumlein do not simply create different varieties of stereo width. They encode different combinations of timing, level, polarity, direction and room information.

A spaced AB pair introduces meaningful arrival-time differences between microphones and can create scale, openness and envelopment.

ORTF combines a moderate physical separation with directional cardioid microphones, creating both interchannel time and level differences.

Coincident systems such as XY, Mid-Side and Blumlein minimise the timing difference introduced by microphone spacing and derive direction primarily through level and polarity relationships.

These differences affect localisation, spaciousness, mono compatibility and the way the recording interacts with loudspeaker reproduction.

Coincident techniques can produce comparatively stable and clearly located virtual sources. Spaced techniques may create broader or more diffuse images and can convey strong spaciousness. Neither outcome is automatically better.

The appropriate technique depends on:

  • the musicians and their physical arrangement;
  • the acoustic character of the venue;
  • the desired listening perspective;
  • the balance between localisation and envelopment;
  • the intended distribution format;
  • and the expected playback environment.

There is no microphone technique that is universally correct.

There is only a technique whose compromises are more or less coherent with the intended result.

Recording Is a Translation

A microphone does not hear like a human being.

It has no head, no outer ears, no perceptual memory and no awareness of the room. It does not compare what it captures with years of experience. It measures sound pressure or pressure gradient according to its physical construction and polar pattern.

Human perception, by contrast, is active.

This means that stereo recording is always a translation.

The task is not simply to place two microphones in front of a performance and assume that the original space has been preserved. The task is to create two signals that, when reproduced through loudspeakers in another room, provide the listener with enough coherent information to reconstruct a plausible scene.

That distinction is fundamental.

The original venue, the microphone array, the recording chain, the loudspeakers, the listening room and the listener are all parts of one perceptual system.

A recording can never transport the original acoustic field intact. It selects, encodes and later stimulates a new reconstruction.

The Listening Room Is Part of the Reproduction

With headphones, each channel is delivered predominantly to one ear. With conventional loudspeakers, both speakers reach both ears.

The left ear receives sound from the left loudspeaker, sound from the right loudspeaker after a different path, and reflections from the listening room. The right ear receives the corresponding combination from the opposite side.

The listener’s brain must interpret this new set of binaural cues and construct the phantom images associated with stereo reproduction.

This is why loudspeaker placement, room acoustics and listening position cannot be separated from the recording itself. They participate in the decoding of its spatial information.

A stable recording cannot correct a fundamentally unsuitable room, and a carefully treated room cannot restore information that was never captured or was destroyed during production.

The recording and reproduction environments form a chain.

Natural Reverberation Is More Than a Tail

In a real acoustic environment, reverberation is not an effect that begins after the direct sound has finished.

It is a continuously evolving field of reflections shaped by the dimensions, materials and geometry of the venue.

Those reflections contain:

  • directional asymmetries;
  • different arrival times;
  • frequency-dependent decay;
  • changes in density over time;
  • and relationships to the position and radiation pattern of every instrument.

A reverberation processor can create extraordinarily convincing spaces, and artificial reverberation is indispensable in many forms of production. But a preset does not reproduce the exact interaction that occurred between particular musicians and a particular room at one unrepeatable moment.

When natural reverberation is captured successfully, it does not feel attached to the performance.

It is the performance continuing into the building.

Natural reverberation is architecture becoming sound.

Recording for the Brain

A realistic recording is not necessarily one that captures the largest possible quantity of information.

It is one that preserves the relationships necessary for perception.

For Direct Sound Records, microphone placement, distance, acoustic environment and minimal signal manipulation are therefore not separate technical choices. They form one recording philosophy.

The objective is not purity for its own sake, and it is not nostalgia for a period before digital production.

It is the preservation of continuity.

When a voice is captured from a believable perspective, its centre depends on more than identical level in the two channels. It also depends on stable spectral, temporal and environmental relationships.

When two instruments occupy different sides of an ensemble, their apparent positions should not be understood merely as pan-control settings. They arise from distance, angle, microphone pattern, radiation, reflections and their relationship with the room.

The engineer must therefore decide what should remain coherent, what can be altered and what must be allowed to exist naturally.

Listening Beyond Detail

This perspective also changes how we evaluate recordings and audio systems.

Instead of asking only, “How much detail can I hear?”, we can ask:

  • How credible is the relationship between the sounds?
  • Does the voice remain stable as its pitch and intensity change?
  • Do instruments possess physical body, or are they merely lateral outlines?
  • Does the acoustic environment belong to the performance?
  • Does depth arise from perspective, or only from added reverberation?
  • Does the central image remain convincing at modest listening levels?
  • Do reflections create continuity and air, or do they blur localisation?
  • Does the recording invite prolonged listening, or impress only for a few moments?

These questions move the discussion beyond spectacular sound.

A recording that preserves spatial relationships does not merely demonstrate the capabilities of a system. It gives the system an opportunity to disappear.

When that happens, we stop concentrating on two loudspeakers producing sound.

We begin to perceive musicians, sounding bodies, distance, architecture and silence.

The Space Between Sounds

The most convincing recordings are not necessarily those that contain the most obvious effects, the widest images or the greatest quantity of isolated detail.

They are often those in which every element appears to belong to the same acoustic reality.

The musicians have scale. The centre has physical stability. The room surrounds rather than decorates. Reflections extend the performance instead of obscuring it. Silence defines the distance between one sounding body and another.

This is the space between sounds.

It cannot be reduced to one measurement, one microphone technique or one idea of phase coherence. It emerges from a network of relationships extending from the original performance to the listener’s brain.

When musicians, room, microphone position, recording chain, distribution format, loudspeakers and listening environment align, something rare can occur.

The recording no longer seems to document a performance from the past.

It makes that performance inhabitable again.

Perhaps this is the deepest purpose of high fidelity: not merely to reproduce sounds, but to reconstruct a space in which those sounds can exist.

References and Further Reading

  1. Middlebrooks, J. C. and Green, D. M. “Sound Localization by Human Listeners.”
    Annual Review of Psychology, 1991.
    View publication
  2. Brughera, A., Dunai, L. and Hartmann, W. M. “Human Interaural Time Difference Thresholds for Sine Tones: The High-Frequency Limit.”
    Journal of the Acoustical Society of America, 2013.
    View publication
  3. Bures, Z. and Marsalek, P. “On the Precision of Neural Computation with Interaural Level Differences in the Lateral Superior Olive.”
    Brain Research, 2013.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning Sound Localization with New Ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Pulkki, V. “Microphone Techniques and Directional Quality of Sound Reproduction.”
    Audio Engineering Society, 2002.
    View AES record
  6. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View AES record
  7. Toole, F. E. “Loudspeakers and Rooms for Stereophonic Sound Reproduction.”
    Audio Engineering Society 8th International Conference, 1990.
    View AES record

An earlier version of this essay was published in Audio Review, issue 487, June 2026. This Direct Sound Records Journal edition has been revised, expanded and technically updated for an international readership.

How We Hear in 3D – The Neuroscience Behind Stereo Perception

1024 1024 Michelangelo

“Phase is not the sound — it is the space between sounds.” Taken as a metaphor, this captures something fundamental about human hearing: the brain does not process the signals arriving at our two ears independently. It continuously compares them, using extraordinarily small differences in time, level and spectral balance to reconstruct the position of sound around us.

In high-end audio, we often concentrate on equipment, formats, resolution and frequency response. Yet beneath every recording and every playback system lies a more fundamental question: how does the human auditory system transform two incoming signals into a convincing three-dimensional world?

Stereo reproduction works not because two loudspeakers recreate the original sound field perfectly, but because they provide the auditory system with enough carefully organised information to create a plausible spatial scene. Understanding that process changes how we think about microphone placement, phase relationships, room acoustics and the meaning of realism in recording.

Hearing Is a Reconstruction

Sound arriving at the ears does not contain a ready-made map of the space around us. The auditory system must infer the direction, distance and environment of a sound source from a collection of acoustic clues.

For horizontal localisation, the most important binaural cues are differences in arrival time and sound level between the two ears. For elevation and front-to-back discrimination, the auditory system also relies heavily on direction-dependent spectral filtering created by the listener’s head, torso and outer ears.

These mechanisms operate together. They should not be understood as three entirely separate systems, nor as rigid frequency zones with precise boundaries. Natural sounds are usually broadband, reflections complicate the incoming signals, and the brain integrates multiple cues over time.

The Principal Cues to Spatial Hearing

1. Interaural Time Differences

When a sound source is positioned to one side of the listener, the sound normally reaches the nearer ear slightly before reaching the farther ear. This difference is known as the interaural time difference, or ITD.

The delays involved are extremely small—often measured in microseconds—but the auditory system is remarkably sensitive to them. For low-frequency sounds, neural activity can follow the temporal structure of the waveform closely enough for the brain to compare the timing received at the two ears.

ITD sensitivity to the fine structure of pure tones becomes progressively less effective as frequency rises, with human sensitivity deteriorating sharply around the region of approximately 1.4 to 1.5 kHz. This should be treated as a broad transition rather than a universal dividing line. High-frequency sounds can still carry timing information through changes in their amplitude envelope.

2. Interaural Level Differences

A sound arriving from one side will also tend to be louder at the nearer ear. At shorter wavelengths, the head obstructs part of the sound travelling towards the farther ear, producing an acoustic shadow. The resulting difference in level is known as the interaural level difference, or ILD.

Level differences generally become more pronounced at higher frequencies because shorter wavelengths are more strongly affected by the head. However, ITD and ILD do not simply exchange responsibility at one exact frequency. For complex sounds, the auditory system can combine timing and level information across several frequency regions.

The relative importance of these cues also changes with the sound itself, its distance, the surrounding reflections and the listener’s hearing.

3. Spectral Cues and the Head-Related Transfer Function

Time and level differences are especially useful for identifying whether a sound is located towards the left or right. They are less able, by themselves, to resolve whether a source is above, below, in front of or behind the listener.

For this, the complex shape of the outer ear becomes essential. The folds of the pinna, together with the head and upper body, alter the spectrum of incoming sound in a direction-dependent way. Some frequencies are reinforced, while others are attenuated or notched.

The complete acoustic transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because every person’s anatomy is different, HRTFs are individual. The brain gradually learns the spectral patterns associated with particular directions. Experiments in which the shape of the outer ear was temporarily altered have shown that localisation initially becomes less accurate, particularly for elevation and front-to-back judgements, but can improve again as listeners adapt to the modified cues.

Where Phase Fits In

The word phase is used in several related but distinct ways in audio, which can easily create confusion.

At low frequencies, a difference in arrival time between the ears can also be described as an interaural phase difference for a periodic waveform. In this context, phase difference is one of the ways the auditory system obtains spatial information.

Neurons within the auditory brainstem, including those associated with the medial superior olive, are specialised for processing extremely small interaural timing differences. Their responses contribute to the neural representation of horizontal sound direction.

However, it would be too simple to conclude that phase is used only to locate a sound and never contributes to what that sound is. The timing structure of a waveform—often described as its temporal fine structure—also contributes to aspects of pitch perception, auditory masking and the separation of sounds in complex listening environments.

The more useful distinction for recording engineers is therefore not between “phase” and “sound identity,” but between the different roles that temporal and phase relationships can play.

Within a stereo recording, relationships between the two channels may affect:

  • the apparent position of a phantom image;
  • the perceived width and stability of the soundstage;
  • the impression of depth and surrounding ambience;
  • the result when the recording is reproduced in mono;
  • frequency-response changes caused by constructive and destructive interference.

Phase is therefore neither an isolated technical curiosity nor a universal explanation for every spatial quality. It is one component within a larger system of timing, level, spectrum, reflection and playback interaction.

Why This Matters for Stereo Recording

Stereo microphone techniques create spatial information in different ways.

A spaced pair such as AB introduces arrival-time differences between the microphones and may also produce level differences. A near-coincident configuration such as ORTF deliberately combines microphone spacing with directional level differences.

Coincident techniques such as XY and Mid-Side minimise the arrival-time difference between microphones and create direction mainly through differences in level. The Blumlein pair, using two coincident figure-of-eight microphones, also derives its directional information from the polar patterns and polarity relationships of the two channels while capturing substantial information from the surrounding acoustic environment.

None of these methods is automatically natural or unnatural in every situation. Each encodes the original acoustic event differently, and each interacts differently with loudspeakers, headphones, room reflections and listener position.

This is particularly important because conventional stereo loudspeaker reproduction does not send the left channel exclusively to the left ear or the right channel exclusively to the right ear. Each loudspeaker reaches both ears, introducing additional timing, level and spectral interactions. The listening room then adds its own reflections.

The task of the recording engineer is therefore not merely to create a wide image. It is to create interchannel relationships that remain meaningful when reproduced through the intended playback system.

From Spatial Effect to Spatial Credibility

A stereo recording can sound spectacularly wide while still producing unstable localisation, exaggerated scale or an uncertain centre image. Conversely, a narrower presentation may feel more convincing because its timing, level and reverberant cues form a more internally consistent spatial picture.

This suggests a more useful question than simply asking whether a recording sounds spacious:

Do the spatial cues reproduced by the system support one another strongly enough for the brain to construct a stable and believable acoustic scene?

Phase coherence is part of that question, but it should not be treated as a single measurement that determines realism on its own. Microphone polar pattern, spacing, angle, source distance, direct-to-reverberant ratio, loudspeaker placement and room acoustics all contribute to the final perception.

What Comes Next

In the next article, I will examine how AB, ORTF, XY, Mid-Side and Blumlein recording techniques encode spatial information differently, and why a technique that produces impressive width does not necessarily produce the most credible depth or localisation.

I will also explore the relationship between coincident microphone techniques, loudspeaker reproduction, room acoustics and the preservation of stable interchannel relationships.

The central principle is simple: recording technology should not be considered separately from human perception. The microphone arrangement, recording space and playback format are parts of one perceptual chain.

References and Further Reading

  1. Brughera, A., Dunai, L. and Hartmann, W. M. “Human interaural time difference thresholds for sine tones: the high-frequency limit.” Journal of the Acoustical Society of America, 2013.
    View publication
  2. Wightman, F. L. and Kistler, D. J. “The dominant role of low-frequency interaural time differences in sound localization.”
    Journal of the Acoustical Society of America, 1992.
    View publication
  3. Salminen, N. H., Tiitinen, H., Yrttiaho, S. and May, P. J. C. “The neural code for interaural time difference in human auditory cortex.”
    Journal of the Acoustical Society of America, 2010.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning sound localization with new ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Moore, B. C. J. “The role of temporal fine structure processing in pitch perception, masking, and speech perception for normal-hearing and hearing-impaired people.”
    Journal of the Association for Research in Otolaryngology, 2008.
    View publication
  6. Eargle, J. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View publication

An earlier version of this article was published on LinkedIn. This Direct Sound Records Journal edition has been revised, expanded and technically updated, with additional context and scientific references.