The Space Between Sounds: How the Brain Reconstructs the Soundstage
https://directsoundrecords.com/wp-content/uploads/2026/07/Screenshot-2026-07-25-at-20.32.25-1024x755.png 1024 755 Michelangelo Michelangelo https://secure.gravatar.com/avatar/89c9d8d7e3cacdd3a1530cf562ec5e1b4de2198dbb8a6dd62b550b23ac888854?s=96&d=mm&r=gIn musical realism, we do not listen only to frequencies, timbres and dynamics. We listen to relationships: differences in arrival time, level, reflection and spatial organisation. When those relationships remain credible, a recording can become more than a collection of sounds. It can become an inhabitable acoustic space.
When listening to a high-fidelity system, we often say that a recording “sounds real.” But what does that actually mean?
Frequency extension matters. So do controlled bass, natural midrange, transient response, low distortion and harmonic detail. Yet realism also depends on something less obvious and considerably more fragile: whether the recording and playback system allow us to perceive a believable relationship between sound sources and the space around them.
A voice is not simply a voice. It is a presence located at a particular distance and position.
A piano is not merely a collection of strings, hammers and resonances. It is a physical body occupying a volume of air, projecting energy into a room and interacting with surrounding surfaces.
A quartet is not simply divided into left, centre and right. It is an acoustic event in which musicians, distances, reflections and silence form one spatial relationship.
The human auditory system does not receive this scene passively. It interprets, compares and reconstructs it. In a sense, the brain triangulates.
The Soundstage Is Not an Effect
In high-end audio, words such as soundstage, depth, focus, air and presence are used constantly. They are useful descriptions, but they can become vague unless we remember that they have a perceptual foundation.
The soundstage is not a picture physically stored inside a recording. Nor is it simply drawn between two loudspeakers.
It is a perceptual reconstruction created from the information reaching the listener’s ears.
The brain interprets several cues simultaneously:
- tiny differences in the time at which sound reaches each ear;
- differences in level between the ears;
- direction-dependent changes in the spectrum;
- the relationship between direct and reflected sound;
- the evolution of reflections and reverberation over time;
- and prior knowledge of familiar voices, instruments and environments.
When these cues support one another, a sound source can appear stable, physical and separate from the loudspeakers.
When they conflict, a recording may still sound impressive, detailed or extremely wide, but the scene can feel unstable or artificial.
The soundstage is therefore not merely an effect.
It is information interpreted by the listener.
How the Brain Reads Acoustic Space
For horizontal localisation, the auditory system relies heavily on differences in timing and level between the two ears.
Interaural Time Differences
When a sound arrives from one side, it normally reaches the nearer ear slightly before the farther ear. These interaural time differences, or ITDs, can be extremely small—sometimes measured in tens of microseconds—yet the auditory system is remarkably sensitive to them.
Timing cues are particularly important at lower frequencies, where neural activity can represent the temporal fine structure of the waveform with sufficient precision. Structures within the auditory brainstem, especially the medial superior olive, contribute to the analysis of these binaural timing relationships.
ITD sensitivity does not end at one perfectly defined frequency, but sensitivity to the fine structure of pure tones deteriorates substantially through the region around 1 to 1.5 kHz. Higher-frequency sounds can still convey timing information through changes in their amplitude envelopes.
Interaural Level Differences
At shorter wavelengths, the head creates a more substantial acoustic shadow. A sound arriving from the left will generally produce a higher level at the left ear than at the right.
This interaural level difference, or ILD, becomes an increasingly useful directional cue as frequency rises. Neural circuits involving the lateral superior olive contribute to its early processing.
In natural listening, timing and level differences do not operate as two isolated systems with a rigid boundary between them. Broadband sounds contain multiple cues, and the brain combines them according to frequency, source position, environment and reliability.
Spectral Cues and the Shape of the Listener
Timing and level differences provide strong information about left and right, but they cannot always distinguish whether a sound is above, below, in front of or behind the listener.
For those dimensions, the shape of the outer ears, head and torso becomes essential.
The folds of the pinnae filter incoming sound differently according to direction. Certain frequencies are reinforced, while others are attenuated or notched. The complete transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.
Because human anatomy varies, each person’s HRTF is individual. Over time, the brain learns the correspondence between these spectral patterns and positions in space.
Experiments in which the outer ears were temporarily reshaped have shown that vertical localisation initially becomes much less accurate. With experience, listeners can learn to interpret the altered cues—evidence that spatial hearing is not merely mechanical, but calibrated and adaptive.
What we perceive is therefore never only the sound emitted by the source. It is the result of a relationship between source, environment, listener and brain.
Where Phase Enters the Picture
In audio engineering, phase is often discussed as a problem.
Signals may be described as “out of phase” when they weaken the centre image, reduce low-frequency energy, produce cancellations or behave unpredictably in mono.
These concerns are real, but they represent only part of the subject.
Phase is not simply an error waiting to be corrected. Phase and timing relationships can also carry important information about position, width, depth and the interaction between sound sources.
For a periodic low-frequency signal, a difference in arrival time between the ears can also be expressed as an interaural phase difference. In this sense, phase contributes directly to spatial localisation.
Within a stereo recording, interchannel timing and phase relationships can influence:
- the position and stability of phantom images;
- the apparent width of the presentation;
- the sense of distance and depth;
- the relationship between direct sound and ambience;
- mono compatibility;
- and frequency-dependent reinforcement or cancellation when signals combine.
However, it would be misleading to say that phase is used only for location and never contributes to the perceived character of a sound. Temporal fine structure is also involved in pitch, masking and the separation of simultaneous sources. In recording and reproduction, phase relationships can additionally alter the spectrum whenever correlated signals combine acoustically or electrically.
The useful distinction is therefore not between phase and sound, but between the different roles that temporal relationships perform.
Phase is not the sound itself, but it can help organise the space in which that sound is perceived.
Coherence Does Not Mean Perfection
The expression phase coherence is frequently used as though it described one measurable quality that a recording either possesses or lacks.
Reality is more complicated.
Every acoustic environment contains delays. Reflections arrive after the direct sound and from different directions. Instruments radiate differently according to frequency. Microphones have frequency-dependent polar patterns. Loudspeakers and rooms introduce further interactions.
A natural acoustic event is not phase-identical at every point in space.
Spatial coherence should therefore not mean eliminating every difference or delay. It means preserving relationships that remain compatible enough for the auditory system to interpret them as belonging to one plausible event.
When this happens, the voice can stabilise at the centre without appearing glued to either speaker. Instruments occupy a readable volume rather than appearing as thin lateral points. The room does not feel like reverberation placed behind the music; it surrounds and continues the performance.
The silence between instruments stops being empty.
It becomes air.
Why Some Recordings Sound Large but Not Real
A wide soundstage is not necessarily a realistic soundstage.
Modern production provides an enormous range of tools for creating size: multiple microphones, pan controls, delay, artificial reverberation, stereo widening, decorrelation and Mid-Side processing.
These tools can be artistically valuable. They can also create a guitar broader than its physical source, a voice floating beyond the loudspeakers or a reverberant field that could never have existed around the original performers.
There is nothing inherently wrong with that. Recording is also an art of construction.
But width and credibility are not synonymous.
The auditory system evaluates more than the apparent size of the scene. It also evaluates whether the timing, spectral, directional and environmental information is mutually plausible.
A very wide image may be initially impressive. Yet if the centre lacks stability, the reverberation does not belong to the sources or the spatial relationships change unnaturally with frequency, the illusion becomes less convincing.
It is similar to viewing a photograph with intense colour and extraordinary sharpness but incorrect perspective. The image attracts attention immediately, yet something feels wrong.
A recording can contain remarkable detail and separation while remaining spatially two-dimensional.
Beautiful, perhaps.
But not alive.
Every Microphone Technique Makes a Decision
For acoustic music, the choice of microphone technique is never neutral.
AB, ORTF, XY, Mid-Side and Blumlein do not simply create different varieties of stereo width. They encode different combinations of timing, level, polarity, direction and room information.
A spaced AB pair introduces meaningful arrival-time differences between microphones and can create scale, openness and envelopment.
ORTF combines a moderate physical separation with directional cardioid microphones, creating both interchannel time and level differences.
Coincident systems such as XY, Mid-Side and Blumlein minimise the timing difference introduced by microphone spacing and derive direction primarily through level and polarity relationships.
These differences affect localisation, spaciousness, mono compatibility and the way the recording interacts with loudspeaker reproduction.
Coincident techniques can produce comparatively stable and clearly located virtual sources. Spaced techniques may create broader or more diffuse images and can convey strong spaciousness. Neither outcome is automatically better.
The appropriate technique depends on:
- the musicians and their physical arrangement;
- the acoustic character of the venue;
- the desired listening perspective;
- the balance between localisation and envelopment;
- the intended distribution format;
- and the expected playback environment.
There is no microphone technique that is universally correct.
There is only a technique whose compromises are more or less coherent with the intended result.
Recording Is a Translation
A microphone does not hear like a human being.
It has no head, no outer ears, no perceptual memory and no awareness of the room. It does not compare what it captures with years of experience. It measures sound pressure or pressure gradient according to its physical construction and polar pattern.
Human perception, by contrast, is active.
This means that stereo recording is always a translation.
The task is not simply to place two microphones in front of a performance and assume that the original space has been preserved. The task is to create two signals that, when reproduced through loudspeakers in another room, provide the listener with enough coherent information to reconstruct a plausible scene.
That distinction is fundamental.
The original venue, the microphone array, the recording chain, the loudspeakers, the listening room and the listener are all parts of one perceptual system.
A recording can never transport the original acoustic field intact. It selects, encodes and later stimulates a new reconstruction.
The Listening Room Is Part of the Reproduction
With headphones, each channel is delivered predominantly to one ear. With conventional loudspeakers, both speakers reach both ears.
The left ear receives sound from the left loudspeaker, sound from the right loudspeaker after a different path, and reflections from the listening room. The right ear receives the corresponding combination from the opposite side.
The listener’s brain must interpret this new set of binaural cues and construct the phantom images associated with stereo reproduction.
This is why loudspeaker placement, room acoustics and listening position cannot be separated from the recording itself. They participate in the decoding of its spatial information.
A stable recording cannot correct a fundamentally unsuitable room, and a carefully treated room cannot restore information that was never captured or was destroyed during production.
The recording and reproduction environments form a chain.
Natural Reverberation Is More Than a Tail
In a real acoustic environment, reverberation is not an effect that begins after the direct sound has finished.
It is a continuously evolving field of reflections shaped by the dimensions, materials and geometry of the venue.
Those reflections contain:
- directional asymmetries;
- different arrival times;
- frequency-dependent decay;
- changes in density over time;
- and relationships to the position and radiation pattern of every instrument.
A reverberation processor can create extraordinarily convincing spaces, and artificial reverberation is indispensable in many forms of production. But a preset does not reproduce the exact interaction that occurred between particular musicians and a particular room at one unrepeatable moment.
When natural reverberation is captured successfully, it does not feel attached to the performance.
It is the performance continuing into the building.
Natural reverberation is architecture becoming sound.
Recording for the Brain
A realistic recording is not necessarily one that captures the largest possible quantity of information.
It is one that preserves the relationships necessary for perception.
For Direct Sound Records, microphone placement, distance, acoustic environment and minimal signal manipulation are therefore not separate technical choices. They form one recording philosophy.
The objective is not purity for its own sake, and it is not nostalgia for a period before digital production.
It is the preservation of continuity.
When a voice is captured from a believable perspective, its centre depends on more than identical level in the two channels. It also depends on stable spectral, temporal and environmental relationships.
When two instruments occupy different sides of an ensemble, their apparent positions should not be understood merely as pan-control settings. They arise from distance, angle, microphone pattern, radiation, reflections and their relationship with the room.
The engineer must therefore decide what should remain coherent, what can be altered and what must be allowed to exist naturally.
Listening Beyond Detail
This perspective also changes how we evaluate recordings and audio systems.
Instead of asking only, “How much detail can I hear?”, we can ask:
- How credible is the relationship between the sounds?
- Does the voice remain stable as its pitch and intensity change?
- Do instruments possess physical body, or are they merely lateral outlines?
- Does the acoustic environment belong to the performance?
- Does depth arise from perspective, or only from added reverberation?
- Does the central image remain convincing at modest listening levels?
- Do reflections create continuity and air, or do they blur localisation?
- Does the recording invite prolonged listening, or impress only for a few moments?
These questions move the discussion beyond spectacular sound.
A recording that preserves spatial relationships does not merely demonstrate the capabilities of a system. It gives the system an opportunity to disappear.
When that happens, we stop concentrating on two loudspeakers producing sound.
We begin to perceive musicians, sounding bodies, distance, architecture and silence.
The Space Between Sounds
The most convincing recordings are not necessarily those that contain the most obvious effects, the widest images or the greatest quantity of isolated detail.
They are often those in which every element appears to belong to the same acoustic reality.
The musicians have scale. The centre has physical stability. The room surrounds rather than decorates. Reflections extend the performance instead of obscuring it. Silence defines the distance between one sounding body and another.
This is the space between sounds.
It cannot be reduced to one measurement, one microphone technique or one idea of phase coherence. It emerges from a network of relationships extending from the original performance to the listener’s brain.
When musicians, room, microphone position, recording chain, distribution format, loudspeakers and listening environment align, something rare can occur.
The recording no longer seems to document a performance from the past.
It makes that performance inhabitable again.
Perhaps this is the deepest purpose of high fidelity: not merely to reproduce sounds, but to reconstruct a space in which those sounds can exist.
References and Further Reading
- Middlebrooks, J. C. and Green, D. M. “Sound Localization by Human Listeners.”
Annual Review of Psychology, 1991.
View publication - Brughera, A., Dunai, L. and Hartmann, W. M. “Human Interaural Time Difference Thresholds for Sine Tones: The High-Frequency Limit.”
Journal of the Acoustical Society of America, 2013.
View publication - Bures, Z. and Marsalek, P. “On the Precision of Neural Computation with Interaural Level Differences in the Lateral Superior Olive.”
Brain Research, 2013.
View publication - Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning Sound Localization with New Ears.”
Nature Neuroscience, 1998.
View publication - Pulkki, V. “Microphone Techniques and Directional Quality of Sound Reproduction.”
Audio Engineering Society, 2002.
View AES record - Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
Journal of the Audio Engineering Society, 1985.
View AES record - Toole, F. E. “Loudspeakers and Rooms for Stereophonic Sound Reproduction.”
Audio Engineering Society 8th International Conference, 1990.
View AES record
An earlier version of this essay was published in Audio Review, issue 487, June 2026. This Direct Sound Records Journal edition has been revised, expanded and technically updated for an international readership.




