Posts Tagged :

high fidelity

Where High Fidelity Really Begins: Four Decisions That Shape a Believable Stereo Recording

1024 582 Michelangelo

High fidelity begins before mastering, before the distribution format and before the physical or digital medium. It begins with the decisions made when the performance is captured: how stereo space is encoded, how phase relationships are managed, how far the microphones are placed from the musicians and which microphones are chosen to translate that perspective.

Listening to a stereo recording involves much more than hearing sound emerge from two loudspeakers.

A voice apparently positioned in the centre does not come from a hidden central speaker. An instrument heard slightly to the left is not physically present at that point in the listening room. The apparent depth extending beyond the wall behind the system is not an acoustic space that has suddenly opened in front of us.

These are perceptual constructions.

The auditory system interprets differences in level, timing, spectrum and phase, together with reflections and the relationship between direct and reverberant sound. From this information, it creates a plausible auditory scene.

Stereo recording is therefore not merely the process of producing two channels. It is the art of creating relationships that the listener’s brain can interpret as space.

A High-End System Cannot Invent the Recording

A recording may be highly detailed, dynamically impressive and extended at both ends of the frequency spectrum while still failing to sound believable.

It may reveal the smallest movements of the musicians but construct no stable image. It may sound extremely wide but imprecise. It may attract attention immediately and gradually become artificial or tiring.

For listeners using loudspeakers, this distinction is fundamental.

A revealing system can expose the spatial quality of a recording, but it cannot manufacture coherence that was never captured. If the central image is unstable, the playback equipment may reveal that instability with greater clarity. If the relationship between instruments and ambience is implausible, additional resolution may make the contradiction more obvious.

The format and reproduction chain matter enormously, but they inherit the decisions made at the beginning.

High fidelity starts at the microphone.

Stereo Is Not Simply Width

In audiophile language, we often speak about soundstage, imaging, depth, focus and air around instruments.

These are useful descriptions, but a believable soundstage is not simply a wide one.

Width can be created in many ways:

  • spaced microphones;
  • pan controls;
  • interchannel delays;
  • Mid-Side processing;
  • stereo reverberation;
  • decorrelation;
  • and dedicated spatial processors.

All of these can be artistically valid. But increasing width does not automatically increase realism.

A believable voice should remain stable in the centre, not only in lateral position but also in physical presence. A piano should have plausible dimensions. A guitar should not become wider than the performer playing it unless that enlargement is an intentional artistic decision. A flute should not be reduced to breath, lips and key noise while losing the integrated sound of the instrument.

The acoustic environment should not feel as though it has been placed behind the musicians as an independent effect. It should appear to belong to the same event.

When a recording is coherent, the listener perceives more than individual sources distributed between two loudspeakers. The listener perceives relationships between the instruments, their apparent distances and the room surrounding them.

That is what makes a soundstage credible.

The Four Decisions Behind a Believable Recording

Four closely connected decisions have a particularly strong influence on spatial realism:

  1. the stereo microphone technique;
  2. the timing and phase relationships between channels;
  3. the distance between microphones and performers;
  4. and the microphone characteristics used at that distance.

None of these choices operates independently.

Changing microphone distance changes the balance between direct sound and room ambience. It also changes proximity effect, source integration, off-axis contribution and the amount of environmental noise captured.

Changing the microphone changes the polar pattern, tonal balance, transient behaviour, self-noise and off-axis response experienced at that distance.

Changing the stereo array changes the interchannel timing and level relationships supplied to the reproduction system.

The art lies in making these decisions support the same perceptual objective.

1. Stereo Technique: Different Ways of Encoding Space

Stereo microphone techniques are not simply alternative arrangements for producing a left and right channel. They encode spatial information in fundamentally different ways.

Spaced Pairs

In an AB arrangement, two microphones are separated physically. A sound arriving from one side normally reaches one microphone before the other, producing an interchannel time difference. Depending on the source, microphone pattern and geometry, differences in level may also occur.

Spaced arrays can create a broad impression of scale and envelopment. In a good concert hall, and particularly with larger ensembles, this can communicate openness, low-frequency spaciousness and the feeling that the performance breathes within a large acoustic environment.

Physical spacing also introduces frequency-dependent phase relationships between the channels. In stereo these can contribute to spaciousness, but they may reduce localisation precision or create tonal changes when the channels are summed or partially combined.

This does not make AB inherently defective. It means that width, envelopment, localisation and mono compatibility must be balanced deliberately.

Near-Coincident Techniques

Near-coincident arrangements such as ORTF, NOS and DIN combine physical microphone spacing with directional polar patterns.

ORTF, for example, uses two cardioid microphones separated by 17 centimetres and angled 110 degrees apart. The resulting stereo image contains both interchannel timing and level differences.

This can offer a productive compromise: greater spaciousness than many coincident cardioid arrangements, together with more definite image positioning than a widely spaced pair.

The timing component is not an accidental defect. It is part of the intended spatial design.

Coincident Techniques

In coincident arrangements such as XY, Mid-Side and Blumlein, the microphone capsules are placed as close as physically possible to the same acoustic point.

Because the direct sound reaches both capsules at approximately the same time, directional information is created primarily through differences in level and polarity produced by the microphones’ polar patterns.

Coincident geometry generally offers:

  • stable localisation;
  • a clearly defined central image;
  • predictable mono compatibility;
  • and fewer time-delay interactions introduced by microphone spacing.

The soundstage can sometimes appear less expansive than that produced by a spaced array, but individual positions may be easier to read.

Practical coincidence is never perfect. The capsules have physical dimensions, and real microphones exhibit frequency-dependent polar and phase responses. Careful construction and positioning still matter.

No Technique Is Universally Superior

The appropriate technique depends on:

  • the size and arrangement of the ensemble;
  • the acoustic character of the venue;
  • the required balance between localisation and spaciousness;
  • the importance of mono compatibility;
  • the intended listening perspective;
  • and the expected playback system.

When scale and strong hall envelopment are priorities, a spaced arrangement may be highly effective. When image stability and coincident timing are central to the project, XY, Mid-Side or Blumlein may offer important advantages.

There is no universally correct technique.

There is only a technique whose characteristics are coherent with the recording’s purpose.

2. Phase: More Than a Technical Problem

In audio, phase is frequently discussed only when something has gone wrong.

We notice it when bass becomes thin, when a centre image loses solidity, when combining microphones produces tonal colouration or when a stereo recording behaves poorly in mono.

These are genuine problems, but phase is also part of spatial information.

Timing and phase relationships between channels can contribute to the perception of width, localisation, ambience and depth. However, the differences captured by two microphones are not identical to the binaural cues produced at two human ears.

Microphones have no head between them, no pinnae and no torso. They do not apply the listener-specific spectral transformations associated with natural localisation. Their spacing and polar patterns create a new encoding intended for reproduction through another system.

This distinction is crucial.

In conventional loudspeaker stereo, each ear hears both speakers. The listening room adds reflections, and the listener’s head modifies the signals again. The brain must interpret this combined information and construct phantom images.

A large interchannel delay can increase spaciousness without necessarily improving image precision. A coincident array reduces the timing differences introduced at capture, but it cannot eliminate every phase interaction in the room, microphone, loudspeaker or recording chain.

The meaningful objective is not perfect phase identity.

It is maintaining interchannel relationships that remain sufficiently consistent for the listener to construct a stable scene.

Phase and the Centre Image

A central phantom image is created when the loudspeakers provide the auditory system with compatible information suggesting that a source lies between them.

Equal level alone is not always sufficient to make that centre feel physical. The spectral and temporal content of the two channels must also support the same perceptual interpretation.

If multiple microphones capture one source with different delays, the resulting interference may change with frequency. The image can become less stable, and tonal character may vary according to the combination of channels and listening position.

This is one reason why microphone count should not be confused with information quality.

More microphones can offer flexibility, control and creative possibilities. They can also create more relationships that must be managed.

3. Microphone Distance: Where Proportion Begins

Microphone distance is fundamental to whether an instrument sounds credible through loudspeakers.

A close position can provide immediacy, clarity and extraordinary detail. It can reveal the movement of piano mechanics, a flautist’s breathing, fingers touching strings, valve noise, bow texture and the precise attack of every note.

These details can be fascinating and musically valuable.

But they are not always proportionate to the way the instrument would be heard from a natural listening position.

The danger is that detail becomes confused with realism.

The Instrument Must Reassemble

Acoustic instruments often radiate different frequency regions from different parts of their bodies and in different directions.

At very close range, a microphone hears one local part of that radiation field. Moving farther away allows those contributions to integrate more fully before reaching the capsule.

The piano can become one sounding body rather than a collection of strings, hammers and mechanical events. The guitar returns to plausible physical dimensions. The flute becomes more than the excitation point at the mouthpiece; it becomes an instrument projecting energy into the room.

Distance allows the instrument to reassemble itself.

Distance Introduces the Room

Moving a microphone away also changes the balance between direct and reflected sound.

More of the room enters the recording. Early reflections influence tone and localisation. Reverberation communicates scale and distance. Background noise and undesirable acoustic characteristics become harder to avoid.

The correct distance is therefore not simply the most natural one in theory. It is the distance at which source integration, clarity, instrumental body and room contribution reach the desired balance.

That point changes with every venue, ensemble and microphone.

Close Is Not Wrong

Close microphone placement should not be treated as inherently artificial.

Many musical genres depend on intimacy, isolation, impact or the ability to balance sources independently. A close perspective may be exactly right for the artistic language of the production.

The problem arises only when a local, magnified perspective is presented as though it were automatically more faithful because it reveals more detail.

Detail describes how much can be perceived.

Realism describes whether those details belong to a plausible whole.

4. Microphone Choice: An Instrument of Proportion

At a realistic recording distance, the microphone is not merely a transparent transducer.

It becomes an instrument of proportion.

Every microphone has a technical personality shaped by factors including:

  • its operating principle;
  • diaphragm or ribbon construction;
  • polar pattern;
  • frequency and phase response;
  • off-axis behaviour;
  • transient response;
  • self-noise and sensitivity;
  • grille and body geometry;
  • electronics and transformers;
  • and manufacturing tolerances.

No microphone is perfectly neutral under every condition.

A microphone that sounds impressive at close range may not be the right choice at several metres. A presence rise that creates attractive clarity nearby may make a distant recording feel thin or overly explicit. A microphone with excellent on-axis response but irregular off-axis behaviour may colour the room contribution as distance increases.

Conversely, a microphone whose polar pattern and tonal balance remain well controlled away from the axis may integrate the direct and reverberant fields more convincingly.

Ribbon and Condenser Microphones

The choice should not be reduced to a contest between ribbon and condenser technology.

Modern condenser microphones can provide extended bandwidth, low self-noise, high sensitivity and carefully controlled directional behaviour. These qualities can be invaluable for distant acoustic recording.

Some ribbon microphones offer a different balance: smooth high-frequency behaviour, figure-of-eight directivity and a substantial sense of instrumental body. These qualities can suit coincident Blumlein recording and certain natural-distance applications particularly well.

But the result depends on the individual microphone, not merely the category printed on its specification sheet.

A ribbon is not automatically warm or natural. A condenser is not automatically bright or analytical.

The meaningful question is:

Does this microphone, at this distance and in this room, preserve the proportions required by the music?

Proximity Effect and Tonal Perspective

Directional pressure-gradient microphones exhibit proximity effect: their low-frequency response increases as the source moves closer.

The effect is generally strongest with figure-of-eight patterns and is also present, to a lesser degree, with cardioid and related directional patterns. Pure pressure-operated omnidirectional microphones do not exhibit conventional proximity effect.

At very close distances, proximity effect may exaggerate bass and make a source appear larger than its natural scale. It can also be used creatively to provide weight, intimacy or authority.

As the microphone moves farther from the source, this low-frequency boost diminishes. The resulting change in tonal balance must be considered alongside the growing contribution of the room.

This is one reason why microphone choice and distance cannot be separated. The correct working distance is not established by geometry alone; it must also produce an appropriate tonal foundation.

The objective is not to enlarge the bass artificially. It is to preserve enough body for the instrument to remain physical without making it implausibly large.

Why Blumlein Deserves Particular Attention

Among coincident techniques, the Blumlein pair occupies a distinctive position.

It uses two figure-of-eight microphones mounted coincidently and angled 90 degrees apart. Direction is encoded primarily through level and polarity differences, while the rear lobes capture a substantial part of the surrounding acoustic environment.

This combination can provide:

  • a stable central image;
  • clearly organised lateral localisation;
  • strong mono compatibility;
  • and a naturally integrated representation of the room.

Blumlein is also demanding.

The room must contribute positively because sound arriving from behind the array is captured strongly. Placement is critical. The ensemble must balance acoustically, and the useful recording angle must suit its arrangement.

A poor room is revealed rather than concealed. An incorrect microphone position can produce too much reverberation, an inappropriate stereo spread or an imbalanced ensemble.

That apparent limitation is also part of the technique’s value.

A minimally manipulative method encourages the main problems to be solved before recording begins rather than postponed until post-production.

The Room Is Not an Effect Added Afterwards

At natural microphone distances, the room inevitably becomes part of the recording.

Not every space deserves that responsibility.

A room may be excessively dry, too small, mechanically noisy, confused in the lower frequencies or dominated by unattractive early reflections. When the room is unsuitable, a purist approach does not transform it into a virtue.

But when the acoustic environment supports the music, natural ambience can provide an unusually coherent relationship between source and space.

The early reflections, reverberant build-up, asymmetries, frequency-dependent decay and interaction with instrumental radiation all belong to one event.

They are not a separate layer added later.

They are part of the way the music happened.

Artificial reverberation can be beautiful, realistic and artistically indispensable. The distinction is not between legitimate natural sound and illegitimate processing.

The distinction is whether the reverberant information supports the same perspective and spatial logic as the direct sound.

When natural reverberation is captured successfully, the room does not sit behind the instruments.

It connects them.

What the Listener Can Evaluate

These ideas are relevant not only to recording engineers. They can also change how an audiophile evaluates recordings and equipment.

Instead of asking only how much detail is audible, how deep the bass extends or how wide the soundstage appears, listen for relationships.

Listen to the Centre

Does a central voice feel stable and physical, or merely like a thin point suspended between the speakers?

Does its position and body remain coherent as pitch, intensity and register change?

Listen to Instrumental Proportion

Do the instruments possess believable dimensions?

Does the piano appear as one body, or as a series of enlarged local details? Does a guitar occupy plausible space? Does a flute retain tone and projection rather than becoming mostly breath and mechanism?

Listen to Complexity

Does the soundstage remain intelligible when the musical texture becomes dense?

A recording may appear sharply separated during a simple passage and lose all spatial organisation when multiple instruments play simultaneously.

Listen to Decay

Do notes decay continuously into the same environment in which they began?

Does a piano chord dissolve naturally into the surrounding space? Does the room respond to the music, or does the reverberation seem to operate as a separate layer?

Listen at Moderate Volume

A spatially coherent recording often remains intelligible without being played loudly. The centre continues to exist, instrumental positions remain readable and the acoustic environment still suggests depth.

Exaggerated spectral balance and artificial spatial effects may depend more strongly on level to remain impressive.

Listen for Credibility, Not Size

A realistic instrument does not need to sound enormous.

It needs to sound plausible.

Many recordings impress by enlarging everything. Enlargement can be artistically exciting, but it is not automatically high fidelity.

The Format Preserves; It Does Not Create

Audiophile discussions often concentrate on formats: vinyl, analogue tape, PCM, DSD and high-resolution distribution.

These discussions matter, but the format is sometimes treated as though it were the original source of realism.

A format can preserve captured information with greater or lesser accuracy. It cannot create spatial information that was never recorded.

If the soundstage is unstable, the medium preserves that instability. If the centre is weak, a particular format may alter the subjective presentation but cannot reconstruct the original microphone geometry. If the ambience bears no coherent relationship to the instruments, greater resolution may simply reveal that separation more clearly.

The format matters—but it comes later.

This does not diminish the value of excellent recording and distribution formats. It clarifies their purpose.

A high-quality medium is most meaningful when it is preserving something worth preserving:

  • musical dynamics;
  • instrumental timbre;
  • temporal relationships;
  • spatial information;
  • natural decay;
  • and believable proportion.

High Fidelity as Coherence

Stereo technique determines how spatial cues are encoded.

Phase and timing relationships influence image stability, width and the behaviour of combined signals.

Microphone distance determines the proportion between source, detail and environment.

Microphone choice determines how that distance is translated tonally and spatially.

The room either contributes meaningfully to the performance or becomes another problem to manage.

The recording format preserves these decisions but cannot replace them.

For Direct Sound Records, this is the foundation of natural acoustic recording. The objective is not to impose a spectacular soundstage but to preserve enough coherent information for the listener to reconstruct one.

A great recording does not necessarily astonish immediately.

It often becomes more convincing over time.

It does not enlarge every instrument. It preserves proportion. It does not use reverberation merely as decoration. It allows the acoustic environment to participate in the music.

Two loudspeakers can accomplish something extraordinary: they can suggest the presence of a space that does not physically exist in front of the listener.

For that illusion to succeed, the recording must contain believable information. It must respect the way we hear and provide the brain not only with individual sounds, but with meaningful relationships between them.

Perhaps this is where a recording becomes genuinely high fidelity—not when it reveals everything, but when it allows us to believe what we are hearing.

Related Articles

  • The Space Between Sounds: How the Brain Reconstructs the Soundstage
  • Why the Blumlein Pair Can Excel in Loudspeaker Playback

References and Further Reading

  1. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1986.
    View AES record
  2. Gerzon, M. A. “The Design of Precisely Coincident Microphone Arrays for Stereo and Surround Sound.”
    Audio Engineering Society 50th Convention, 1975.
    View AES record
  3. Politis, A., Laitinen, M.-V., Ahonen, J. and Pulkki, V.
    “Parametric Spatial Audio Processing of Spaced Microphone Array Recordings for Multichannel Reproduction.”
    Journal of the Audio Engineering Society, 2015.
    View AES record
  4. DPA Microphones. “Stereo Recording Techniques and Setups.”
    Read technical guide
  5. DPA Microphones. “ORTF.”
    View technical definition
  6. Neumann. “What Is the Proximity Effect?”
    Read technical guide
  7. Bock, T. M. and Keele, D. B.
    “The Effects of Interaural Crosstalk on Stereo Reproduction and Minimizing Interaural Crosstalk in Nearfield Monitoring.”
    Audio Engineering Society 81st Convention, 1986.
    View AES record
  8. Schneider, M. “MS Mastering of Stereo Microphone Signals.”
    Audio Engineering Society 132nd Convention, 2012.
    View AES record

An earlier version of this article was published in Audio Review, July/August 2026, and was subsequently adapted for LinkedIn. This Direct Sound Records Journal edition has been substantially revised, expanded and technically updated.

Space and Sound

The Space Between Sounds: How the Brain Reconstructs the Soundstage

1024 755 Michelangelo

In musical realism, we do not listen only to frequencies, timbres and dynamics. We listen to relationships: differences in arrival time, level, reflection and spatial organisation. When those relationships remain credible, a recording can become more than a collection of sounds. It can become an inhabitable acoustic space.

When listening to a high-fidelity system, we often say that a recording “sounds real.” But what does that actually mean?

Frequency extension matters. So do controlled bass, natural midrange, transient response, low distortion and harmonic detail. Yet realism also depends on something less obvious and considerably more fragile: whether the recording and playback system allow us to perceive a believable relationship between sound sources and the space around them.

A voice is not simply a voice. It is a presence located at a particular distance and position.

A piano is not merely a collection of strings, hammers and resonances. It is a physical body occupying a volume of air, projecting energy into a room and interacting with surrounding surfaces.

A quartet is not simply divided into left, centre and right. It is an acoustic event in which musicians, distances, reflections and silence form one spatial relationship.

The human auditory system does not receive this scene passively. It interprets, compares and reconstructs it. In a sense, the brain triangulates.

The Soundstage Is Not an Effect

In high-end audio, words such as soundstage, depth, focus, air and presence are used constantly. They are useful descriptions, but they can become vague unless we remember that they have a perceptual foundation.

The soundstage is not a picture physically stored inside a recording. Nor is it simply drawn between two loudspeakers.

It is a perceptual reconstruction created from the information reaching the listener’s ears.

The brain interprets several cues simultaneously:

  • tiny differences in the time at which sound reaches each ear;
  • differences in level between the ears;
  • direction-dependent changes in the spectrum;
  • the relationship between direct and reflected sound;
  • the evolution of reflections and reverberation over time;
  • and prior knowledge of familiar voices, instruments and environments.

When these cues support one another, a sound source can appear stable, physical and separate from the loudspeakers.

When they conflict, a recording may still sound impressive, detailed or extremely wide, but the scene can feel unstable or artificial.

The soundstage is therefore not merely an effect.

It is information interpreted by the listener.

How the Brain Reads Acoustic Space

For horizontal localisation, the auditory system relies heavily on differences in timing and level between the two ears.

Interaural Time Differences

When a sound arrives from one side, it normally reaches the nearer ear slightly before the farther ear. These interaural time differences, or ITDs, can be extremely small—sometimes measured in tens of microseconds—yet the auditory system is remarkably sensitive to them.

Timing cues are particularly important at lower frequencies, where neural activity can represent the temporal fine structure of the waveform with sufficient precision. Structures within the auditory brainstem, especially the medial superior olive, contribute to the analysis of these binaural timing relationships.

ITD sensitivity does not end at one perfectly defined frequency, but sensitivity to the fine structure of pure tones deteriorates substantially through the region around 1 to 1.5 kHz. Higher-frequency sounds can still convey timing information through changes in their amplitude envelopes.

Interaural Level Differences

At shorter wavelengths, the head creates a more substantial acoustic shadow. A sound arriving from the left will generally produce a higher level at the left ear than at the right.

This interaural level difference, or ILD, becomes an increasingly useful directional cue as frequency rises. Neural circuits involving the lateral superior olive contribute to its early processing.

In natural listening, timing and level differences do not operate as two isolated systems with a rigid boundary between them. Broadband sounds contain multiple cues, and the brain combines them according to frequency, source position, environment and reliability.

Spectral Cues and the Shape of the Listener

Timing and level differences provide strong information about left and right, but they cannot always distinguish whether a sound is above, below, in front of or behind the listener.

For those dimensions, the shape of the outer ears, head and torso becomes essential.

The folds of the pinnae filter incoming sound differently according to direction. Certain frequencies are reinforced, while others are attenuated or notched. The complete transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because human anatomy varies, each person’s HRTF is individual. Over time, the brain learns the correspondence between these spectral patterns and positions in space.

Experiments in which the outer ears were temporarily reshaped have shown that vertical localisation initially becomes much less accurate. With experience, listeners can learn to interpret the altered cues—evidence that spatial hearing is not merely mechanical, but calibrated and adaptive.

What we perceive is therefore never only the sound emitted by the source. It is the result of a relationship between source, environment, listener and brain.

Where Phase Enters the Picture

In audio engineering, phase is often discussed as a problem.

Signals may be described as “out of phase” when they weaken the centre image, reduce low-frequency energy, produce cancellations or behave unpredictably in mono.

These concerns are real, but they represent only part of the subject.

Phase is not simply an error waiting to be corrected. Phase and timing relationships can also carry important information about position, width, depth and the interaction between sound sources.

For a periodic low-frequency signal, a difference in arrival time between the ears can also be expressed as an interaural phase difference. In this sense, phase contributes directly to spatial localisation.

Within a stereo recording, interchannel timing and phase relationships can influence:

  • the position and stability of phantom images;
  • the apparent width of the presentation;
  • the sense of distance and depth;
  • the relationship between direct sound and ambience;
  • mono compatibility;
  • and frequency-dependent reinforcement or cancellation when signals combine.

However, it would be misleading to say that phase is used only for location and never contributes to the perceived character of a sound. Temporal fine structure is also involved in pitch, masking and the separation of simultaneous sources. In recording and reproduction, phase relationships can additionally alter the spectrum whenever correlated signals combine acoustically or electrically.

The useful distinction is therefore not between phase and sound, but between the different roles that temporal relationships perform.

Phase is not the sound itself, but it can help organise the space in which that sound is perceived.

Coherence Does Not Mean Perfection

The expression phase coherence is frequently used as though it described one measurable quality that a recording either possesses or lacks.

Reality is more complicated.

Every acoustic environment contains delays. Reflections arrive after the direct sound and from different directions. Instruments radiate differently according to frequency. Microphones have frequency-dependent polar patterns. Loudspeakers and rooms introduce further interactions.

A natural acoustic event is not phase-identical at every point in space.

Spatial coherence should therefore not mean eliminating every difference or delay. It means preserving relationships that remain compatible enough for the auditory system to interpret them as belonging to one plausible event.

When this happens, the voice can stabilise at the centre without appearing glued to either speaker. Instruments occupy a readable volume rather than appearing as thin lateral points. The room does not feel like reverberation placed behind the music; it surrounds and continues the performance.

The silence between instruments stops being empty.

It becomes air.

Why Some Recordings Sound Large but Not Real

A wide soundstage is not necessarily a realistic soundstage.

Modern production provides an enormous range of tools for creating size: multiple microphones, pan controls, delay, artificial reverberation, stereo widening, decorrelation and Mid-Side processing.

These tools can be artistically valuable. They can also create a guitar broader than its physical source, a voice floating beyond the loudspeakers or a reverberant field that could never have existed around the original performers.

There is nothing inherently wrong with that. Recording is also an art of construction.

But width and credibility are not synonymous.

The auditory system evaluates more than the apparent size of the scene. It also evaluates whether the timing, spectral, directional and environmental information is mutually plausible.

A very wide image may be initially impressive. Yet if the centre lacks stability, the reverberation does not belong to the sources or the spatial relationships change unnaturally with frequency, the illusion becomes less convincing.

It is similar to viewing a photograph with intense colour and extraordinary sharpness but incorrect perspective. The image attracts attention immediately, yet something feels wrong.

A recording can contain remarkable detail and separation while remaining spatially two-dimensional.

Beautiful, perhaps.

But not alive.

Every Microphone Technique Makes a Decision

For acoustic music, the choice of microphone technique is never neutral.

AB, ORTF, XY, Mid-Side and Blumlein do not simply create different varieties of stereo width. They encode different combinations of timing, level, polarity, direction and room information.

A spaced AB pair introduces meaningful arrival-time differences between microphones and can create scale, openness and envelopment.

ORTF combines a moderate physical separation with directional cardioid microphones, creating both interchannel time and level differences.

Coincident systems such as XY, Mid-Side and Blumlein minimise the timing difference introduced by microphone spacing and derive direction primarily through level and polarity relationships.

These differences affect localisation, spaciousness, mono compatibility and the way the recording interacts with loudspeaker reproduction.

Coincident techniques can produce comparatively stable and clearly located virtual sources. Spaced techniques may create broader or more diffuse images and can convey strong spaciousness. Neither outcome is automatically better.

The appropriate technique depends on:

  • the musicians and their physical arrangement;
  • the acoustic character of the venue;
  • the desired listening perspective;
  • the balance between localisation and envelopment;
  • the intended distribution format;
  • and the expected playback environment.

There is no microphone technique that is universally correct.

There is only a technique whose compromises are more or less coherent with the intended result.

Recording Is a Translation

A microphone does not hear like a human being.

It has no head, no outer ears, no perceptual memory and no awareness of the room. It does not compare what it captures with years of experience. It measures sound pressure or pressure gradient according to its physical construction and polar pattern.

Human perception, by contrast, is active.

This means that stereo recording is always a translation.

The task is not simply to place two microphones in front of a performance and assume that the original space has been preserved. The task is to create two signals that, when reproduced through loudspeakers in another room, provide the listener with enough coherent information to reconstruct a plausible scene.

That distinction is fundamental.

The original venue, the microphone array, the recording chain, the loudspeakers, the listening room and the listener are all parts of one perceptual system.

A recording can never transport the original acoustic field intact. It selects, encodes and later stimulates a new reconstruction.

The Listening Room Is Part of the Reproduction

With headphones, each channel is delivered predominantly to one ear. With conventional loudspeakers, both speakers reach both ears.

The left ear receives sound from the left loudspeaker, sound from the right loudspeaker after a different path, and reflections from the listening room. The right ear receives the corresponding combination from the opposite side.

The listener’s brain must interpret this new set of binaural cues and construct the phantom images associated with stereo reproduction.

This is why loudspeaker placement, room acoustics and listening position cannot be separated from the recording itself. They participate in the decoding of its spatial information.

A stable recording cannot correct a fundamentally unsuitable room, and a carefully treated room cannot restore information that was never captured or was destroyed during production.

The recording and reproduction environments form a chain.

Natural Reverberation Is More Than a Tail

In a real acoustic environment, reverberation is not an effect that begins after the direct sound has finished.

It is a continuously evolving field of reflections shaped by the dimensions, materials and geometry of the venue.

Those reflections contain:

  • directional asymmetries;
  • different arrival times;
  • frequency-dependent decay;
  • changes in density over time;
  • and relationships to the position and radiation pattern of every instrument.

A reverberation processor can create extraordinarily convincing spaces, and artificial reverberation is indispensable in many forms of production. But a preset does not reproduce the exact interaction that occurred between particular musicians and a particular room at one unrepeatable moment.

When natural reverberation is captured successfully, it does not feel attached to the performance.

It is the performance continuing into the building.

Natural reverberation is architecture becoming sound.

Recording for the Brain

A realistic recording is not necessarily one that captures the largest possible quantity of information.

It is one that preserves the relationships necessary for perception.

For Direct Sound Records, microphone placement, distance, acoustic environment and minimal signal manipulation are therefore not separate technical choices. They form one recording philosophy.

The objective is not purity for its own sake, and it is not nostalgia for a period before digital production.

It is the preservation of continuity.

When a voice is captured from a believable perspective, its centre depends on more than identical level in the two channels. It also depends on stable spectral, temporal and environmental relationships.

When two instruments occupy different sides of an ensemble, their apparent positions should not be understood merely as pan-control settings. They arise from distance, angle, microphone pattern, radiation, reflections and their relationship with the room.

The engineer must therefore decide what should remain coherent, what can be altered and what must be allowed to exist naturally.

Listening Beyond Detail

This perspective also changes how we evaluate recordings and audio systems.

Instead of asking only, “How much detail can I hear?”, we can ask:

  • How credible is the relationship between the sounds?
  • Does the voice remain stable as its pitch and intensity change?
  • Do instruments possess physical body, or are they merely lateral outlines?
  • Does the acoustic environment belong to the performance?
  • Does depth arise from perspective, or only from added reverberation?
  • Does the central image remain convincing at modest listening levels?
  • Do reflections create continuity and air, or do they blur localisation?
  • Does the recording invite prolonged listening, or impress only for a few moments?

These questions move the discussion beyond spectacular sound.

A recording that preserves spatial relationships does not merely demonstrate the capabilities of a system. It gives the system an opportunity to disappear.

When that happens, we stop concentrating on two loudspeakers producing sound.

We begin to perceive musicians, sounding bodies, distance, architecture and silence.

The Space Between Sounds

The most convincing recordings are not necessarily those that contain the most obvious effects, the widest images or the greatest quantity of isolated detail.

They are often those in which every element appears to belong to the same acoustic reality.

The musicians have scale. The centre has physical stability. The room surrounds rather than decorates. Reflections extend the performance instead of obscuring it. Silence defines the distance between one sounding body and another.

This is the space between sounds.

It cannot be reduced to one measurement, one microphone technique or one idea of phase coherence. It emerges from a network of relationships extending from the original performance to the listener’s brain.

When musicians, room, microphone position, recording chain, distribution format, loudspeakers and listening environment align, something rare can occur.

The recording no longer seems to document a performance from the past.

It makes that performance inhabitable again.

Perhaps this is the deepest purpose of high fidelity: not merely to reproduce sounds, but to reconstruct a space in which those sounds can exist.

References and Further Reading

  1. Middlebrooks, J. C. and Green, D. M. “Sound Localization by Human Listeners.”
    Annual Review of Psychology, 1991.
    View publication
  2. Brughera, A., Dunai, L. and Hartmann, W. M. “Human Interaural Time Difference Thresholds for Sine Tones: The High-Frequency Limit.”
    Journal of the Acoustical Society of America, 2013.
    View publication
  3. Bures, Z. and Marsalek, P. “On the Precision of Neural Computation with Interaural Level Differences in the Lateral Superior Olive.”
    Brain Research, 2013.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning Sound Localization with New Ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Pulkki, V. “Microphone Techniques and Directional Quality of Sound Reproduction.”
    Audio Engineering Society, 2002.
    View AES record
  6. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View AES record
  7. Toole, F. E. “Loudspeakers and Rooms for Stereophonic Sound Reproduction.”
    Audio Engineering Society 8th International Conference, 1990.
    View AES record

An earlier version of this essay was published in Audio Review, issue 487, June 2026. This Direct Sound Records Journal edition has been revised, expanded and technically updated for an international readership.