Posts By :

Michelangelo

Vinyl Playback Fingerprint

1024 570 Michelangelo

DSR Technical Journal · Methods paper · DSR-VPF v1.0

Vinyl Playback Fingerprint

A Persistent Multidimensional Method for the Objective Characterisation and Comparison of Analogue Record Playback

Method
DSR Vinyl Playback Fingerprint Method (DSR-VPF) v1.0
Core profile
VPF-Core-1K v1.0
Licence
CC BY 4.0

Document status: Publication version 1.0. The object model, VPF-Core-1K dimensions, display transforms, quality states and comparison rules are fixed for this version. Empirical repeatability and reproducibility validation remains ongoing.

Suggested citation: Canonico, M. (2026). Vinyl Playback Fingerprint: A Persistent Multidimensional Method for the Objective Characterisation and Comparison of Analogue Record Playback. DSR Technical Journal. DSR Vinyl Playback Fingerprint Method, version 1.0.

Abstract

Objective analysis of vinyl playback commonly reports speed error, wow and flutter, channel balance, crosstalk, distortion, harmonic content and noise as separate results. Those measurements are useful, but their isolation makes it difficult to preserve the technical identity of a measurement session, determine whether a later session is meaningfully different, or compare records, pressings and playback systems without losing context. This paper defines the DSR Vinyl Playback Fingerprint Method (DSR-VPF) v1.0: a persistent, versioned, multidimensional measurement object that combines heterogeneous objective results with their validity state, repeatability or uncertainty information, acquisition provenance and comparison rules. The Fingerprint is not the radar graph used to display it; the graph is one projection of an underlying machine-readable object. The first fixed profile, VPF-Core-1K v1.0, comprises speed accuracy, weighted wow and flutter, channel balance, stereo separation, directional crosstalk asymmetry and total harmonic distortion. It fixes dimension order, diagnostic display transforms, quality-to-admissibility mapping, equal metric weights and a minimum numerical-comparison coverage of two thirds. Display bounds are stable coordinates, not pass/fail tolerances or sound-quality thresholds. Invalid or unstable dimensions are not converted into apparently meaningful points. Fingerprint comparison is permitted only when a documented compatibility condition is met, and any aggregate distance is accompanied by coverage, original metric deltas and quality states. Groove Scope is designated as the first reference implementation. DSR-VPF v1.0 establishes the open method architecture and initial profile; empirical repeatability and inter-device reproducibility validation remains ongoing.

Keywords: vinyl playback; turntable measurement; cartridge measurement; wow and flutter; crosstalk; distortion; measurement fingerprint; multidimensional comparison; reproducible measurement; Groove Scope; DSR-VPF

1. Introduction

Vinyl playback is unusually resistant to being reduced to one number. A speed result says little about channel geometry; crosstalk says little about rotational stability; a low total harmonic distortion value does not reveal whether the platter is running at the correct mean speed. Conventional reports therefore present a collection of measurements. This is correct metrologically, but weak as a persistent comparative language: the reader must reconstruct the identity of a session from several pages, and two reports are easily compared selectively rather than systematically.

The purpose of the Vinyl Playback Fingerprint is to provide that comparative language without pretending that heterogeneous measurements are interchangeable. It preserves the original values and units, records whether each value is fit for comparison, maps an explicitly selected subset into a stable multidimensional representation, and binds the result to the measurement conditions and method version that produced it.

The method is intentionally not a universal sound-quality score. Vinyl playback is a coupled mechanical and electrical process, and a single ranking would conceal both causality and uncertainty. The VPF is instead a structured technical identity for a documented playback event. Its most useful question is not ‘Which turntable wins?’ but ‘How, where and with what confidence does this measurement differ from the reference measurement?’

Pipeline from reference signal and playback chain through capture, metric extraction, validity and provenance to a persistent fingerprint object and comparative graphical projections.
Figure 1. Conceptual workflow. The persistent Fingerprint is created after metric extraction and quality control; the graph is a projection of that object, not the object itself.

The method supports three distinct applications: longitudinal observation of one setup, controlled comparison of playback systems, and controlled comparison of records or pressings. These applications share a data structure but do not share identical admissibility conditions. The distinction is central to DSR-VPF.

2. Scope and terminology

2.1 Vinyl Playback Fingerprint

A Vinyl Playback Fingerprint is a persistent, versioned data object that represents a defined set of objective measurements from one documented vinyl playback session, together with validity, uncertainty or repeatability, provenance and comparison metadata. It may be rendered as a radar chart, parallel-coordinate plot, deviation panel, table or another visual form. Changing the rendering does not create a new measurement, provided the underlying object and profile version remain unchanged.

2.2 Why ‘playback’ matters

The captured signal is not produced by the disc alone. It is shaped by the groove, stylus, cartridge, tonearm, turntable drive and bearing, alignment, tracking force, anti-skate, phono equalisation, electrical loading, gain structure, analogue-to-digital converter, clock and measurement environment. Unless those contributions are controlled or characterised, the correct object is a playback fingerprint rather than an intrinsic record fingerprint.

The shorter term Record Fingerprint may be used only for comparisons made through a declared reference chain whose relevant settings, calibration and repeatability are held constant. Even then, the result is an operational characterisation of the record under that reference method, not a metaphysical extraction of the record’s one true identity. Vinyl has enough personality already; the terminology does not need to add more.

2.3 Distinction from content-identification audio fingerprints

In information retrieval, an audio fingerprint normally means a compact content-based signature designed to identify a recording despite noise or transformation [1]. DSR-VPF has a different objective. It does not identify a song. It characterises measured playback behaviour. The term fingerprint is used because the stored multidimensional pattern is intended to be persistent, recognisable and comparable across documented sessions.

3. Existing measurement approaches and the methodological gap

Standards and specialist tools already cover important parts of vinyl measurement. IEC 60098 defines characteristics and agreed measurement methods for analogue audio disc records and reproducing equipment [2]. IEC 60386 addresses speed fluctuations, and its 1988 amendment introduced a preferred two-sigma method [3, 4]. Dr. Feickert Adjust+ and AnalogMagik provide specialised turntable and cartridge measurements [5, 6]. Industrial and academic projects have gone further into pressing analysis, spectral comparison and automated quality assurance [7, 8, 9].

DSR-VPF does not claim priority over any of those measurements, algorithms or graphical conventions. Its narrower contribution is to make a versioned, persistent multidimensional object - rather than an individual test result, a set-up screen, a spectrum or an anomaly alert - the explicit unit of cross-session comparison.

Table 1. Relevant prior approaches and the boundary of the DSR-VPF contribution. Descriptions are limited to public documentation and should not be read as claims about undisclosed internal capabilities.
ApproachPublicly described capabilityBoundary relative to DSR-VPF
IEC 60098 / IEC 60386Definitions and agreed methods for analogue-disc characteristics and speed fluctuation.Standards for individual characteristics; not a persistent heterogeneous fingerprint object.
Dr. Feickert Adjust+Graphical cartridge-azimuth analysis using crosstalk and phase, alongside turntable-related measurements.Multiple tests and graphs; no public definition located of one versioned multidimensional session object designed as the unit of comparison.
AnalogMagikBroad cartridge and turntable optimisation including speed, balance, azimuth/VTA, anti-skate, loading, gain, VTF, vibration, resonance and wow/flutter.Multi-parameter workflow; public documentation presents parameter-by-parameter optimisation rather than a persistent unified fingerprint.
Zaworski, 2022Large pressing study measuring noise, clicks, wow/flutter, stereo bleed and THD; includes rotation-period groove segmentation.Objective multi-metric research, but no persistent graphical identity was identified as the comparison unit.
Stamper Discs FonographFrequency/amplitude comparison of two recordings, including lacquer, DMM, pressing and playback-chain comparisons.The public method is spectral comparison within one measurement domain rather than a heterogeneous measurement object.
Precision Record Pressing AQAAutomated audio analysis, visual inspection and machine learning; synchronised playback compared with a digital master.Industrial QC and anomaly detection; no publicly specified DSR-VPF-like object or open comparison schema was located.

The structured public-documentation review was completed on 3 September 2026 and is summarised in Appendix C. It combined exact-phrase searches with conceptual searches across standards, scholarly repositories, product documentation, industrial quality-control descriptions and public patent web indexes.

No search can prove the absence of an earlier private, unpublished, non-indexed or differently described system. The originality statement in Section 15 is therefore bounded by the public record reviewed and is not a patentability opinion.

4. Measurement model and measurand

A digital capture of vinyl playback may be represented conceptually as a function of the record or test signal R, the mechanical playback chain P, the electrical and conversion chain C, the environment E, the session settings theta, and residual error epsilon. The analysis method A derives measurements from the observed waveform y(t).

The practical consequence is that every fingerprint must name its measurand and scope. A ‘turntable comparison’ is only meaningful when the record, cartridge-related variables and capture path are held sufficiently constant. A ‘pressing comparison’ is only meaningful when the playback and capture chain are held sufficiently constant. A ‘before/after alignment comparison’ requires the changed parameter to be named and other important variables to be controlled.

DSR-VPF therefore treats context as part of the measurement, not as decorative metadata. A fingerprint stripped of its test source, track, playback chain, capture format and method version may still be visually interesting, but it is not admissible for quantitative comparison.

5. The Fingerprint as a persistent data object

DSR-VPF separates measurement, interpretation and display. One compact representation is:

Fingerprint object equation
\mathcal{F}=(id,\mathbf{x},\mathbf{z},\mathbf{q},\mathbf{u},\mathbf{p},v)

where id is an immutable fingerprint identifier; x is the vector of raw measurement values in their original units; z contains profile-defined display or comparison coordinates; q records quality and admissibility; u records uncertainty or repeatability information; p contains provenance and measurement context; and v identifies the method, metric-profile, projection and analysis versions.

This definition creates several safeguards. Raw values remain authoritative. A later change to a display scale cannot silently rewrite history. Invalid measurements remain visible as invalid rather than becoming zero. A comparison algorithm can determine whether two objects are compatible before calculating a distance. Finally, the same object can support more than one visual projection without implying that the data have changed.

  • Multidimensional: the object contains measurements of different phenomena and units.
  • Persistent: it is saved with an immutable identity and can be recalled independently of the live measurement screen.
  • Versioned: method, metric-profile, normalisation, projection and application versions are explicit.
  • Comparable: admissibility and compatibility rules are machine-readable.
  • Auditable: raw values, validity, uncertainty and provenance can be inspected without relying on the shape of a graph.

6. VPF-Core-1K v1.0 metric profile

DSR-VPF may support several metric profiles. VPF-Core-1K v1.0 is the first fixed profile for a compact 1 kHz reference workflow. It is deliberately small enough to remain readable and broad enough to represent rotational, level, stereo and distortion behaviour. Version 1.0 fixes the six dimension identifiers and order, stored quantities, graphical transforms, display bounds, equal weights and default quality mapping. Signal-acquisition and analysis algorithms remain separately versioned so that later improvements do not silently alter historical objects.

Table 2. Fixed dimensions of VPF-Core-1K v1.0.
DimensionStored raw quantityCore projectionAdmissibility note
Speed accuracyMeasured fundamental frequency, nominal frequency, signed speed error (%), calculated RPM.Absolute deviation from nominal; signed value retained in the record.Capture-clock and test-tone accuracy must be documented or bounded.
Weighted wow and flutterWeighted W&F (%) with named standard, weighting and statistic.Increasing deviation axis.Method label such as IEC, DIN or AES must not be collapsed into a generic W&F field.
Channel balanceSigned L-R level difference (dB).Absolute imbalance; sign retained.Requires stable level, defined windowing and adequate headroom.
Stereo separationL-to-R and R-to-L crosstalk ratios plus average separation (dB).Separation deficiency relative to the fixed 40 dB display reference.Both directions and test-track orientation must be stored.
Crosstalk symmetryAbsolute difference between directional separation values (dB).Increasing asymmetry axis.Not interchangeable with average separation.
Total harmonic distortionPer-channel and/or combined THD (%), harmonic limit K, bandwidth and frame-quality statistics.Increasing distortion axis.Excluded when clean continuous tone segments are insufficient or variability exceeds the analysis-method rule.

6.1 Speed accuracy

For a reference tone whose recorded nominal frequency is f0 and measured fundamental is fhat, the signed speed error is:

Signed speed error equation
e_s=100\left(\frac{\hat f}{f_0}-1\right)\%

For a disc nominally rotating at r0, the calculated rotational speed is rhat = r0(1 + es/100). The signed result must be retained because running fast and running slow are diagnostically different, even when the graphical projection uses absolute deviation.

6.2 Weighted wow and flutter

The fingerprint stores the reported value only together with the named weighting curve, detector or statistic, analysis bandwidth, carrier frequency and algorithm version. A value described merely as ‘wow and flutter’ is insufficient for strict comparison because different carriers, weightings and statistics can yield different numbers. The use of a 1 kHz carrier in VPF-Core-1K does not by itself establish conformity with IEC 60386 or another standard that may prescribe additional conditions. An implementation may claim a named standard method only when all relevant requirements are met; otherwise it must identify the result as a versioned DSR-VPF weighted estimate. Where a two-sigma or weighted-peak statistic is used, that choice is explicit [3, 4].

6.3 Channel balance

Channel balance is stored as a signed level difference, conventionally 20 log10(AL/AR), after the analysis has defined how amplitudes are estimated. The projection uses the magnitude of the imbalance, while the raw sign identifies which channel is higher.

6.4 Separation and directional crosstalk

The record stores left-to-right and right-to-left values separately. Because conventions differ, both the signed crosstalk ratio and the corresponding positive separation magnitude should be available. Average separation and directional asymmetry answer different questions and remain separate dimensions.

6.5 Distortion and harmonic profile

For a fundamental amplitude A1 and harmonics A2 to AK, THD may be calculated as the root-sum-square of harmonic amplitudes divided by the fundamental, expressed as a percentage. The fingerprint must declare K, bandwidth, windowing and aggregation method. Harmonic values H2-H5 may be retained as auxiliary dimensions even when the core projection uses a combined THD quantity.

6.6 Fixed diagnostic transforms

The profile uses the following transforms for its radar-deviation projection. All six dimensions have equal profile weight wi = 1.0. The bounds stabilise graphical meaning across sessions; they are not pass/fail tolerances, product specifications, audibility thresholds or claims of good and bad sound.

Table 3. Fixed diagnostic transforms and display bounds for VPF-Core-1K v1.0.
Dimension IDVPF-Core-1K v1.0 transformCentreOuter display bound
speed_errorz = min(|e_s| / 3.00, 1)0% error3.00% absolute error
weighted_wow_flutterz = min(WF / 0.300, 1)0%0.300%
channel_balancez = min(|B_LR| / 3.00, 1)0 dB imbalance3.00 dB imbalance
separation_deficiencyz = min(max(40 - S_avg, 0) / 30, 1)40 dB or greater10 dB or lower
crosstalk_asymmetryz = min(|S_LR - S_RL| / 12, 1)0 dB difference12 dB difference
thdz = min(THD / 3.00, 1)0%3.00%

The profile may additionally store RPM range, raw speed-variation statistics, unweighted wow and flutter components, individual harmonics, residual-noise diagnostics and frame-validity measures. Auxiliary data enrich interpretation but do not automatically become radar axes. More spokes do not necessarily mean more knowledge; sometimes they just mean a busier spider.

7. Capture and extraction requirements

A VPF-compatible capture begins with a declared test source and track. Commercial test records can provide reference and channel-specific signals intended to examine cartridge and turntable behaviour; for example, Ortofon describes its Test Record as containing signals for analysing cartridge performance and its interaction with the tonearm and turntable [10]. DSR-VPF does not assume that any physical test record is perfect. Edition, side, track, condition and known calibration information must therefore be stored.

  1. Document the complete playback chain: turntable, tonearm, cartridge/stylus, tracking force, anti-skate, alignment, azimuth, VTA/SRA where relevant, phono stage, loading, gain and power-supply configuration.
  2. Document the capture chain: interface, analogue routing, sample rate, bit depth, clock source, channel mapping and gain settings.
  3. Acquire with adequate headroom and without automatic gain control, sample-rate conversion or channel processing that is not declared.
  4. Identify stable analysis regions and reject or flag frames affected by clicks, dropouts, mistracking, overload, loss of phase tracking or insufficient continuous tone.
  5. Retain the analysis-method version and enough diagnostic data to explain why each dimension was admitted, downgraded or rejected.

The method does not require publication of copyrighted raw test-record audio. Derived measurements, short lawful visualisations, checksums and provenance can support auditability while respecting the rights attached to the source material.

8. Quality states, repeatability and uncertainty

The International Vocabulary of Metrology distinguishes repeatability as measurement precision under a stated set of repeatability conditions [11]. That idea is especially important for vinyl, where disc centring, groove contamination, stylus seating and transient defects can alter a capture. The Fingerprint therefore stores a quality state for every dimension rather than assuming that one successful export makes every number equally trustworthy.

Table 4. Fixed per-dimension quality states and default admissibility weights.
Quality stateMeaningVPF-Core-1K default q
primary_validMeets the analysis method’s signal, stability and metadata rules.1.0 - included at full profile weight.
supporting_validUsable, but carries a declared limitation or lower evidential status.0.5 - included at reduced admissibility weight.
diagnostic_onlyPotentially informative for inspection, not admitted to aggregate comparison.0.0 - displayed with status, excluded from distance.
variable_or_unreliableCapture variability, contamination or algorithm confidence is outside the analysis rule.0.0 - no numeric plot point; never replaced with zero.
unavailableThe metric was not measured or could not be derived.0.0 - explicitly missing.

A single valid capture may generate a fingerprint, but publication-grade controlled comparisons should state the repeat count and, where practical, use at least three valid captures per condition. This is a methodological recommendation rather than a conformance requirement in v1.0. Robust summaries such as the median and median absolute deviation may be preferable when occasional surface events remain possible. Mean and standard deviation may also be reported where the distribution and sample size justify them.

Repeatability is not the same as reproducibility. Reproducibility testing changes relevant conditions - for example operator, interface, location or reference deck - and asks how much additional variation appears. ISO 5725-2 provides a general framework for estimating repeatability and reproducibility of measurement methods, including preliminary use for methods not yet fully standardised [12]. DSR-VPF validation should use these principles without claiming standards compliance before the required study has been completed.

Where uncertainty cannot yet be expressed as a formal expanded uncertainty, the object should still preserve empirical repeatability intervals, valid-frame proportions and known systematic contributors. Honest partial uncertainty is more useful than a false aura of laboratory certainty.

9. Normalisation and graphical projection

The raw dimensions use incompatible units and directions. A graphical projection therefore requires a profile-defined transform. For a deviation-type metric, one general form is:

Normalisation equation
z_i=\min\left(\frac{|g_i(x_i)-t_i|}{b_i},1\right)

Here gi is a declared transformation, ti is the target or reference coordinate, and bi is a fixed display bound. Table 3 instantiates these terms for VPF-Core-1K v1.0. The bound is not estimated from whichever two fingerprints happen to be on screen. Data-dependent scaling would make shapes look comparable while changing their meaning from one comparison to the next.

The default diagnostic projection places the target or lowest defined deviation at the centre and increasing deviation outward. Separation, for which larger raw values are normally preferable, is converted into a separation-deficiency axis. Signed information remains available in labels and detail views. The display must expose raw values because radial distances are harder to compare accurately than values on a common linear scale.

Radar plots are useful as compact glyphs but can overemphasise polygon area, impose artificial relationships between neighbouring axes and become cluttered when many traces are overlaid. Polygon area must therefore not be used as a quality score, axis order is fixed by profile version, and detailed comparison must also provide aligned values or a delta table. DSR-VPF permits alternative projections where they improve readability, provided the projection identifier is stored.

10. Comparison rules

10.1 Compatibility before distance

Two fingerprints are compared quantitatively only when a compatibility predicate confirms that their profiles, source conditions and intended comparison mode are sufficiently aligned. A visually selectable overlay is not automatically a valid scientific comparison.

Table 5. DSR-VPF v1.0 comparison classes.
ClassIntended useMinimum compatibility
A - LongitudinalBefore/after or ageing observation of one record and playback system.Same record/copy, track, playback chain and method profile; declared intervention and short list of permitted changes.
B - Playback-systemComparison of turntables, arms, cartridges, phono settings or other system variables.Same test source/track and compatible capture method; controlled variables and changed device(s) explicitly named.
C - Record/pressingComparison of copies, pressings, stampers or production stages.Same qualified reference playback/capture chain and method; matched programme region or prescribed test signal.
D - Descriptive onlyExploratory viewing of unlike or incompletely documented sessions.Overlay may be shown with warning; no aggregate distance or superiority claim.

10.2 Missing-data-aware distance

When compatibility is satisfied, a weighted distance can summarise the separation between two normalised vectors. The distance must ignore unavailable or inadmissible dimensions rather than treating them as zero:

Weighted fingerprint distance equation
D_w(A,B)=\sqrt{\frac{\sum_i q_{Ai}q_{Bi}w_i(z_{Ai}-z_{Bi})^2}{\sum_i q_{Ai}q_{Bi}w_i}}

The associated comparison coverage is:

Comparison coverage equation
C(A,B)=\frac{\sum_i q_{Ai}q_{Bi}w_i}{\sum_i w_i}

In VPF-Core-1K v1.0 every profile weight wi equals 1.0 and q follows Table 4. The numerical distance Dw lies between zero and one. It may be reported only when C is at least 2/3 and the relevant Class A, B or C compatibility requirements are met. Below that coverage the comparison is Class D descriptive only; at zero coverage no quantitative comparison exists.

A distance without coverage is incomplete, and distance is not a quality score. Every numerical comparison must also expose signed raw deltas, quality states and repeatability information. Agreement between methods or devices should be examined with techniques appropriate to the design; Bland and Altman’s work remains a useful warning that correlation alone does not demonstrate agreement [13].

11. Worked example using existing Groove Scope archive records

The following example uses two pre-profile Groove Scope 1 kHz records preserved in the Direct Sound Records public measurement archive [14, 15, 16]. Both source records remain drafts, were not acquired under the complete DSR-VPF protocol, and omit several setup fields. Their overlay is therefore Class D descriptive: it demonstrates data handling and visual logic, not an equipment ranking or a valid numerical fingerprint distance.

Table 6. Pre-profile Groove Scope archive data used for the illustrative projection.
MetricGS-2026-0001GS-2026-0002Interpretive status
Speed error-0.71%+2.41%Stable/primary in both source records; sign retained.
Calculated speed33.10 RPM34.14 RPMMeasured under stylus load.
Weighted W&F0.045% DIN-weighted0.037% DIN-shapedLabels require method harmonisation before strict comparison.
Channel balance-0.63 dB+0.82 dBDiagnostic in source records; direction retained.
Average separation30.4 dB26.8 dBDirectional values also stored.
Directional difference5.4 dB3.83 dBSupporting directional diagnostic.
THD0.31% combinedNot admittedGS-2026-0002 distortion was classified variable and unsuitable for setup decisions.
Radar chart comparing two pre-profile Groove Scope measurements across speed deviation, weighted wow and flutter, channel imbalance, separation deficiency, crosstalk asymmetry and THD; the second record has an invalid THD point.
Figure 2. Descriptive projection of GS-2026-0001 and GS-2026-0002 using the fixed VPF-Core-1K v1.0 transforms. Outward means greater deviation, not better or worse sound. The missing GS-2026-0002 THD point demonstrates the quality mask: variable data are not silently plotted as zero. Because the source records are not fully compatible, no aggregate distance is reported.

The example shows why a fingerprint needs more than a polygon. GS-2026-0002 has lower reported weighted wow and flutter but substantially larger mean speed error. GS-2026-0001 has stronger average separation, while GS-2026-0002 has smaller directional asymmetry. The second distortion result cannot be used. None of those facts supports a universal declaration that one complete system is superior. The useful outcome is a structured map of where the sessions differ and where the evidence is incomplete.

12. Intended applications

12.1 Longitudinal setup and maintenance

A baseline fingerprint can be repeated after cartridge alignment, tracking-force adjustment, anti-skate change, belt replacement, bearing service, power-supply change or stylus wear. The fingerprint helps prevent selective attention: an intervention that improves one metric while worsening another remains visible. It can also indicate when an apparent improvement is smaller than the method’s repeatability.

12.2 Cartridge and playback-system comparison

Under a common test source and controlled capture chain, fingerprints can summarise differences between cartridges, tonearms, turntables or phono configurations. The method does not remove the need for listening or explain subjective preference. It provides a repeatable technical companion to those observations.

12.3 Record, pressing and copy comparison

With a qualified reference chain, copies of the same release, test pressings, production runs, alternative compounds or records before and after cleaning can be compared. This application is the point at which the term Record Fingerprint becomes operationally defensible. Rotation-synchronous extensions may later preserve spatially recurring defects or noise patterns; Zaworski’s observation of coherence in rotation-length groove segments provides a relevant research precedent for that direction [7].

12.4 Manufacturing and archival quality assurance

The method could complement industrial master comparison and defect detection by supplying an open, compact summary of named metrics and their changes. It is not a replacement for full-side listening, spectral analysis, visual inspection or automated anomaly detection such as publicly described in Fonograph and AQA [8, 9].

12.5 Public datasets and machine-assisted analysis

A collection of versioned fingerprint objects can support search, clustering, longitudinal maintenance histories and future statistical studies. These uses become credible only when metadata and compatibility rules are preserved. Without them, a large catalogue becomes a large collection of attractive but incomparable shapes.

13. Limitations and failure modes

DSR-VPF is a measurement framework, not an escape hatch from the physical limitations of vinyl. Important limitations include:

  • Test-source error. A reference tone’s cut frequency, eccentricity, wear and contamination can influence the result.
  • Clock error. Frequency-derived speed estimates inherit error from the capture clock unless that clock is verified or corrected.
  • Playback-chain confounding. Cartridge alignment, compliance, loading, gain, RIAA response, arm behaviour and stylus condition all contribute to the capture.
  • Transient contamination. Clicks, mistracking, overload and dropouts can corrupt harmonic and channel estimates unless detected and handled.
  • Position dependence. Outer-, middle- and inner-groove measurements are not automatically interchangeable.
  • Environmental sensitivity. Structure-borne vibration, acoustic feedback and electrical noise can alter results.
  • Visual bias. Radar area and axis adjacency can create an impression unsupported by the raw numbers.
  • Profile dependence. A Fingerprint shape is meaningful only with its profile and normalisation version.
  • Causal overreach. A changed fingerprint identifies a changed session; it does not uniquely identify the responsible component without controlled intervention.

The method must therefore be presented as technical observation rather than certification unless the complete measurement system has been validated for a certification purpose. It must not be used to manufacture league tables from incomparable user submissions.

14. Reference implementation and data stewardship

Groove Scope is designated as the first reference implementation of DSR-VPF. The existing Direct Sound Records Groove Scope Measurements archive already assigns immutable identifiers, preserves human-readable records and machine-readable JSON, records publication status, and anticipates separate application and analysis-method versions [14]. Its publication policy explicitly rejects silent rewriting of historical results and warns against league tables made from incomparable systems [17].

A DSR-VPF v1.0 JSON Schema and conforming example object accompany this publication. The archive schema can incorporate a fingerprint object containing method name, profile identifier, profile version, projection version, raw and transformed dimensions, quality states, repeatability or uncertainty, comparison class and provenance. Historical measurements may be reanalysed only by adding a new analysis or record version while preserving the original report and values.

Publishing the paper, schema and example data under stable identifiers supports the FAIR goals of making digital research objects findable, accessible, interoperable and reusable [18]. The DSR Technical Journal provides the human explanation; a preservation repository can provide a persistent citation; a versioned source repository provides the machine-readable contract; and Groove Scope supplies the reference implementation.

15. Originality and priority statement

The structured public-documentation review described in Appendix C was completed on 3 September 2026. It identified extensive prior work in individual vinyl measurements, multi-parameter setup, graphical display, statistical pressing analysis, spectral comparison and automated quality assurance [2, 3, 4, 5, 6, 7, 8, 9].

No reviewed source was found to explicitly define all of the following as one public method: (1) heterogeneous objective vinyl-playback metrics; (2) a persistent session-level data object; (3) a versioned normalisation or graphical profile; (4) per-dimension validity, admissibility and uncertainty or repeatability; and (5) explicit cross-session comparison rules in which that object is the unit of comparison.

This statement does not claim invention of the underlying measurements, radar charts, content-identification audio fingerprints, record-quality research, spectral comparison or automated quality assurance. It is a bounded priority and attribution statement about their integration and formal use in DSR-VPF. It is not a legal opinion on patent novelty, inventiveness, freedom to operate or trademark availability, and it should be revised if an earlier qualifying public source is identified.

16. Validation and revision programme

Publication version 1.0 freezes the DSR-VPF object architecture and VPF-Core-1K profile so that implementations and datasets have a stable reference. It does not imply that the complete measurement system is a calibrated standard or that inter-laboratory repeatability and reproducibility have already been established.

  1. Specify and version each supported test signal, analysis window, estimator, bandwidth, weighting, sign convention and automated quality rule.
  2. Characterise source and capture-clock error, including a traceable or independently verified frequency reference where practical.
  3. Measure within-session repeatability across multiple valid captures and define rules for robust aggregation.
  4. Test intermediate precision by varying day, operator, record reseating and other realistic conditions.
  5. Assess inter-device and inter-interface reproducibility where Groove Scope supports more than one capture configuration.
  6. Validate sensitivity through controlled interventions expected to affect specific dimensions, while checking for collateral changes.
  7. Review the v1.0 display bounds against empirical distributions and intended use; any altered transform or bound must receive a new profile version and must not rewrite v1.0 objects.
  8. Evaluate agreement between software versions or reference systems with delta analysis and, where suitable, Bland-Altman methods rather than correlation alone [13].
  9. Publish validation datasets, schema revisions, worked examples, change policy and known limitations under stable version identifiers.

Validation may lead to more than one profile. A compact consumer-facing 1 kHz profile, a laboratory reference profile and a pressing-comparison profile need not use identical dimensions. What remains common is the DSR-VPF object model: raw data, validity, uncertainty, provenance, versioning and controlled comparison. Later profiles and method revisions must coexist with, rather than silently mutate, version 1.0.

17. Conclusion

The Vinyl Playback Fingerprint addresses a practical gap between isolated measurement reports and meaningful longitudinal or comparative analysis. Its contribution is not a new speed test, distortion formula or chart type. It is the definition of a persistent, versioned and auditable multidimensional object in which raw vinyl-playback measurements, quality states, uncertainty and provenance remain inseparable from the graphical identity used for comparison.

By separating the measurement object from its visual projection, refusing to convert unreliable data into decorative certainty, and requiring compatibility before distance, DSR-VPF can serve enthusiasts, reviewers, engineers, record labels and pressing plants without collapsing complex analogue behaviour into a simplistic score.

This publication establishes DSR-VPF v1.0 and VPF-Core-1K v1.0 as stable public references. Groove Scope is designated as the first reference implementation, and the Direct Sound Records archive provides a natural public home for versioned examples. The next phase is empirical validation and implementation; any substantive change will be published under a new version rather than applied retroactively.

Declarations

Competing interests: The author is the developer of Groove Scope and the creator of the DSR Vinyl Playback Fingerprint Method. Direct Sound Records publishes the associated measurement archive.

Funding: No external funding was received for this work.

Author contributions: Conceptualisation, methodology, software direction, data curation, visualisation, writing and final approval: Michelangelo Canonico.

Data availability: The illustrative source records are publicly available in the Direct Sound Records Groove Scope Measurements repository [14, 15, 16]. The publication package includes the DSR-VPF v1.0 JSON Schema, the fixed VPF-Core-1K profile and a conforming example object.

Method status: DSR-VPF v1.0 is a published technical method specification. It is not an international standard, accredited test method or certification procedure. Empirical validation remains ongoing and later revisions must be versioned.

Licence: This paper and accompanying open specification are released under Creative Commons Attribution 4.0 International (CC BY 4.0).

Generative-AI assistance: Generative AI tools were used for editorial assistance, literature discovery and document preparation under the author’s supervision. All measurements, scientific claims, interpretations and final responsibility remain with the author.

References

  1. Cano, P., Batlle, E., Kalker, T. and Haitsma, J. (2005). A Review of Audio Fingerprinting. Journal of VLSI Signal Processing Systems, 41, 271-284. Source
  2. International Electrotechnical Commission. IEC 60098:2020: Analogue audio disk records and reproducing equipment. Source
  3. International Electrotechnical Commission. IEC 60386:1972: Method of measurement of speed fluctuations in sound recording and reproducing equipment. Source
  4. International Electrotechnical Commission. IEC 60386:1972/AMD1:1988: Amendment 1 - new preferred two-sigma method. Source
  5. Dr. Feickert Analogue / ProgTec GmbH. (2008). Adjust+: Azimuth/Crosstalk Instructions and Whys. Revision 17 September 2008. Source
  6. AnalogMagik. AnalogMagik Version 2 Cartridge Setup Software - official product and test-function documentation. Source
  7. Zaworski, C. (2022). The optimization of a modern day record press. Master’s thesis, University of Waterloo. Handle 10012/17874. Source
  8. Stamper Discs. Fonograph - objective audio comparison. Public technical description. Source
  9. Daley, S. (2023). PRP Announces New Audio Quality Assurance System. Precision Record Pressing, 14 September 2023. Source
  10. Ortofon. Ortofon Test Record - official product and use description. Source
  11. Joint Committee for Guides in Metrology. JCGM 200:2012: International Vocabulary of Metrology - Basic and general concepts and associated terms, 3rd edition. Source
  12. International Organization for Standardization. ISO 5725-2:2025: Accuracy (trueness and precision) of measurement methods and results - Part 2. Source
  13. Bland, J. M. and Altman, D. G. (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 1(8476), 307-310. PMID 2868172. Source
  14. Canonico, M. and Direct Sound Records. Groove Scope Measurements - public measurement archive. Source
  15. Direct Sound Records. GS-2026-0001: Rega P3 / RB330 / Audio-Technica AT-OC9XML / Audio Research SP20 - 1 kHz Reference Check. Source
  16. Direct Sound Records. GS-2026-0002: Ortofon OM10 / Thorens TD160 Super / SME III - 1 kHz Reference Check. Source
  17. Direct Sound Records. Groove Scope Measurements: Publication and measurement policy. Source
  18. Wilkinson, M. D. et al. (2016). The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data, 3, 160018. Source

Appendix A. Minimum metadata required by DSR-VPF v1.0

A conforming DSR-VPF v1.0 object must provide, or explicitly mark as unavailable where the schema permits, the following information:

  • Immutable fingerprint and measurement identifiers
  • Fingerprint method, profile, normalisation/projection and analysis versions
  • Publication status and record version
  • Measurement date, local time and time zone
  • Comparison class and declared measurand
  • Turntable, tonearm, cartridge/stylus and power supply
  • Tracking force, anti-skate, alignment, azimuth and VTA/SRA where applicable
  • Phono stage, gain, loading and relevant filters
  • Test record, edition, side, track, condition and nominal signal
  • Audio interface, routing, gain, sample rate, bit depth and clock information
  • Per-dimension raw values, units, sign convention and estimator
  • Per-dimension quality state, admissibility weight, valid-frame information and repeatability/uncertainty
  • Normalised coordinates and fixed profile transforms/bounds
  • Known limitations, interventions and environmental notes
  • Source report, checksums, licence and provenance

Appendix B. Example machine-readable DSR-VPF v1.0 object

The example re-expresses selected values from GS-2026-0001 through the fixed VPF-Core-1K v1.0 display transforms. Because the source session predates the complete protocol and remains a draft archive record, it is eligible only for descriptive Class D use.

{
  "fingerprint_id": "VPF-2026-0001",
  "measurement_id": "GS-2026-0001",
  "method": "DSR-VPF",
  "method_version": "1.0",
  "profile": "VPF-Core-1K",
  "profile_version": "1.0",
  "projection": "radar-deviation-1.0",
  "comparison_class_eligibility": ["D"],
  "dimensions": [
    {"id":"speed_error","raw_value":-0.71,"unit":"%",
     "quality":"primary_valid","q":1.0,"z":0.2367},
    {"id":"weighted_wow_flutter","raw_value":0.045,"unit":"%",
     "quality":"primary_valid","q":1.0,"z":0.1500},
    {"id":"channel_balance","raw_value":-0.63,"unit":"dB",
     "quality":"diagnostic_only","q":0.0,"z":0.2100},
    {"id":"separation_deficiency","raw_value":30.4,"unit":"dB separation",
     "quality":"supporting_valid","q":0.5,"z":0.3200},
    {"id":"crosstalk_asymmetry","raw_value":5.4,"unit":"dB",
     "quality":"supporting_valid","q":0.5,"z":0.4500},
    {"id":"thd","raw_value":0.31,"unit":"%",
     "quality":"supporting_valid","q":0.5,"z":0.1033}
  ],
  "provenance": {
    "source_record":"GS-2026-0001",
    "source_record_status":"draft",
    "status":"illustrative_pre_profile_reanalysis",
    "note":"Not acquired under the complete DSR-VPF v1.0 protocol."
  }
}

Appendix C. Public-documentation search protocol

The priority statement in Section 15 is based on a structured review of public documentation completed on 3 September 2026. The review was designed to find systems that combined the same defining elements as DSR-VPF, not merely systems that used the words “audio fingerprint” or displayed individual vinyl measurements.

  • Sources reviewed included IEC and ISO catalogues, scholarly and university repositories, manufacturer and product documentation, pressing-plant quality-control descriptions, public web patent indexes and the existing Direct Sound Records archive.
  • Exact-phrase searches included “vinyl playback fingerprint”, “record playback fingerprint”, “vinyl fingerprint” with measurement terms, and combinations of turntable, cartridge, multidimensional, persistent, graphical, comparison and fingerprint.
  • Conceptual searches covered multi-parameter vinyl measurement, radar or spider plots, stored measurement profiles, record-pressing quality metrics, spectral comparison, automated vinyl quality assurance and rotation-synchronous analysis.
  • A potential anticipation had to combine heterogeneous objective vinyl-playback metrics, a persistent session-level object, a versioned graphical or normalisation profile, per-dimension validity or uncertainty, and explicit cross-session comparison rules.
  • Located precedents were retained in the review even when they did not satisfy all criteria; Table 1 states the relevant boundary rather than dismissing their contribution.
Resonance ab

Once Upon a Time, There Was the Rega Planar: The Resonance Hidden Beneath the Music

1024 639 Michelangelo

A turntable does not operate in isolation. Tonearm mass, cartridge compliance, mounting hardware, furniture, floors and walls all participate in the result. Resonance Lab was created to make those relationships visible—before uncertain listening impressions become unnecessary purchases or endless adjustments.

Once upon a time, there was Rega.

Not Rega as an object of endless forum debate. Not Rega as a flag to be waved in the eternal arguments between belt drive and direct drive, low mass and high mass, suspended designs and rigid plinths.

Rega as an idea.

A simple, almost stubborn idea: remove what is unnecessary, make rigid what must remain still and avoid allowing unwanted energy to accumulate within the structure.

Perhaps that is why a Rega Planar 3 can still appear both modern and slightly old-fashioned. Modern because it is visually restrained, mechanically purposeful and free from excessive mass. Old-fashioned because it recalls a period of British hi-fi in which products often appeared to have been designed to solve practical problems rather than to resemble industrial monuments.

These were not domestic altars constructed from huge slabs of metal and acrylic. They were comparatively light, intelligent instruments intended to operate in real homes.

And this is where the story becomes interesting.

Lightness Is a Design Decision

The Rega Planar 3 is simple, but it is not simplistic.

Rega describes the Planar 3 as using a lightweight laminated plinth reinforced between the tonearm mounting and main bearing by its double-brace structure. The intention is to increase rigidity where it is required without turning the complete plinth into a large energy-storing mass.

This philosophy differs from the approach of a turntable that attempts to resist vibration primarily through weight.

A low-mass, high-rigidity design aims to minimise stored energy and reduce the duration of unwanted resonances within the structure. Rather than trying to become an immovable object, it attempts to manage energy quickly and predictably.

But every engineering philosophy creates conditions under which it performs best.

A lightweight turntable may respond differently to its support and surrounding structure than a very heavy, highly damped design. The equipment table, floor and wall are not automatically external to the turntable system. Under some conditions, they become part of it.

A Rega must therefore be given the right environment in which to behave like a Rega.

That does not always happen.

A Planar 3 on a Suspended Wooden Floor

The system in this case consisted of a modern Rega Planar 3 with its RB330 tonearm and an Audio-Technica AT-OC9XML moving-coil cartridge.

It was an interesting combination. The Microlinear stylus and boron cantilever of the AT-OC9XML offered excellent tracking potential, while the RB330 provided the rigid, low-friction platform around which the Planar 3 had been designed.

The turntable was positioned on a good-quality equipment table.

The table, however, stood on the suspended wooden floor of an English house.

The result was good, but not memorable.

It was controlled, detailed and pleasant, yet it did not quite deliver the immediacy, rhythm and physical presence often associated with a well-installed Rega.

The bass was present but did not feel completely secure. The soundstage opened, but it did not always seem to lock firmly into place. Voices were clear, yet the central image lacked some of the natural solidity that can transform competent reproduction into a convincing musical event.

None of this amounted to an obvious malfunction.

The stylus did not jump. There was no dramatic feedback, no clearly audible mechanical noise and no single defect that could be isolated immediately.

The system simply appeared to play with a small hesitation hidden beneath the music.

The Traditional Audiophile Response

The first instinct in situations like this is often to begin changing things.

We change the interconnect. We try another platter mat. We suspect the cartridge. We adjust tracking force, vertical tracking angle, azimuth and anti-skating. Eventually, we spend an afternoon staring at the tonearm as though it were about to confess.

Audiophiles know this ritual well.

When something does not sound quite right, the mind generates possible explanations faster than they can be tested. Some are technically reasonable. Others are elegant methods of converting uncertainty into maintenance.

Before changing anything, however, it is worth asking a simpler question:

What is the expected mechanical behaviour of the tonearm and cartridge combination?

The Tonearm and Cartridge Resonance

A cartridge suspension behaves like a spring, while the effective moving mass of the tonearm, cartridge and mounting hardware behaves like a mass attached to that spring.

Together they form a resonant mechanical system.

The commonly used estimate is:

fr = 159 / √(M × C)

where:

  • fr is the estimated resonance frequency in hertz;
  • M is the total moving mass in grams;
  • C is the cartridge’s dynamic compliance around the resonance region, expressed in µm/mN or an equivalent compliance unit.

For this system, the published and estimated inputs were:

  • RB330 effective mass: 11 g;
  • AT-OC9XML cartridge mass: 7.6 g;
  • mounting screws and washers: approximately 1–1.5 g.

This produces a total moving mass of approximately 19.6–20.1 g.

The Compliance Problem

The next figure is less straightforward.

Audio-Technica specifies the AT-OC9XML’s dynamic compliance as 16 × 10⁻⁶ cm/dyne at 100 Hz.

Tonearm and cartridge resonance, however, normally occurs much lower, generally somewhere around the single-digit or low-double-digit hertz region. A compliance value measured at 100 Hz cannot simply be inserted into the formula as though it described the suspension identically at 10 Hz.

Compliance is frequency-dependent, and manufacturers do not all publish it under the same test conditions.

Enthusiasts and designers therefore sometimes apply a practical conversion multiplier to Japanese 100 Hz specifications. Values between approximately 1.5 and 2 are commonly explored, but this is a heuristic—not a universal physical conversion law.

Using a factor of 1.7 gives an estimated 10 Hz compliance of:

16 × 1.7 = 27.2 µm/mN

Entering that estimate into the resonance formula gives:

159 / √(19.6 × 27.2) ≈ 6.9 Hz

With the slightly greater mass estimate of 20.1 g, the result becomes approximately:

159 / √(20.1 × 27.2) ≈ 6.8 Hz

The calculation therefore suggests a resonance around 6.8–6.9 Hz under that particular compliance assumption.

But the decimal places must not seduce us into believing that the estimate is more precise than the input data.

If different plausible conversion assumptions are explored, the predicted resonance can move approximately between 6.3 and 7.3 Hz. The real cartridge suspension may also differ from the nominal specification, and mounting conditions can affect the result.

The honest conclusion is therefore not:

“This combination resonates at exactly 6.86 Hz.”

It is:

“This combination is likely to operate near the lower end of the generally preferred resonance region and deserves closer attention.”

What a Low Estimate Actually Means

A predicted resonance below the centre of the preferred range does not automatically mean that the tonearm and cartridge are unusable together.

Ortofon currently describes approximately 7–12 Hz as an optimal region, with 10 Hz as a useful target. It also notes that values around 6.5–7 Hz may still be usable without problems.

This makes the distinction between a warning and a verdict extremely important.

A result near 6.8 or 6.9 Hz does not say:

“Remove this cartridge immediately.”

It says:

  • the combination may be more sensitive to record warps;
  • subsonic energy deserves attention;
  • the turntable support may become particularly important;
  • structural movement should not be dismissed;
  • and calculated behaviour should ideally be checked against observation or measurement.

The system was not necessarily wrong.

It was potentially delicate.

Where Resonance Lab Becomes Useful

This is the purpose of Resonance Lab.

It does not replace listening, and it does not convert analogue reproduction into a simple pass-or-fail calculation.

It provides a structured way to examine the relationship between tonearm effective mass, cartridge weight, mounting hardware and compliance.

Most importantly, it allows the user to see how assumptions change the result.

In a case such as the AT-OC9XML, the compliance conversion should not be hidden behind an apparently unquestionable number. It should be explored.

What happens if the effective compliance is closer to 24 µm/mN?

What happens if it is closer to 27 or 32?

What is the effect of an additional gram of mounting mass?

Would a lighter cartridge move the resonance significantly?

Does the result remain comfortably inside the desired region, or is it strongly dependent on uncertain inputs?

This is more informative than asking whether two products are merely “compatible.”

Compatibility is rarely binary.

A combination may be:

  • comfortably matched;
  • technically usable but sensitive to its environment;
  • dependent on uncertain compliance data;
  • or sufficiently extreme to justify reconsideration.

Resonance Lab makes that uncertainty visible.

What Resonance Lab Does Not Do

A calculation cannot measure a moving floor.

It cannot determine the actual structural resonance of an equipment rack, quantify footfall vibration or prove that a wall shelf will improve every system.

It also cannot know the exact low-frequency compliance of an individual cartridge unless that value has been measured under relevant conditions.

Resonance Lab therefore does not diagnose environmental vibration directly.

Its role is different: it helps identify whether the arm and cartridge combination makes environmental vibration a plausible and technically consistent part of the investigation.

In this case, it did not prove that the wooden floor was responsible.

It made the floor impossible to ignore.

The Floor Enters the System

Suspended wooden floors are elastic structures.

They move under footsteps and distribute low-frequency mechanical energy through joists, boards, furniture and equipment supports. The degree of movement depends on the building, span, construction, loading and position within the room.

This is not automatically a defect. It is simply the behaviour of the structure.

Anyone who has lived in an older English house knows that the building cannot be understood only by looking at it. It creaks, moves, breathes and responds.

An analogue turntable responds too.

When a tonearm and cartridge system is already operating near a relatively low resonance frequency, low-frequency structural movement may become more relevant. The interaction need not be dramatic enough to throw the stylus from the groove.

It may instead appear as:

  • less articulate bass;
  • a centre image that does not feel completely settled;
  • reduced rhythmic certainty;
  • slightly unstable spatial focus;
  • or a vague sense that the performance is not firmly grounded.

These descriptions are subjective listening observations, not unique diagnostic signatures. Similar impressions can have many causes.

But when the calculated arm and cartridge behaviour points towards greater low-frequency sensitivity, the support structure becomes a rational variable to test before purchasing new components.

The Load-Bearing Wall

The turntable was moved from the equipment table to a high-quality wall shelf fixed to a solid load-bearing wall.

The listening change was substantial.

The bass did not simply become more abundant. It became easier to follow. Low notes felt less hesitant and more clearly connected to the musical line.

Transient definition improved. The soundstage became more stable. Voices gained a firmer centre, and instruments occupied more credible positions.

Most importantly, the reproduction acquired a stronger sense of continuity.

The music no longer seemed to pass through a succession of small, invisible obstacles.

It flowed.

This observation does not establish a controlled scientific comparison, and it does not prove that every Rega turntable should be wall-mounted.

It does, however, align with Rega’s own approach. The company produces a lightweight, rigid wall bracket specifically for the Planar 1, Planar 2, Planar 3 and Planar 6, describing it as a vibration-isolation solution intended to complement its lightweight turntables.

The shelf did not add musical information.

Good mechanical engineering rarely adds magic.

It removes interference.

A Case Study, Not a Universal Rule

Not every wooden floor is unsuitable for a turntable.

Not every equipment table performs poorly.

Not every Rega must be placed on a wall shelf, and not every wall is structurally appropriate for supporting one.

A concrete floor and a stable rack may provide excellent conditions. A poorly installed wall shelf may create its own problems. A different tonearm and cartridge combination may be less sensitive to low-frequency movement.

The purpose of this case is not to produce another audiophile commandment.

It is to demonstrate a better sequence of reasoning:

  1. Describe the listening problem without immediately deciding its cause.
  2. Check the published mechanical specifications.
  3. Model the tonearm and cartridge resonance.
  4. Identify uncertainty in the compliance data.
  5. Explore plausible scenarios rather than trusting one exact number.
  6. Consider the support and building structure.
  7. Change one meaningful variable.
  8. Listen again—and measure where possible.

This is far more useful than changing three accessories simultaneously and then attempting to remember which one supposedly transformed the system.

Why I Created Resonance Lab

Tonearm and cartridge resonance has remained unnecessarily mysterious for too long.

The subject often sits somewhere between textbook equations, manufacturer specifications expressed under different conditions, enthusiast-produced compatibility charts and forum discussions in which every confident statement is followed by another confident statement claiming the opposite.

Resonance should not be an initiation ritual.

It should not be reserved for people who enjoy calculations more than music.

It should be a practical tool for better decisions.

I created Resonance Lab to make the relationship understandable and explorable.

The app does not tell the user what to hear. It helps organise the variables that may explain what they are hearing.

It can help reveal whether a combination is:

  • comfortably within a preferred region;
  • near a boundary;
  • highly dependent on an uncertain compliance conversion;
  • or potentially sensitive to environmental conditions.

It can also help prevent expensive misdiagnoses.

A new cartridge will not solve a moving floor if the replacement produces the same mechanical relationship. A different mat will not correct an unsuitable arm and cartridge match. A heavier mounting plate may move the resonance in the wrong direction.

Sometimes the most important upgrade is not another component.

It is a clearer understanding of the system already in front of us.

From Calculation to Reality

The calculated resonance frequency is a starting point, not the final truth.

Where possible, it should be complemented by real-world observation or measurement using a suitable test record and analysis method.

A measured result can reveal the actual resonance peak of the installed system, including the behaviour of the individual cartridge suspension rather than only its nominal specification.

Calculation and measurement serve different purposes:

  • calculation helps evaluate combinations before purchase and explore alternatives;
  • measurement reveals how the installed system actually behaves;
  • listening tells us whether that behaviour is musically significant in the complete system.

None should be forced to perform the role of the others.

The strongest diagnosis emerges when all three point in the same direction.

Understanding Rega Without Worshipping It

Understanding this case does not require worshipping Rega as a brand.

It requires understanding why its design choices make sense.

A Rega is not “simple” in the impoverished sense of the word. It follows a specific path based on lightness, rigidity and controlled energy behaviour.

But a specific design philosophy also requires a suitable context.

Place the turntable on a structure that moves, and under certain conditions it may tell you.

Give it a stable mechanical reference, and it may stop defending itself and begin to communicate the music more freely.

Perhaps this is one reason Rega has retained such a distinctive identity. When the system is working well, the turntable does not seem to be trying to impress the listener.

It simply allows the performance to pass through.

The Relationship Beneath the Music

A turntable is not a collection of isolated products.

The tonearm, cartridge, compliance, mounting screws, support, floor, walls and furniture all participate in its mechanical behaviour.

Sometimes their influence is obvious.

At other times, it does not add a recognisable defect. It removes certainty.

That is the deeper purpose of Resonance Lab.

It does not replace the ears, and it does not promise to solve every analogue problem through a formula.

It makes relationships visible.

In this case, the app showed that the RB330 and AT-OC9XML combination was not absurd, but potentially sensitive. That made the suspended floor a credible variable rather than a piece of inherited audiophile folklore.

The wall shelf then became more than an accessory recommended by tradition.

It became a mechanical response to a specific hypothesis.

Once upon a time, there was Rega.

But Rega is still here.

To hear it properly, we may simply need to understand where it came from, what it is trying to achieve and the environment in which we are asking it to perform.

Because in analogue reproduction, sound never comes from one component alone.

It comes from a relationship.

And sometimes, to rediscover the music, we do not need to replace the turntable.

We simply need to remove the floor from the conversation.

Explore Resonance Lab

Resonance Lab helps vinyl enthusiasts explore tonearm and cartridge compatibility, estimate system resonance and compare how changes in mass or compliance may influence the result.


Learn more about Resonance Lab


Download Resonance Lab from the App Store

References and Technical Sources

  1. Rega Research. “Planar 3.”

    View official product information
  2. Rega Research. “RB330 Tonearm.”

    View official specifications
  3. Audio-Technica. “AT-OC9XML Dual Moving-Coil Stereo Cartridge.”

    View official specifications
  4. Ortofon. “Matching Cartridges with Tonearms.”

    View resonance formula and guidance
  5. Rega Research. “Turntable Wall Bracket.”

    View official product information

This article describes a real-world setup and subjective listening observations supported by resonance modelling. Calculated resonance values are estimates and depend on the accuracy and measurement frequency of the compliance data. They should not be interpreted as a substitute for direct measurement of the installed system.

Where High Fidelity Really Begins: Four Decisions That Shape a Believable Stereo Recording

1024 582 Michelangelo

High fidelity begins before mastering, before the distribution format and before the physical or digital medium. It begins with the decisions made when the performance is captured: how stereo space is encoded, how phase relationships are managed, how far the microphones are placed from the musicians and which microphones are chosen to translate that perspective.

Listening to a stereo recording involves much more than hearing sound emerge from two loudspeakers.

A voice apparently positioned in the centre does not come from a hidden central speaker. An instrument heard slightly to the left is not physically present at that point in the listening room. The apparent depth extending beyond the wall behind the system is not an acoustic space that has suddenly opened in front of us.

These are perceptual constructions.

The auditory system interprets differences in level, timing, spectrum and phase, together with reflections and the relationship between direct and reverberant sound. From this information, it creates a plausible auditory scene.

Stereo recording is therefore not merely the process of producing two channels. It is the art of creating relationships that the listener’s brain can interpret as space.

A High-End System Cannot Invent the Recording

A recording may be highly detailed, dynamically impressive and extended at both ends of the frequency spectrum while still failing to sound believable.

It may reveal the smallest movements of the musicians but construct no stable image. It may sound extremely wide but imprecise. It may attract attention immediately and gradually become artificial or tiring.

For listeners using loudspeakers, this distinction is fundamental.

A revealing system can expose the spatial quality of a recording, but it cannot manufacture coherence that was never captured. If the central image is unstable, the playback equipment may reveal that instability with greater clarity. If the relationship between instruments and ambience is implausible, additional resolution may make the contradiction more obvious.

The format and reproduction chain matter enormously, but they inherit the decisions made at the beginning.

High fidelity starts at the microphone.

Stereo Is Not Simply Width

In audiophile language, we often speak about soundstage, imaging, depth, focus and air around instruments.

These are useful descriptions, but a believable soundstage is not simply a wide one.

Width can be created in many ways:

  • spaced microphones;
  • pan controls;
  • interchannel delays;
  • Mid-Side processing;
  • stereo reverberation;
  • decorrelation;
  • and dedicated spatial processors.

All of these can be artistically valid. But increasing width does not automatically increase realism.

A believable voice should remain stable in the centre, not only in lateral position but also in physical presence. A piano should have plausible dimensions. A guitar should not become wider than the performer playing it unless that enlargement is an intentional artistic decision. A flute should not be reduced to breath, lips and key noise while losing the integrated sound of the instrument.

The acoustic environment should not feel as though it has been placed behind the musicians as an independent effect. It should appear to belong to the same event.

When a recording is coherent, the listener perceives more than individual sources distributed between two loudspeakers. The listener perceives relationships between the instruments, their apparent distances and the room surrounding them.

That is what makes a soundstage credible.

The Four Decisions Behind a Believable Recording

Four closely connected decisions have a particularly strong influence on spatial realism:

  1. the stereo microphone technique;
  2. the timing and phase relationships between channels;
  3. the distance between microphones and performers;
  4. and the microphone characteristics used at that distance.

None of these choices operates independently.

Changing microphone distance changes the balance between direct sound and room ambience. It also changes proximity effect, source integration, off-axis contribution and the amount of environmental noise captured.

Changing the microphone changes the polar pattern, tonal balance, transient behaviour, self-noise and off-axis response experienced at that distance.

Changing the stereo array changes the interchannel timing and level relationships supplied to the reproduction system.

The art lies in making these decisions support the same perceptual objective.

1. Stereo Technique: Different Ways of Encoding Space

Stereo microphone techniques are not simply alternative arrangements for producing a left and right channel. They encode spatial information in fundamentally different ways.

Spaced Pairs

In an AB arrangement, two microphones are separated physically. A sound arriving from one side normally reaches one microphone before the other, producing an interchannel time difference. Depending on the source, microphone pattern and geometry, differences in level may also occur.

Spaced arrays can create a broad impression of scale and envelopment. In a good concert hall, and particularly with larger ensembles, this can communicate openness, low-frequency spaciousness and the feeling that the performance breathes within a large acoustic environment.

Physical spacing also introduces frequency-dependent phase relationships between the channels. In stereo these can contribute to spaciousness, but they may reduce localisation precision or create tonal changes when the channels are summed or partially combined.

This does not make AB inherently defective. It means that width, envelopment, localisation and mono compatibility must be balanced deliberately.

Near-Coincident Techniques

Near-coincident arrangements such as ORTF, NOS and DIN combine physical microphone spacing with directional polar patterns.

ORTF, for example, uses two cardioid microphones separated by 17 centimetres and angled 110 degrees apart. The resulting stereo image contains both interchannel timing and level differences.

This can offer a productive compromise: greater spaciousness than many coincident cardioid arrangements, together with more definite image positioning than a widely spaced pair.

The timing component is not an accidental defect. It is part of the intended spatial design.

Coincident Techniques

In coincident arrangements such as XY, Mid-Side and Blumlein, the microphone capsules are placed as close as physically possible to the same acoustic point.

Because the direct sound reaches both capsules at approximately the same time, directional information is created primarily through differences in level and polarity produced by the microphones’ polar patterns.

Coincident geometry generally offers:

  • stable localisation;
  • a clearly defined central image;
  • predictable mono compatibility;
  • and fewer time-delay interactions introduced by microphone spacing.

The soundstage can sometimes appear less expansive than that produced by a spaced array, but individual positions may be easier to read.

Practical coincidence is never perfect. The capsules have physical dimensions, and real microphones exhibit frequency-dependent polar and phase responses. Careful construction and positioning still matter.

No Technique Is Universally Superior

The appropriate technique depends on:

  • the size and arrangement of the ensemble;
  • the acoustic character of the venue;
  • the required balance between localisation and spaciousness;
  • the importance of mono compatibility;
  • the intended listening perspective;
  • and the expected playback system.

When scale and strong hall envelopment are priorities, a spaced arrangement may be highly effective. When image stability and coincident timing are central to the project, XY, Mid-Side or Blumlein may offer important advantages.

There is no universally correct technique.

There is only a technique whose characteristics are coherent with the recording’s purpose.

2. Phase: More Than a Technical Problem

In audio, phase is frequently discussed only when something has gone wrong.

We notice it when bass becomes thin, when a centre image loses solidity, when combining microphones produces tonal colouration or when a stereo recording behaves poorly in mono.

These are genuine problems, but phase is also part of spatial information.

Timing and phase relationships between channels can contribute to the perception of width, localisation, ambience and depth. However, the differences captured by two microphones are not identical to the binaural cues produced at two human ears.

Microphones have no head between them, no pinnae and no torso. They do not apply the listener-specific spectral transformations associated with natural localisation. Their spacing and polar patterns create a new encoding intended for reproduction through another system.

This distinction is crucial.

In conventional loudspeaker stereo, each ear hears both speakers. The listening room adds reflections, and the listener’s head modifies the signals again. The brain must interpret this combined information and construct phantom images.

A large interchannel delay can increase spaciousness without necessarily improving image precision. A coincident array reduces the timing differences introduced at capture, but it cannot eliminate every phase interaction in the room, microphone, loudspeaker or recording chain.

The meaningful objective is not perfect phase identity.

It is maintaining interchannel relationships that remain sufficiently consistent for the listener to construct a stable scene.

Phase and the Centre Image

A central phantom image is created when the loudspeakers provide the auditory system with compatible information suggesting that a source lies between them.

Equal level alone is not always sufficient to make that centre feel physical. The spectral and temporal content of the two channels must also support the same perceptual interpretation.

If multiple microphones capture one source with different delays, the resulting interference may change with frequency. The image can become less stable, and tonal character may vary according to the combination of channels and listening position.

This is one reason why microphone count should not be confused with information quality.

More microphones can offer flexibility, control and creative possibilities. They can also create more relationships that must be managed.

3. Microphone Distance: Where Proportion Begins

Microphone distance is fundamental to whether an instrument sounds credible through loudspeakers.

A close position can provide immediacy, clarity and extraordinary detail. It can reveal the movement of piano mechanics, a flautist’s breathing, fingers touching strings, valve noise, bow texture and the precise attack of every note.

These details can be fascinating and musically valuable.

But they are not always proportionate to the way the instrument would be heard from a natural listening position.

The danger is that detail becomes confused with realism.

The Instrument Must Reassemble

Acoustic instruments often radiate different frequency regions from different parts of their bodies and in different directions.

At very close range, a microphone hears one local part of that radiation field. Moving farther away allows those contributions to integrate more fully before reaching the capsule.

The piano can become one sounding body rather than a collection of strings, hammers and mechanical events. The guitar returns to plausible physical dimensions. The flute becomes more than the excitation point at the mouthpiece; it becomes an instrument projecting energy into the room.

Distance allows the instrument to reassemble itself.

Distance Introduces the Room

Moving a microphone away also changes the balance between direct and reflected sound.

More of the room enters the recording. Early reflections influence tone and localisation. Reverberation communicates scale and distance. Background noise and undesirable acoustic characteristics become harder to avoid.

The correct distance is therefore not simply the most natural one in theory. It is the distance at which source integration, clarity, instrumental body and room contribution reach the desired balance.

That point changes with every venue, ensemble and microphone.

Close Is Not Wrong

Close microphone placement should not be treated as inherently artificial.

Many musical genres depend on intimacy, isolation, impact or the ability to balance sources independently. A close perspective may be exactly right for the artistic language of the production.

The problem arises only when a local, magnified perspective is presented as though it were automatically more faithful because it reveals more detail.

Detail describes how much can be perceived.

Realism describes whether those details belong to a plausible whole.

4. Microphone Choice: An Instrument of Proportion

At a realistic recording distance, the microphone is not merely a transparent transducer.

It becomes an instrument of proportion.

Every microphone has a technical personality shaped by factors including:

  • its operating principle;
  • diaphragm or ribbon construction;
  • polar pattern;
  • frequency and phase response;
  • off-axis behaviour;
  • transient response;
  • self-noise and sensitivity;
  • grille and body geometry;
  • electronics and transformers;
  • and manufacturing tolerances.

No microphone is perfectly neutral under every condition.

A microphone that sounds impressive at close range may not be the right choice at several metres. A presence rise that creates attractive clarity nearby may make a distant recording feel thin or overly explicit. A microphone with excellent on-axis response but irregular off-axis behaviour may colour the room contribution as distance increases.

Conversely, a microphone whose polar pattern and tonal balance remain well controlled away from the axis may integrate the direct and reverberant fields more convincingly.

Ribbon and Condenser Microphones

The choice should not be reduced to a contest between ribbon and condenser technology.

Modern condenser microphones can provide extended bandwidth, low self-noise, high sensitivity and carefully controlled directional behaviour. These qualities can be invaluable for distant acoustic recording.

Some ribbon microphones offer a different balance: smooth high-frequency behaviour, figure-of-eight directivity and a substantial sense of instrumental body. These qualities can suit coincident Blumlein recording and certain natural-distance applications particularly well.

But the result depends on the individual microphone, not merely the category printed on its specification sheet.

A ribbon is not automatically warm or natural. A condenser is not automatically bright or analytical.

The meaningful question is:

Does this microphone, at this distance and in this room, preserve the proportions required by the music?

Proximity Effect and Tonal Perspective

Directional pressure-gradient microphones exhibit proximity effect: their low-frequency response increases as the source moves closer.

The effect is generally strongest with figure-of-eight patterns and is also present, to a lesser degree, with cardioid and related directional patterns. Pure pressure-operated omnidirectional microphones do not exhibit conventional proximity effect.

At very close distances, proximity effect may exaggerate bass and make a source appear larger than its natural scale. It can also be used creatively to provide weight, intimacy or authority.

As the microphone moves farther from the source, this low-frequency boost diminishes. The resulting change in tonal balance must be considered alongside the growing contribution of the room.

This is one reason why microphone choice and distance cannot be separated. The correct working distance is not established by geometry alone; it must also produce an appropriate tonal foundation.

The objective is not to enlarge the bass artificially. It is to preserve enough body for the instrument to remain physical without making it implausibly large.

Why Blumlein Deserves Particular Attention

Among coincident techniques, the Blumlein pair occupies a distinctive position.

It uses two figure-of-eight microphones mounted coincidently and angled 90 degrees apart. Direction is encoded primarily through level and polarity differences, while the rear lobes capture a substantial part of the surrounding acoustic environment.

This combination can provide:

  • a stable central image;
  • clearly organised lateral localisation;
  • strong mono compatibility;
  • and a naturally integrated representation of the room.

Blumlein is also demanding.

The room must contribute positively because sound arriving from behind the array is captured strongly. Placement is critical. The ensemble must balance acoustically, and the useful recording angle must suit its arrangement.

A poor room is revealed rather than concealed. An incorrect microphone position can produce too much reverberation, an inappropriate stereo spread or an imbalanced ensemble.

That apparent limitation is also part of the technique’s value.

A minimally manipulative method encourages the main problems to be solved before recording begins rather than postponed until post-production.

The Room Is Not an Effect Added Afterwards

At natural microphone distances, the room inevitably becomes part of the recording.

Not every space deserves that responsibility.

A room may be excessively dry, too small, mechanically noisy, confused in the lower frequencies or dominated by unattractive early reflections. When the room is unsuitable, a purist approach does not transform it into a virtue.

But when the acoustic environment supports the music, natural ambience can provide an unusually coherent relationship between source and space.

The early reflections, reverberant build-up, asymmetries, frequency-dependent decay and interaction with instrumental radiation all belong to one event.

They are not a separate layer added later.

They are part of the way the music happened.

Artificial reverberation can be beautiful, realistic and artistically indispensable. The distinction is not between legitimate natural sound and illegitimate processing.

The distinction is whether the reverberant information supports the same perspective and spatial logic as the direct sound.

When natural reverberation is captured successfully, the room does not sit behind the instruments.

It connects them.

What the Listener Can Evaluate

These ideas are relevant not only to recording engineers. They can also change how an audiophile evaluates recordings and equipment.

Instead of asking only how much detail is audible, how deep the bass extends or how wide the soundstage appears, listen for relationships.

Listen to the Centre

Does a central voice feel stable and physical, or merely like a thin point suspended between the speakers?

Does its position and body remain coherent as pitch, intensity and register change?

Listen to Instrumental Proportion

Do the instruments possess believable dimensions?

Does the piano appear as one body, or as a series of enlarged local details? Does a guitar occupy plausible space? Does a flute retain tone and projection rather than becoming mostly breath and mechanism?

Listen to Complexity

Does the soundstage remain intelligible when the musical texture becomes dense?

A recording may appear sharply separated during a simple passage and lose all spatial organisation when multiple instruments play simultaneously.

Listen to Decay

Do notes decay continuously into the same environment in which they began?

Does a piano chord dissolve naturally into the surrounding space? Does the room respond to the music, or does the reverberation seem to operate as a separate layer?

Listen at Moderate Volume

A spatially coherent recording often remains intelligible without being played loudly. The centre continues to exist, instrumental positions remain readable and the acoustic environment still suggests depth.

Exaggerated spectral balance and artificial spatial effects may depend more strongly on level to remain impressive.

Listen for Credibility, Not Size

A realistic instrument does not need to sound enormous.

It needs to sound plausible.

Many recordings impress by enlarging everything. Enlargement can be artistically exciting, but it is not automatically high fidelity.

The Format Preserves; It Does Not Create

Audiophile discussions often concentrate on formats: vinyl, analogue tape, PCM, DSD and high-resolution distribution.

These discussions matter, but the format is sometimes treated as though it were the original source of realism.

A format can preserve captured information with greater or lesser accuracy. It cannot create spatial information that was never recorded.

If the soundstage is unstable, the medium preserves that instability. If the centre is weak, a particular format may alter the subjective presentation but cannot reconstruct the original microphone geometry. If the ambience bears no coherent relationship to the instruments, greater resolution may simply reveal that separation more clearly.

The format matters—but it comes later.

This does not diminish the value of excellent recording and distribution formats. It clarifies their purpose.

A high-quality medium is most meaningful when it is preserving something worth preserving:

  • musical dynamics;
  • instrumental timbre;
  • temporal relationships;
  • spatial information;
  • natural decay;
  • and believable proportion.

High Fidelity as Coherence

Stereo technique determines how spatial cues are encoded.

Phase and timing relationships influence image stability, width and the behaviour of combined signals.

Microphone distance determines the proportion between source, detail and environment.

Microphone choice determines how that distance is translated tonally and spatially.

The room either contributes meaningfully to the performance or becomes another problem to manage.

The recording format preserves these decisions but cannot replace them.

For Direct Sound Records, this is the foundation of natural acoustic recording. The objective is not to impose a spectacular soundstage but to preserve enough coherent information for the listener to reconstruct one.

A great recording does not necessarily astonish immediately.

It often becomes more convincing over time.

It does not enlarge every instrument. It preserves proportion. It does not use reverberation merely as decoration. It allows the acoustic environment to participate in the music.

Two loudspeakers can accomplish something extraordinary: they can suggest the presence of a space that does not physically exist in front of the listener.

For that illusion to succeed, the recording must contain believable information. It must respect the way we hear and provide the brain not only with individual sounds, but with meaningful relationships between them.

Perhaps this is where a recording becomes genuinely high fidelity—not when it reveals everything, but when it allows us to believe what we are hearing.

Related Articles

  • The Space Between Sounds: How the Brain Reconstructs the Soundstage
  • Why the Blumlein Pair Can Excel in Loudspeaker Playback

References and Further Reading

  1. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1986.
    View AES record
  2. Gerzon, M. A. “The Design of Precisely Coincident Microphone Arrays for Stereo and Surround Sound.”
    Audio Engineering Society 50th Convention, 1975.
    View AES record
  3. Politis, A., Laitinen, M.-V., Ahonen, J. and Pulkki, V.
    “Parametric Spatial Audio Processing of Spaced Microphone Array Recordings for Multichannel Reproduction.”
    Journal of the Audio Engineering Society, 2015.
    View AES record
  4. DPA Microphones. “Stereo Recording Techniques and Setups.”
    Read technical guide
  5. DPA Microphones. “ORTF.”
    View technical definition
  6. Neumann. “What Is the Proximity Effect?”
    Read technical guide
  7. Bock, T. M. and Keele, D. B.
    “The Effects of Interaural Crosstalk on Stereo Reproduction and Minimizing Interaural Crosstalk in Nearfield Monitoring.”
    Audio Engineering Society 81st Convention, 1986.
    View AES record
  8. Schneider, M. “MS Mastering of Stereo Microphone Signals.”
    Audio Engineering Society 132nd Convention, 2012.
    View AES record

An earlier version of this article was published in Audio Review, July/August 2026, and was subsequently adapted for LinkedIn. This Direct Sound Records Journal edition has been substantially revised, expanded and technically updated.

Space and Sound

The Space Between Sounds: How the Brain Reconstructs the Soundstage

1024 755 Michelangelo

In musical realism, we do not listen only to frequencies, timbres and dynamics. We listen to relationships: differences in arrival time, level, reflection and spatial organisation. When those relationships remain credible, a recording can become more than a collection of sounds. It can become an inhabitable acoustic space.

When listening to a high-fidelity system, we often say that a recording “sounds real.” But what does that actually mean?

Frequency extension matters. So do controlled bass, natural midrange, transient response, low distortion and harmonic detail. Yet realism also depends on something less obvious and considerably more fragile: whether the recording and playback system allow us to perceive a believable relationship between sound sources and the space around them.

A voice is not simply a voice. It is a presence located at a particular distance and position.

A piano is not merely a collection of strings, hammers and resonances. It is a physical body occupying a volume of air, projecting energy into a room and interacting with surrounding surfaces.

A quartet is not simply divided into left, centre and right. It is an acoustic event in which musicians, distances, reflections and silence form one spatial relationship.

The human auditory system does not receive this scene passively. It interprets, compares and reconstructs it. In a sense, the brain triangulates.

The Soundstage Is Not an Effect

In high-end audio, words such as soundstage, depth, focus, air and presence are used constantly. They are useful descriptions, but they can become vague unless we remember that they have a perceptual foundation.

The soundstage is not a picture physically stored inside a recording. Nor is it simply drawn between two loudspeakers.

It is a perceptual reconstruction created from the information reaching the listener’s ears.

The brain interprets several cues simultaneously:

  • tiny differences in the time at which sound reaches each ear;
  • differences in level between the ears;
  • direction-dependent changes in the spectrum;
  • the relationship between direct and reflected sound;
  • the evolution of reflections and reverberation over time;
  • and prior knowledge of familiar voices, instruments and environments.

When these cues support one another, a sound source can appear stable, physical and separate from the loudspeakers.

When they conflict, a recording may still sound impressive, detailed or extremely wide, but the scene can feel unstable or artificial.

The soundstage is therefore not merely an effect.

It is information interpreted by the listener.

How the Brain Reads Acoustic Space

For horizontal localisation, the auditory system relies heavily on differences in timing and level between the two ears.

Interaural Time Differences

When a sound arrives from one side, it normally reaches the nearer ear slightly before the farther ear. These interaural time differences, or ITDs, can be extremely small—sometimes measured in tens of microseconds—yet the auditory system is remarkably sensitive to them.

Timing cues are particularly important at lower frequencies, where neural activity can represent the temporal fine structure of the waveform with sufficient precision. Structures within the auditory brainstem, especially the medial superior olive, contribute to the analysis of these binaural timing relationships.

ITD sensitivity does not end at one perfectly defined frequency, but sensitivity to the fine structure of pure tones deteriorates substantially through the region around 1 to 1.5 kHz. Higher-frequency sounds can still convey timing information through changes in their amplitude envelopes.

Interaural Level Differences

At shorter wavelengths, the head creates a more substantial acoustic shadow. A sound arriving from the left will generally produce a higher level at the left ear than at the right.

This interaural level difference, or ILD, becomes an increasingly useful directional cue as frequency rises. Neural circuits involving the lateral superior olive contribute to its early processing.

In natural listening, timing and level differences do not operate as two isolated systems with a rigid boundary between them. Broadband sounds contain multiple cues, and the brain combines them according to frequency, source position, environment and reliability.

Spectral Cues and the Shape of the Listener

Timing and level differences provide strong information about left and right, but they cannot always distinguish whether a sound is above, below, in front of or behind the listener.

For those dimensions, the shape of the outer ears, head and torso becomes essential.

The folds of the pinnae filter incoming sound differently according to direction. Certain frequencies are reinforced, while others are attenuated or notched. The complete transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because human anatomy varies, each person’s HRTF is individual. Over time, the brain learns the correspondence between these spectral patterns and positions in space.

Experiments in which the outer ears were temporarily reshaped have shown that vertical localisation initially becomes much less accurate. With experience, listeners can learn to interpret the altered cues—evidence that spatial hearing is not merely mechanical, but calibrated and adaptive.

What we perceive is therefore never only the sound emitted by the source. It is the result of a relationship between source, environment, listener and brain.

Where Phase Enters the Picture

In audio engineering, phase is often discussed as a problem.

Signals may be described as “out of phase” when they weaken the centre image, reduce low-frequency energy, produce cancellations or behave unpredictably in mono.

These concerns are real, but they represent only part of the subject.

Phase is not simply an error waiting to be corrected. Phase and timing relationships can also carry important information about position, width, depth and the interaction between sound sources.

For a periodic low-frequency signal, a difference in arrival time between the ears can also be expressed as an interaural phase difference. In this sense, phase contributes directly to spatial localisation.

Within a stereo recording, interchannel timing and phase relationships can influence:

  • the position and stability of phantom images;
  • the apparent width of the presentation;
  • the sense of distance and depth;
  • the relationship between direct sound and ambience;
  • mono compatibility;
  • and frequency-dependent reinforcement or cancellation when signals combine.

However, it would be misleading to say that phase is used only for location and never contributes to the perceived character of a sound. Temporal fine structure is also involved in pitch, masking and the separation of simultaneous sources. In recording and reproduction, phase relationships can additionally alter the spectrum whenever correlated signals combine acoustically or electrically.

The useful distinction is therefore not between phase and sound, but between the different roles that temporal relationships perform.

Phase is not the sound itself, but it can help organise the space in which that sound is perceived.

Coherence Does Not Mean Perfection

The expression phase coherence is frequently used as though it described one measurable quality that a recording either possesses or lacks.

Reality is more complicated.

Every acoustic environment contains delays. Reflections arrive after the direct sound and from different directions. Instruments radiate differently according to frequency. Microphones have frequency-dependent polar patterns. Loudspeakers and rooms introduce further interactions.

A natural acoustic event is not phase-identical at every point in space.

Spatial coherence should therefore not mean eliminating every difference or delay. It means preserving relationships that remain compatible enough for the auditory system to interpret them as belonging to one plausible event.

When this happens, the voice can stabilise at the centre without appearing glued to either speaker. Instruments occupy a readable volume rather than appearing as thin lateral points. The room does not feel like reverberation placed behind the music; it surrounds and continues the performance.

The silence between instruments stops being empty.

It becomes air.

Why Some Recordings Sound Large but Not Real

A wide soundstage is not necessarily a realistic soundstage.

Modern production provides an enormous range of tools for creating size: multiple microphones, pan controls, delay, artificial reverberation, stereo widening, decorrelation and Mid-Side processing.

These tools can be artistically valuable. They can also create a guitar broader than its physical source, a voice floating beyond the loudspeakers or a reverberant field that could never have existed around the original performers.

There is nothing inherently wrong with that. Recording is also an art of construction.

But width and credibility are not synonymous.

The auditory system evaluates more than the apparent size of the scene. It also evaluates whether the timing, spectral, directional and environmental information is mutually plausible.

A very wide image may be initially impressive. Yet if the centre lacks stability, the reverberation does not belong to the sources or the spatial relationships change unnaturally with frequency, the illusion becomes less convincing.

It is similar to viewing a photograph with intense colour and extraordinary sharpness but incorrect perspective. The image attracts attention immediately, yet something feels wrong.

A recording can contain remarkable detail and separation while remaining spatially two-dimensional.

Beautiful, perhaps.

But not alive.

Every Microphone Technique Makes a Decision

For acoustic music, the choice of microphone technique is never neutral.

AB, ORTF, XY, Mid-Side and Blumlein do not simply create different varieties of stereo width. They encode different combinations of timing, level, polarity, direction and room information.

A spaced AB pair introduces meaningful arrival-time differences between microphones and can create scale, openness and envelopment.

ORTF combines a moderate physical separation with directional cardioid microphones, creating both interchannel time and level differences.

Coincident systems such as XY, Mid-Side and Blumlein minimise the timing difference introduced by microphone spacing and derive direction primarily through level and polarity relationships.

These differences affect localisation, spaciousness, mono compatibility and the way the recording interacts with loudspeaker reproduction.

Coincident techniques can produce comparatively stable and clearly located virtual sources. Spaced techniques may create broader or more diffuse images and can convey strong spaciousness. Neither outcome is automatically better.

The appropriate technique depends on:

  • the musicians and their physical arrangement;
  • the acoustic character of the venue;
  • the desired listening perspective;
  • the balance between localisation and envelopment;
  • the intended distribution format;
  • and the expected playback environment.

There is no microphone technique that is universally correct.

There is only a technique whose compromises are more or less coherent with the intended result.

Recording Is a Translation

A microphone does not hear like a human being.

It has no head, no outer ears, no perceptual memory and no awareness of the room. It does not compare what it captures with years of experience. It measures sound pressure or pressure gradient according to its physical construction and polar pattern.

Human perception, by contrast, is active.

This means that stereo recording is always a translation.

The task is not simply to place two microphones in front of a performance and assume that the original space has been preserved. The task is to create two signals that, when reproduced through loudspeakers in another room, provide the listener with enough coherent information to reconstruct a plausible scene.

That distinction is fundamental.

The original venue, the microphone array, the recording chain, the loudspeakers, the listening room and the listener are all parts of one perceptual system.

A recording can never transport the original acoustic field intact. It selects, encodes and later stimulates a new reconstruction.

The Listening Room Is Part of the Reproduction

With headphones, each channel is delivered predominantly to one ear. With conventional loudspeakers, both speakers reach both ears.

The left ear receives sound from the left loudspeaker, sound from the right loudspeaker after a different path, and reflections from the listening room. The right ear receives the corresponding combination from the opposite side.

The listener’s brain must interpret this new set of binaural cues and construct the phantom images associated with stereo reproduction.

This is why loudspeaker placement, room acoustics and listening position cannot be separated from the recording itself. They participate in the decoding of its spatial information.

A stable recording cannot correct a fundamentally unsuitable room, and a carefully treated room cannot restore information that was never captured or was destroyed during production.

The recording and reproduction environments form a chain.

Natural Reverberation Is More Than a Tail

In a real acoustic environment, reverberation is not an effect that begins after the direct sound has finished.

It is a continuously evolving field of reflections shaped by the dimensions, materials and geometry of the venue.

Those reflections contain:

  • directional asymmetries;
  • different arrival times;
  • frequency-dependent decay;
  • changes in density over time;
  • and relationships to the position and radiation pattern of every instrument.

A reverberation processor can create extraordinarily convincing spaces, and artificial reverberation is indispensable in many forms of production. But a preset does not reproduce the exact interaction that occurred between particular musicians and a particular room at one unrepeatable moment.

When natural reverberation is captured successfully, it does not feel attached to the performance.

It is the performance continuing into the building.

Natural reverberation is architecture becoming sound.

Recording for the Brain

A realistic recording is not necessarily one that captures the largest possible quantity of information.

It is one that preserves the relationships necessary for perception.

For Direct Sound Records, microphone placement, distance, acoustic environment and minimal signal manipulation are therefore not separate technical choices. They form one recording philosophy.

The objective is not purity for its own sake, and it is not nostalgia for a period before digital production.

It is the preservation of continuity.

When a voice is captured from a believable perspective, its centre depends on more than identical level in the two channels. It also depends on stable spectral, temporal and environmental relationships.

When two instruments occupy different sides of an ensemble, their apparent positions should not be understood merely as pan-control settings. They arise from distance, angle, microphone pattern, radiation, reflections and their relationship with the room.

The engineer must therefore decide what should remain coherent, what can be altered and what must be allowed to exist naturally.

Listening Beyond Detail

This perspective also changes how we evaluate recordings and audio systems.

Instead of asking only, “How much detail can I hear?”, we can ask:

  • How credible is the relationship between the sounds?
  • Does the voice remain stable as its pitch and intensity change?
  • Do instruments possess physical body, or are they merely lateral outlines?
  • Does the acoustic environment belong to the performance?
  • Does depth arise from perspective, or only from added reverberation?
  • Does the central image remain convincing at modest listening levels?
  • Do reflections create continuity and air, or do they blur localisation?
  • Does the recording invite prolonged listening, or impress only for a few moments?

These questions move the discussion beyond spectacular sound.

A recording that preserves spatial relationships does not merely demonstrate the capabilities of a system. It gives the system an opportunity to disappear.

When that happens, we stop concentrating on two loudspeakers producing sound.

We begin to perceive musicians, sounding bodies, distance, architecture and silence.

The Space Between Sounds

The most convincing recordings are not necessarily those that contain the most obvious effects, the widest images or the greatest quantity of isolated detail.

They are often those in which every element appears to belong to the same acoustic reality.

The musicians have scale. The centre has physical stability. The room surrounds rather than decorates. Reflections extend the performance instead of obscuring it. Silence defines the distance between one sounding body and another.

This is the space between sounds.

It cannot be reduced to one measurement, one microphone technique or one idea of phase coherence. It emerges from a network of relationships extending from the original performance to the listener’s brain.

When musicians, room, microphone position, recording chain, distribution format, loudspeakers and listening environment align, something rare can occur.

The recording no longer seems to document a performance from the past.

It makes that performance inhabitable again.

Perhaps this is the deepest purpose of high fidelity: not merely to reproduce sounds, but to reconstruct a space in which those sounds can exist.

References and Further Reading

  1. Middlebrooks, J. C. and Green, D. M. “Sound Localization by Human Listeners.”
    Annual Review of Psychology, 1991.
    View publication
  2. Brughera, A., Dunai, L. and Hartmann, W. M. “Human Interaural Time Difference Thresholds for Sine Tones: The High-Frequency Limit.”
    Journal of the Acoustical Society of America, 2013.
    View publication
  3. Bures, Z. and Marsalek, P. “On the Precision of Neural Computation with Interaural Level Differences in the Lateral Superior Olive.”
    Brain Research, 2013.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning Sound Localization with New Ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Pulkki, V. “Microphone Techniques and Directional Quality of Sound Reproduction.”
    Audio Engineering Society, 2002.
    View AES record
  6. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View AES record
  7. Toole, F. E. “Loudspeakers and Rooms for Stereophonic Sound Reproduction.”
    Audio Engineering Society 8th International Conference, 1990.
    View AES record

An earlier version of this essay was published in Audio Review, issue 487, June 2026. This Direct Sound Records Journal edition has been revised, expanded and technically updated for an international readership.

Blumlein - Stereo Recording

Why the Blumlein Pair Can Excel in Loudspeaker Playback

1024 575 Michelangelo

Stereo realism does not necessarily emerge from making an image as wide as possible. It emerges when the directional, tonal and reverberant cues in a recording support one another strongly enough to create a stable and believable acoustic space.

In the first article of this series, I explored how human hearing reconstructs a three-dimensional auditory world from differences in arrival time, level and spectral filtering at the two ears.

This second article moves from perception to production: how do different stereo microphone techniques encode spatial information, and why can the choice of array become especially important when a recording is reproduced through loudspeakers?

As a recording engineer and the founder of Direct Sound Records, my central objective is not simply to create an impressive stereo effect. It is to preserve the relationship between the musicians, the acoustic environment and the listener in a way that remains convincing during reproduction.

Among the many available stereo techniques, the Blumlein pair remains one of the most revealing—and one of the most demanding.

There Is No Universal “Best” Stereo Technique

Before considering Blumlein, an important distinction must be made: no stereo microphone technique wins in every situation.

The appropriate array depends on several factors:

  • the size and arrangement of the ensemble;
  • the acoustic quality of the room;
  • the desired relationship between direct and reverberant sound;
  • the required stereo width and localisation precision;
  • mono compatibility;
  • the intended playback system;
  • and the artistic purpose of the recording.

A spaced pair may be ideal when a broad sense of scale and low-frequency spaciousness is required. ORTF can provide an effective balance of width, localisation and ambience. Mid-Side offers control over stereo width after recording. XY can provide a stable image with strong mono compatibility.

Blumlein has its own strengths and limitations. Its value lies not in being universally superior, but in the particular way it connects direct sound, room ambience and coincident stereo geometry.

How Stereo Microphone Arrays Encode Space

Stereo microphone techniques can be broadly understood according to the cues they create between the left and right channels.

Spaced Pairs

In an AB arrangement, two microphones are separated physically. A sound arriving from one side will generally reach one microphone before the other, producing an interchannel time difference. Depending on the microphones and source position, there may also be a difference in level.

Spaced arrays can produce a broad and enveloping presentation. They can be particularly effective for large ensembles, organs, orchestras and situations in which the acoustic environment is an important part of the experience.

However, the time differences between channels can influence mono compatibility and may produce frequency-dependent reinforcement or cancellation when the channels are combined. Increasing microphone spacing can also weaken centre localisation if the geometry is not appropriate for the source and listening conditions.

These are design trade-offs, not proof that spaced recording is inherently defective.

Near-Coincident Arrays

Near-coincident techniques such as ORTF deliberately combine microphone spacing with directional microphone patterns.

ORTF uses two cardioid microphones separated by approximately 17 centimetres and angled 110 degrees apart. The resulting stereo image contains both interchannel timing and level differences.

This combination often produces greater spaciousness than a fully coincident cardioid pair while retaining more definite localisation than a widely spaced AB array. It is one reason ORTF has remained a widely used technique for classical music, ensembles and location recording.

Describing ORTF simply as “phasey” overlooks the fact that its timing differences are intentional components of its spatial design.

Coincident Arrays

In a coincident array, the microphone capsules are positioned as close as physically possible to the same acoustic point.

Because direct sound reaches the two capsules at almost the same time, stereo direction is created primarily through differences in level and polarity rather than substantial arrival-time differences.

XY, Mid-Side and Blumlein are all coincident techniques, although they use different polar patterns and encode the surrounding sound field differently.

The coincident geometry generally provides predictable mono compatibility and reduces the possibility of time-delay-related cancellations when the channels are summed. It can also create a clearly defined centre image and stable localisation within the normal listening area.

Real microphones are not mathematically perfect points, however. Capsule dimensions, vertical displacement, polar-pattern differences and off-axis response mean that no practical array is perfectly coincident or perfectly phase coherent at every frequency.

What Loudspeaker Playback Changes

Headphone and loudspeaker reproduction deliver stereo signals to the listener in fundamentally different ways.

With conventional headphones, the left channel is delivered predominantly to the left ear and the right channel to the right ear. With two loudspeakers, each loudspeaker reaches both ears.

The left ear therefore hears:

  • the left loudspeaker directly;
  • the right loudspeaker through an additional acoustic path;
  • and reflections from the listening room.

The right ear receives the corresponding combination from the opposite side.

This acoustic crosstalk is not an accidental failure of stereo. It is part of conventional two-channel loudspeaker reproduction. The brain uses the resulting combination of timing, level and spectral cues to perceive phantom images between and sometimes beyond the loudspeakers.

However, the reconstruction is sensitive to geometry. Moving away from the central listening position changes the relative distances from the two loudspeakers and therefore changes the arrival-time and level relationships at the ears. The phantom image tends to shift towards the nearer speaker.

The loudspeakers, room and listener must therefore be considered as one reproduction system. A recording does not carry an independent three-dimensional space that remains unchanged under every playback condition.

Enter the Blumlein Pair

The Blumlein pair uses two figure-of-eight microphones mounted coincidently and angled 90 degrees apart.

The technique is associated with Alan Dower Blumlein, whose pioneering 1931 patent described fundamental principles of stereophonic recording and reproduction.

In a correctly arranged Blumlein pair, the microphone diaphragms occupy almost the same acoustic point. Directional information is encoded primarily through the different levels and polarities produced by the two figure-of-eight patterns.

A figure-of-eight microphone is equally sensitive to sound arriving from the front and rear, while strongly rejecting sound arriving from its sides. Consequently, the array captures both the performance in front of the microphones and a substantial amount of acoustic information from behind them.

This is a defining characteristic of Blumlein—not a minor detail.

Why Blumlein Can Sound So Convincing

Coincident Timing for Direct Sound

Because the two capsules are positioned at approximately the same point, direct sound from an instrument reaches both microphones almost simultaneously.

This minimises interchannel arrival-time differences introduced by the microphone spacing itself. The stereo image is produced predominantly through level and polarity relationships.

For loudspeaker reproduction, this can create a precise centre image and clearly organised lateral positions, particularly when the ensemble and array are positioned carefully.

Strong Mono Compatibility

When the two channels of a coincident recording are combined, corresponding direct sounds normally align more predictably than they do in a widely spaced array.

This does not mean that every part of a Blumlein recording will combine perfectly. Reflections arrive from many directions and at many times, while real microphones have tolerances and frequency-dependent polar behaviour.

Nevertheless, the coincident geometry generally gives Blumlein excellent mono compatibility compared with arrays that rely heavily on microphone spacing.

Natural Integration of the Room

The rear lobes of the figure-of-eight microphones capture reverberant energy and sound arriving from behind the array.

In a good acoustic environment, this can create a remarkably integrated sense of depth. The room does not feel like a synthetic effect added behind the musicians. It becomes part of the same spatial event.

The direct sound establishes the performers, while the reflected energy communicates the dimensions, character and decay of the venue.

When those relationships are balanced correctly, the listener may perceive not merely a wide line between two loudspeakers, but a coherent acoustic scene extending behind and around the performers.

Spatial Information Without Microphone Spacing

Blumlein can generate a substantial stereo image without separating the microphones horizontally.

This is particularly attractive when the engineer wants clear directional information while minimising time-of-arrival differences between channels.

The result can feel cohesive because the direct sound and room information are captured from a single acoustic viewpoint.

What Phase Coherence Really Means Here

The expression phase coherence is often used loosely in audio. In the context of a coincident stereo array, it is more useful to speak about the consistency of interchannel timing relationships.

Blumlein does not remove phase from a recording. Every acoustic event contains complex phase relationships, and every room creates reflections with different delays, levels and spectra.

What the array minimises is the additional time difference that would otherwise be introduced by placing the two microphones at separate locations.

This distinction matters.

A Blumlein recording can still contain:

  • phase differences created by room reflections;
  • microphone-response differences;
  • polarity differences inherent in the figure-of-eight geometry;
  • and complex interference between direct and reverberant sound.

Its strength is not “perfect phase purity.” Its strength is that both channels observe the direct acoustic event from approximately the same point in space.

Why Blumlein Is Also Demanding

The characteristics that make Blumlein revealing also make it unforgiving.

The Room Must Deserve to Be Recorded

Because figure-of-eight microphones capture strongly from both front and rear, an unattractive room will not politely disappear.

Flutter echoes, mechanical noise, audience movement, heating systems and poorly controlled reflections can become prominent parts of the recording.

Blumlein works best when the acoustic environment contributes positively to the performance.

Placement Is Critical

The balance between ensemble width, direct sound and reverberation depends strongly on the distance and orientation of the array.

Positioning the microphones too close may produce an image that is excessively wide or exclude important sources from the useful recording angle. Placing them too far away may allow reverberation to dominate and reduce clarity.

Small movements can significantly change the result. This is why Blumlein rewards careful listening and deliberate placement rather than formula alone.

Rear Sound Is Part of the Recording

The rear lobes do not distinguish between beautiful reverberation and unwanted noise.

Musicians, audience members, equipment and reflective surfaces behind the microphones all become part of the captured field. The engineer must therefore consider the entire environment around the array, not only what lies in front of it.

The Listening Position Still Matters

Blumlein does not eliminate the limitations of two-loudspeaker stereo.

A listener moving significantly away from the central position will still experience changes in timing and level from the loudspeakers, and the stereo image will shift accordingly.

Coincident recording can provide a coherent source signal, but it cannot make conventional stereo reproduction independent of loudspeaker and listener geometry.

When I Choose Blumlein

In my work, Blumlein becomes especially compelling when:

  • the musicians are acoustically balanced in the room;
  • the venue has a distinctive and musically valuable acoustic;
  • the ensemble fits naturally within the array’s useful recording angle;
  • the intention is to preserve a complete performance rather than construct one later;
  • and loudspeaker playback is an important reference.

I would not choose it automatically when the room is problematic, when strong isolation is required, when sources must be balanced independently, or when the ensemble geometry demands a wider or more flexible array.

In those situations, ORTF, AB, Mid-Side, XY, supplementary microphones or a hybrid approach may be more appropriate.

The technique should serve the acoustic event—not the engineer’s ideology.

Spatial Width Is Not the Same as Realism

A recording can create an enormous stereo image and still feel artificial.

Width may be produced by long interchannel delays, decorrelation, processing or exaggerated ambience. These effects can be exciting, but they do not necessarily communicate a believable relationship between performers and space.

Blumlein offers a different proposition. Its most successful recordings do not merely place sounds from left to right. They establish a unified perspective from which the listener can infer:

  • where the musicians are positioned;
  • how far away they appear;
  • how the room surrounds them;
  • and how direct and reflected sound belong to the same event.

This is why the technique can feel less like an audio effect and more like a view into an acoustic performance.

A Reference, Not a Religion

The Blumlein pair deserves its status as one of the foundational stereo microphone techniques. Its coincident geometry, figure-of-eight patterns and integration of direct and reverberant sound can produce extraordinary depth, localisation and spatial coherence.

But the strongest case for Blumlein does not require dismissing other approaches.

AB can communicate scale and spaciousness that a coincident pair may not reproduce in the same way. ORTF can offer a persuasive compromise between width and localisation. Mid-Side provides valuable control after recording. XY can be practical, focused and robust.

The real achievement lies in understanding how each technique encodes space—and choosing the one whose compromises best serve the music, venue and intended reproduction system.

For Direct Sound Records, the objective is not to manufacture an impressive stereo image. It is to preserve the acoustic relationships that make a performance feel present, intelligible and emotionally credible.

When the room, musicians and microphone position align, Blumlein can be one of the most direct ways of achieving that objective.

References and Further Reading

  1. Blumlein, A. D. Improvements in and Relating to Sound-Transmission, Sound-Recording and Sound-Reproducing Systems. British Patent GB394325A, filed 1931 and published 1933.
    View patent
  2. Eargle, J. M. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View AES record
  3. Ceoen, C. “Basic Stereo Microphone Perspectives—A Review.”
    Journal of the Audio Engineering Society, 1985.
    View AES record
  4. Toole, F. E. “Loudspeakers and Rooms for Stereophonic Sound Reproduction.”
    Audio Engineering Society 8th International Conference, 1990.
    View AES record
  5. Kendall, G. S. “The Effects of Interaural Crosstalk on Stereo Reproduction and Minimizing Interaural Crosstalk in Nearfield Monitoring by the Use of a Physical Barrier: Part 1.”
    Audio Engineering Society 81st Convention, 1986.
    View AES record
  6. Lee, H. and Gribben, C. “On the Optimum Listening Position and Listening Angle in a Two-Channel Stereophonic Reproduction System.”
    Audio Engineering Society.
    View AES record

An earlier version of this article was published on LinkedIn. This Direct Sound Records Journal edition has been revised, expanded and technically updated, with additional context and references.

How We Hear in 3D – The Neuroscience Behind Stereo Perception

1024 1024 Michelangelo

“Phase is not the sound — it is the space between sounds.” Taken as a metaphor, this captures something fundamental about human hearing: the brain does not process the signals arriving at our two ears independently. It continuously compares them, using extraordinarily small differences in time, level and spectral balance to reconstruct the position of sound around us.

In high-end audio, we often concentrate on equipment, formats, resolution and frequency response. Yet beneath every recording and every playback system lies a more fundamental question: how does the human auditory system transform two incoming signals into a convincing three-dimensional world?

Stereo reproduction works not because two loudspeakers recreate the original sound field perfectly, but because they provide the auditory system with enough carefully organised information to create a plausible spatial scene. Understanding that process changes how we think about microphone placement, phase relationships, room acoustics and the meaning of realism in recording.

Hearing Is a Reconstruction

Sound arriving at the ears does not contain a ready-made map of the space around us. The auditory system must infer the direction, distance and environment of a sound source from a collection of acoustic clues.

For horizontal localisation, the most important binaural cues are differences in arrival time and sound level between the two ears. For elevation and front-to-back discrimination, the auditory system also relies heavily on direction-dependent spectral filtering created by the listener’s head, torso and outer ears.

These mechanisms operate together. They should not be understood as three entirely separate systems, nor as rigid frequency zones with precise boundaries. Natural sounds are usually broadband, reflections complicate the incoming signals, and the brain integrates multiple cues over time.

The Principal Cues to Spatial Hearing

1. Interaural Time Differences

When a sound source is positioned to one side of the listener, the sound normally reaches the nearer ear slightly before reaching the farther ear. This difference is known as the interaural time difference, or ITD.

The delays involved are extremely small—often measured in microseconds—but the auditory system is remarkably sensitive to them. For low-frequency sounds, neural activity can follow the temporal structure of the waveform closely enough for the brain to compare the timing received at the two ears.

ITD sensitivity to the fine structure of pure tones becomes progressively less effective as frequency rises, with human sensitivity deteriorating sharply around the region of approximately 1.4 to 1.5 kHz. This should be treated as a broad transition rather than a universal dividing line. High-frequency sounds can still carry timing information through changes in their amplitude envelope.

2. Interaural Level Differences

A sound arriving from one side will also tend to be louder at the nearer ear. At shorter wavelengths, the head obstructs part of the sound travelling towards the farther ear, producing an acoustic shadow. The resulting difference in level is known as the interaural level difference, or ILD.

Level differences generally become more pronounced at higher frequencies because shorter wavelengths are more strongly affected by the head. However, ITD and ILD do not simply exchange responsibility at one exact frequency. For complex sounds, the auditory system can combine timing and level information across several frequency regions.

The relative importance of these cues also changes with the sound itself, its distance, the surrounding reflections and the listener’s hearing.

3. Spectral Cues and the Head-Related Transfer Function

Time and level differences are especially useful for identifying whether a sound is located towards the left or right. They are less able, by themselves, to resolve whether a source is above, below, in front of or behind the listener.

For this, the complex shape of the outer ear becomes essential. The folds of the pinna, together with the head and upper body, alter the spectrum of incoming sound in a direction-dependent way. Some frequencies are reinforced, while others are attenuated or notched.

The complete acoustic transformation between a sound source and the listener’s ears is described by the head-related transfer function, or HRTF.

Because every person’s anatomy is different, HRTFs are individual. The brain gradually learns the spectral patterns associated with particular directions. Experiments in which the shape of the outer ear was temporarily altered have shown that localisation initially becomes less accurate, particularly for elevation and front-to-back judgements, but can improve again as listeners adapt to the modified cues.

Where Phase Fits In

The word phase is used in several related but distinct ways in audio, which can easily create confusion.

At low frequencies, a difference in arrival time between the ears can also be described as an interaural phase difference for a periodic waveform. In this context, phase difference is one of the ways the auditory system obtains spatial information.

Neurons within the auditory brainstem, including those associated with the medial superior olive, are specialised for processing extremely small interaural timing differences. Their responses contribute to the neural representation of horizontal sound direction.

However, it would be too simple to conclude that phase is used only to locate a sound and never contributes to what that sound is. The timing structure of a waveform—often described as its temporal fine structure—also contributes to aspects of pitch perception, auditory masking and the separation of sounds in complex listening environments.

The more useful distinction for recording engineers is therefore not between “phase” and “sound identity,” but between the different roles that temporal and phase relationships can play.

Within a stereo recording, relationships between the two channels may affect:

  • the apparent position of a phantom image;
  • the perceived width and stability of the soundstage;
  • the impression of depth and surrounding ambience;
  • the result when the recording is reproduced in mono;
  • frequency-response changes caused by constructive and destructive interference.

Phase is therefore neither an isolated technical curiosity nor a universal explanation for every spatial quality. It is one component within a larger system of timing, level, spectrum, reflection and playback interaction.

Why This Matters for Stereo Recording

Stereo microphone techniques create spatial information in different ways.

A spaced pair such as AB introduces arrival-time differences between the microphones and may also produce level differences. A near-coincident configuration such as ORTF deliberately combines microphone spacing with directional level differences.

Coincident techniques such as XY and Mid-Side minimise the arrival-time difference between microphones and create direction mainly through differences in level. The Blumlein pair, using two coincident figure-of-eight microphones, also derives its directional information from the polar patterns and polarity relationships of the two channels while capturing substantial information from the surrounding acoustic environment.

None of these methods is automatically natural or unnatural in every situation. Each encodes the original acoustic event differently, and each interacts differently with loudspeakers, headphones, room reflections and listener position.

This is particularly important because conventional stereo loudspeaker reproduction does not send the left channel exclusively to the left ear or the right channel exclusively to the right ear. Each loudspeaker reaches both ears, introducing additional timing, level and spectral interactions. The listening room then adds its own reflections.

The task of the recording engineer is therefore not merely to create a wide image. It is to create interchannel relationships that remain meaningful when reproduced through the intended playback system.

From Spatial Effect to Spatial Credibility

A stereo recording can sound spectacularly wide while still producing unstable localisation, exaggerated scale or an uncertain centre image. Conversely, a narrower presentation may feel more convincing because its timing, level and reverberant cues form a more internally consistent spatial picture.

This suggests a more useful question than simply asking whether a recording sounds spacious:

Do the spatial cues reproduced by the system support one another strongly enough for the brain to construct a stable and believable acoustic scene?

Phase coherence is part of that question, but it should not be treated as a single measurement that determines realism on its own. Microphone polar pattern, spacing, angle, source distance, direct-to-reverberant ratio, loudspeaker placement and room acoustics all contribute to the final perception.

What Comes Next

In the next article, I will examine how AB, ORTF, XY, Mid-Side and Blumlein recording techniques encode spatial information differently, and why a technique that produces impressive width does not necessarily produce the most credible depth or localisation.

I will also explore the relationship between coincident microphone techniques, loudspeaker reproduction, room acoustics and the preservation of stable interchannel relationships.

The central principle is simple: recording technology should not be considered separately from human perception. The microphone arrangement, recording space and playback format are parts of one perceptual chain.

References and Further Reading

  1. Brughera, A., Dunai, L. and Hartmann, W. M. “Human interaural time difference thresholds for sine tones: the high-frequency limit.” Journal of the Acoustical Society of America, 2013.
    View publication
  2. Wightman, F. L. and Kistler, D. J. “The dominant role of low-frequency interaural time differences in sound localization.”
    Journal of the Acoustical Society of America, 1992.
    View publication
  3. Salminen, N. H., Tiitinen, H., Yrttiaho, S. and May, P. J. C. “The neural code for interaural time difference in human auditory cortex.”
    Journal of the Acoustical Society of America, 2010.
    View publication
  4. Hofman, P. M., Van Riswick, J. G. A. and Van Opstal, A. J. “Relearning sound localization with new ears.”
    Nature Neuroscience, 1998.
    View publication
  5. Moore, B. C. J. “The role of temporal fine structure processing in pitch perception, masking, and speech perception for normal-hearing and hearing-impaired people.”
    Journal of the Association for Research in Otolaryngology, 2008.
    View publication
  6. Eargle, J. “An Overview of Stereo Recording Techniques for Popular Music.”
    Journal of the Audio Engineering Society, 1985.
    View publication

An earlier version of this article was published on LinkedIn. This Direct Sound Records Journal edition has been revised, expanded and technically updated, with additional context and scientific references.