SWORD LAB / DIGITAL AUDIO

Digital audio from first principles

Sampling, Nyquist, bit depth, quantisation, dither, jitter, latency and clocking

Digital audio does not turn a waveform into a staircase that a loudspeaker “plays step by step”. A band-limited continuous signal can be represented by discrete samples and reconstructed continuously. Understanding the sampling theorem, quantisation and implementation limits removes many persistent audio myths.

1. Sampling

PCM represents signal amplitude at discrete time instants. Under the Nyquist-Shannon sampling theorem, a signal band-limited below half the sampling frequency can be reconstructed from its samples in principle. The DAC output is not inherently a staircase after proper reconstruction filtering.

fsignal,max < fs2

Symbols: fsignal,max is the highest frequency present in the band-limited input and fs sampling frequency. In practice, an analogue anti-alias filter needs transition bandwidth, so useful signal bandwidth must remain below the theoretical Nyquist limit.

2. Why 44.1 kHz?

44.1 kHz provides a Nyquist frequency of 22.05 kHz, allowing the conventional audio band to fit below Nyquist with a transition region for anti-alias/reconstruction filtering. The rate also has historical roots in early digital-audio/video recording systems. 48 kHz became common in professional video and broadcast workflows.

3. Aliasing

Energy above Nyquist folds into the baseband if it reaches the sampler unfiltered. For a component f, aliases appear at frequencies related to |f-k·fₛ|. Because aliasing creates false in-band components, analogue anti-alias filtering and/or oversampling are fundamental parts of converters.

4. Bit depth and quantisation

Amplitude must be quantised to a finite set of codes. For an ideal N-bit uniform quantiser and a full-scale sine with uniformly distributed quantisation error, the classic signal-to-quantisation-noise ratio is:

SNR ≈ 6.02·N + 1.76 dB

Symbols: SNR is ideal quantisation signal-to-noise ratio and N bit depth. The estimate assumes an ideal N-bit quantiser driven by a full-scale sine wave; each added bit theoretically improves SNR by about 6.02 dB.

The 1.76 dB term is not universal; it follows specifically from a full-scale sine and the assumptions behind the ideal quantiser model. The derivation compares sine RMS power with error variance Q²/12.

5. Dither

At very low signal levels, undithered quantisation error becomes correlated with the signal and behaves as distortion. Adding carefully chosen noise randomises the error. TPDF (Triangular Probability Density Function) dither is widely used because it removes signal-correlated quantisation distortion and noise modulation under the standard model. The price is a predictable increase in noise floor - roughly 4.77 dB relative to the undithered uniform-error variance in the common comparison.

6. Noise shaping

Noise shaping uses feedback to redistribute quantisation-noise energy toward frequency regions where it is less audible, while preserving overall information. It is especially useful when reducing to 16-bit delivery formats. It should not be confused with removing noise; the energy is spectrally rearranged.

7. Oversampling

Oversampling moves converter image/alias constraints farther from the audio band, allowing gentler analogue filters and sophisticated digital filtering. Internal oversampling is also used in nonlinear DSP and true-peak metering because processing can create or reveal intersample peaks that are invisible in the original sample values.

8. Jitter

Jitter is timing uncertainty of sampling/clock edges. A first-order error relation is:

ΔU ≈ dudt · Δt

Symbols: ΔU is instantaneous amplitude error caused by timing jitter, du/dt the analogue signal slope and Δt sampling-time error. The same timing error creates a larger voltage error where the signal changes more steeply.

Thus timing error matters more when the waveform slope is steep - typically at higher audio frequencies and higher amplitudes. Modern competent converters usually suppress clock jitter sufficiently that other analogue limitations dominate, but poor clock distribution can still create measurable sidebands and noise.

9. 44.1 vs 48 vs 96 vs 192 kHz

Higher sample rates ease transition-band filtering and can reduce latency for a given sample buffer, and they are useful in some DSP/nonlinear processing. They also multiply storage, bandwidth and CPU cost and do not automatically improve audible quality in a properly designed playback chain. Choose a rate for workflow and processing requirements, not numerology.

10. Why 24 bit is useful in production

24-bit acquisition and processing provide enormous numerical margin relative to practical analogue noise floors. This lets engineers record with conservative headroom without sacrificing meaningful resolution. The real-world dynamic range of converters is far below the theoretical 144 dB, but still comfortably exceeds most acoustic environments.

11. Integer and floating-point processing

Fixed-point formats have a hard full-scale ceiling. IEEE-754 32-bit floating point has about 24 bits of significand precision but an exponent range giving a numerical dynamic range far above 1500 dB. Internal floating-point mix buses can therefore represent signals above 0 dBFS numerically without immediate integer clipping, provided they are brought back within range before a fixed-point output, DAC or file conversion.

IMPORTANT

Floating point is not “infinite headroom”. Plugins can have internal nonlinear limits, and the final fixed-point/DAC stage still has a physical ceiling.

12. Intersample peaks and true peak

The continuous reconstructed waveform can exceed the highest stored sample value. These intersample peaks matter in mastering and conversion. ITU-R BS.1770-based true-peak metering uses oversampled estimation - commonly 4× or more in implementations - to approximate the reconstructed peak rather than simply reading sample peaks.

13. Latency and buffer size

Buffer latency scales approximately with samples divided by sample rate. A 256-sample one-way buffer at 48 kHz represents about 5.33 ms before adding driver, conversion, DSP and operating-system delays. Round-trip performance must be measured as an end-to-end budget, not inferred from one buffer setting.

14. Sample-rate conversion quality

Sample-rate conversion is filtering plus resampling. Good asynchronous SRC can provide extremely low distortion and aliasing; poor conversion can produce passband ripple, imaging or alias products. Repeated unnecessary conversions should be avoided, but modern high-quality SRC is not inherently destructive in an audible sense.

15. Word clock and multiple digital devices

Synchronous digital links require a common timing relationship. Devices can derive clock from the incoming stream or use a dedicated clock network depending on protocol. Multiple unsynchronised sources require asynchronous sample-rate conversion or buffering strategies; otherwise clicks, dropouts or sample slips occur.

16. When is dither required?

Dither is relevant when reducing bit depth or otherwise quantising to a coarser fixed-point representation. It is generally applied once, at the final quantisation stage. Adding dither repeatedly at every processing step only raises noise without benefit.

17. Digital silence and analogue noise

A stream of exact digital zeros is mathematically silent, but a real DAC output still contains analogue noise, converter noise, power-supply noise and environmental interference. Digital dynamic-range numbers and actual acoustic noise at the listener are therefore different system metrics.

Sources and professional background

  • Nyquist-Shannon sampling theorem and standard DSP texts
  • IEEE 754 floating-point standard
  • ITU-R BS.1770 - loudness and true-peak measurement
  • Converter and sample-rate-conversion engineering literature

Professional and legal notice

I prepare the technical descriptions, calculations, examples, diagrams and other information published in SWORD LAB for educational and informational purposes. When compiling the material I aim for technical accuracy, correct presentation of the underlying relationships and careful use of the available professional knowledge.

Nevertheless, the information may contain inaccuracies, errors or simplifications that cannot be applied unchanged to a specific system or environment. The calculations and engineering examples are generally based on stated or implicit assumptions. Real systems are also affected by the actual parameters of the equipment, system topology, environmental conditions, measurement method, installation practice, applicable standards, legislation and manufacturer requirements.

The material I publish does not constitute design documentation, an expert opinion, an installation instruction, a safety instruction or individual professional advice. It does not replace manufacturer documentation, current regulations and standards, or - where required - the examination, measurement or design work of a suitably qualified and authorised professional.

Technical standards, product data and technologies change over time. Before design, installation, measurement, operation, repair or equipment selection, I therefore recommend checking the current primary and authoritative sources.

I make every reasonable effort to prepare the material carefully; however, to the extent permitted by applicable law, I accept no liability for direct or indirect damage, loss, malfunction or interruption resulting from the use, misinterpretation or incorrect application of information published on this site, or from interventions carried out on that basis.

The purpose of SWORD LAB is to help explain engineering relationships and the physical and technical processes behind sound reinforcement, electroacoustics, digital audio and DJ technology. It is not intended to replace on-site investigation, measurement or engineering design of a specific system.

© 2026 SWORD · All rights reserved. Copyright notice