WAV, FLAC, MP3, AAC and CD
Containers, codecs, lossless and lossy compression - what actually happens to the music?
File extensions are often treated as shorthand for “quality”, but they refer to different layers of the signal chain. WAV is usually a container for PCM, FLAC is lossless compression, MP3 and AAC are perceptual lossy codecs, and CD-DA is a physical/digital delivery format. A useful comparison starts by separating container, codec, sample format, bitrate and mastering.
1. PCM
Pulse Code Modulation stores or transports quantised sample values. Sample rate determines time sampling; bit depth determines the number of available amplitude codes. PCM by itself is not a file extension and does not imply a particular mastering quality.
2. WAV
WAV is primarily a RIFF container. It commonly contains uncompressed linear PCM, but the container can hold other formats as well. A WAV file is therefore not automatically “better” than another file; what matters is the encoded audio and its provenance.
3. FLAC
FLAC compresses PCM losslessly. Decoding a valid FLAC produces the same PCM sample values that were encoded. If a WAV and FLAC originate from the same PCM master, the decoded audio is bit-identical. The difference is storage efficiency and metadata capability, not sound quality.
4. MP3
MP3 is a perceptual lossy codec. It analyses signal energy in time/frequency regions and allocates bits according to psychoacoustic masking models. MP3 uses a hybrid 32-band polyphase filterbank plus MDCT. At lower bitrates, quantisation and coding decisions can produce pre-echo, swishing, stereo-image changes or high-frequency simplification.
5. AAC
AAC is a newer family of perceptual codecs with more flexible transform tools than MP3. It is predominantly MDCT-based and can switch between long and short windows more effectively around transients, improving time/frequency allocation and reducing pre-echo at a given bitrate. AAC is not universally “transparent”; implementation, profile and bitrate still matter.
6. Bitrate
Bitrate is data per unit time, not direct audio quality. For lossless codecs, bitrate varies with signal compressibility while decoded PCM remains identical. For lossy codecs, higher bitrate generally permits fewer coding compromises, but codec generation and encoder quality also matter.
7. VBR and CBR
Constant Bit Rate allocates roughly the same rate over time; Variable Bit Rate can spend more bits on complex passages and fewer on easy ones. For quality-oriented encoding, well-implemented VBR is often more efficient because musical complexity is not constant.
8. Transcoding
Lossless-to-lossless conversion preserves decoded PCM when performed correctly. Lossy-to-lossy transcoding cannot restore discarded information and can compound artefacts because the second encoder sees already modified material. Converting an MP3 to FLAC only creates a larger lossless copy of the decoded MP3; it does not recreate the original master.
9. ReplayGain and loudness normalisation
Loudness normalisation adjusts playback gain from measured loudness metadata; it is not the same as dynamic-range compression. ReplayGain and streaming loudness systems aim to reduce large programme-to-programme loudness differences while leaving the audio dynamics intact unless limiting is required to prevent peaks.
10. Can you hear the difference?
The only reliable way to answer a subtle codec question is controlled, level-matched, blind comparison. Expectation bias is strong in audio. For a 50/50 ABX task, the one-sided chance probability of at least 12 correct responses in 16 independent trials is:
Symbols: X is the number of correct responses, C(16,i) the binomial coefficient and 0.5 the chance success probability on each trial. About 3.84% is below a 5% significance threshold, but the result is meaningful only with level matching, no identity cues and a pre-defined test protocol.
11. Psychoacoustic coding in more depth
Lossy coding removes or coarsely represents components judged less perceptually important because of simultaneous and temporal masking. The encoder uses a model to estimate masking thresholds, transforms the signal into frequency-domain coefficients and allocates quantisation noise where it should be less audible. Artefacts arise when the model or bit budget is insufficient.
12. Pre-echo
Transform coding represents blocks of time. A sharp transient inside a long transform window can spread quantisation noise backward in time, creating audible pre-echo before the transient. Short-block switching reduces the temporal spread. AAC’s more flexible window structure is one reason it tends to perform better than MP3 on difficult transients at comparable low/moderate bitrates.
13. Joint stereo
Two mechanisms must be separated. Mid/Side stereo can be represented as:
L = M+S, R = M-S
Symbols: L and R are the left and right channels, M their common middle component and S their difference. The inverse equations show that this transform is reversible by itself. Intensity stereo, in contrast, is perceptual and can simplify high-frequency spatial information at aggressive settings.
14. CD error correction
Red Book CD-DA uses 44.1 kHz, 16-bit stereo PCM at the logical audio layer. The physical disc adds substantial coding and error resilience. CIRC (Cross-Interleaved Reed-Solomon Code) combines error correction with interleaving so burst defects such as scratches can often be reconstructed. EFM (Eight-to-Fourteen Modulation) maps data to run-length-limited channel symbols that support reliable optical clock recovery.
Error concealment is used when corruption exceeds correction capability; therefore “the CD always returns bit-perfect audio no matter how damaged” is not a correct statement.
15. Metadata and file management for DJs
For a working library, technical quality includes reliable metadata, consistent naming, verified backups and known provenance. Keep a lossless archival master where possible, derive delivery copies from that master, and avoid repeated lossy transcodes. Store cue/beat-grid information in a way that can be backed up independently of one application database.
16. Normalisation versus compression
Peak or loudness normalisation applies gain; it does not change internal dynamics unless a ceiling is exceeded and limiting/clipping follows. Dynamic compression changes gain continuously as a function of signal level and time constants. Mixing the terms obscures what processing actually occurred.
17. How to test a codec correctly
- Start from the same source master.
- Decode both candidates through the same playback path.
- Match level precisely.
- Use blind or ABX switching.
- Repeat enough trials for statistical meaning.
- Use difficult programme material as well as normal music.
A sighted comparison of files with different mastering or level is not a codec test.
Sources and professional background
- ISO/IEC MPEG audio standards and codec literature
- FLAC format documentation
- IEC 60908 / Red Book CD-DA technical background
- Psychoacoustic masking, MDCT and ABX methodology literature
Professional and legal notice
I prepare the technical descriptions, calculations, examples, diagrams and other information published in SWORD LAB for educational and informational purposes. When compiling the material I aim for technical accuracy, correct presentation of the underlying relationships and careful use of the available professional knowledge.
Nevertheless, the information may contain inaccuracies, errors or simplifications that cannot be applied unchanged to a specific system or environment. The calculations and engineering examples are generally based on stated or implicit assumptions. Real systems are also affected by the actual parameters of the equipment, system topology, environmental conditions, measurement method, installation practice, applicable standards, legislation and manufacturer requirements.
The material I publish does not constitute design documentation, an expert opinion, an installation instruction, a safety instruction or individual professional advice. It does not replace manufacturer documentation, current regulations and standards, or - where required - the examination, measurement or design work of a suitably qualified and authorised professional.
Technical standards, product data and technologies change over time. Before design, installation, measurement, operation, repair or equipment selection, I therefore recommend checking the current primary and authoritative sources.
I make every reasonable effort to prepare the material carefully; however, to the extent permitted by applicable law, I accept no liability for direct or indirect damage, loss, malfunction or interruption resulting from the use, misinterpretation or incorrect application of information published on this site, or from interventions carried out on that basis.
The purpose of SWORD LAB is to help explain engineering relationships and the physical and technical processes behind sound reinforcement, electroacoustics, digital audio and DJ technology. It is not intended to replace on-site investigation, measurement or engineering design of a specific system.
