What do we actually hear?
From vibration to auditory perception - the physics and physiology of human hearing, and how hearing changes with age
The most important “terminal device” in sound engineering is not the loudspeaker but the human listener. At the end of every signal chain a physical vibration becomes neural information and, finally, conscious auditory perception. Understanding that chain prevents us from confusing frequency with pitch, sound pressure with loudness, or a linear measurement with the nonlinear behaviour of hearing.
1. Everything begins with vibration
A vibration is a time-varying state of a physical system around an equilibrium. The simplest useful model is the mass-spring-damper system. An ideal harmonic displacement can be written as:
Symbols: x(t) is instantaneous displacement, A amplitude, f frequency in hertz, t time in seconds, φ initial phase in radians, and T period in seconds. The first expression describes ideal harmonic motion; the second states that one period is the reciprocal of frequency.
Here A is amplitude, f frequency, t time and φ initial phase. A 100 Hz vibration completes 100 cycles per second, so one period is 10 ms; at 1 kHz it is 1 ms and at 10 kHz only 0.1 ms.
Amplitude, energy and damping
Amplitude describes excursion, not energy by itself. Energy also depends on mass, stiffness, velocity and other system parameters. Real systems lose energy through mechanical friction, air resistance, electrical resistance and structural losses, so a free vibration decays with time.
Natural frequency and resonance
Elastic systems have natural frequencies. When excitation approaches one of them, response can rise substantially. The same physical idea appears in loudspeaker suspensions, bass-reflex alignments, room modes, microphone diaphragms and the ear canal.
Resonance is not automatically a fault. It is deliberately used, for example, in bass-reflex tuning. It becomes problematic when it is uncontrolled, has excessive Q, creates unwanted gain in the passband or produces long decay.
The linear mass-spring-damper model is:
H(s) = X(s)F(s) = 1m·s² + c·s + k
Symbols: x is displacement, dx/dt velocity, d²x/dt² acceleration, m mass, c damping coefficient, k spring stiffness and F(t) external force. H(s) is the force-to-displacement transfer function, X(s) and F(s) are the corresponding Laplace transforms, and s is complex frequency. The three terms in the differential equation represent inertia, damping and restoring force.
Its ideal natural angular frequency is ω₀ = √(k/m), while Q = √(mk)/c. The same system language is useful for loudspeakers, the tympanic membrane and many acoustic resonators.
2. When does vibration become sound?
A vibrating object only becomes audible to us when mechanical energy reaches the hearing system through a medium. In air, sound propagates predominantly as a longitudinal pressure wave: air particles oscillate locally while pressure and density variations travel through the medium. In vacuum there is no material medium to carry a conventional acoustic pressure wave. In solids, both longitudinal and transverse elastic waves may occur.
Speed of sound
Sound speed depends on the medium. In air near room temperature it is about 343 m/s. A useful first-order approximation between 0 and 30 °C is:
Symbols: in this equation c is sound speed in metres per second and T°C is air temperature in degrees Celsius. It is an approximation for dry air over ordinary temperature ranges.
Humidity and gas composition introduce additional corrections.
Wavelength
Symbols: λ is wavelength in metres, c sound speed in metres per second and f frequency in hertz. At a fixed sound speed, higher frequency means shorter wavelength.
| Frequency | Wavelength at 343 m/s | Practical meaning |
|---|---|---|
| 20 Hz | 17.15 m | longer than many rooms |
| 100 Hz | 3.43 m | comparable with room dimensions |
| 1 kHz | 34.3 cm | midrange scale |
| 10 kHz | 3.43 cm | head and pinna geometry strongly matter |
| 20 kHz | 1.72 cm | strong directivity and geometric effects |
3. Frequency is not simply “low” or “high”
For a pure sine wave, frequency and perceived pitch are closely related, but musical sounds are not pure sines. They contain a fundamental, harmonics, noise components, transients and a time envelope. Two sounds can share the same fundamental and still have completely different timbre.
The auditory system can also infer periodicity. In the missing fundamental phenomenon, a 100 Hz pitch may still be perceived even if the 100 Hz spectral component is weak or absent, provided that harmonics at 200, 300, 400 Hz and so on establish the periodic pattern.
“Humans hear exactly from 20 Hz to 20 kHz.” No. This is a conventional nominal range. Thresholds are strongly frequency-, level-, age- and person-dependent, with no sharp wall at either end.
4. What do we measure when we measure sound pressure?
A microphone senses small alternating pressure variations superimposed on atmospheric pressure. Sound-pressure level is based on the RMS pressure:
Symbols: Lp is sound-pressure level in decibels, prms measured RMS sound pressure and p0 the 20 µPa reference pressure in air. The factor 20 appears because acoustic power is proportional to pressure squared.
0 dB SPL therefore does not mean “no sound”; it is the chosen reference, approximately close to the threshold of a young healthy listener around 1 kHz. Pressure is a physical quantity; loudness is a perceptual result and does not track pressure linearly.
5. The ear is not a microphone - it is an active biological system
The auditory periphery is conventionally divided into outer, middle and inner ear. These structures do much more than relay vibration: they filter, transform impedance and encode frequency-dependent information.
Outer ear
The pinna and ear canal provide direction- and frequency-dependent filtering. Ear-canal resonance contributes several dB of gain in the few-kilohertz region, important for speech sensitivity and localisation.
Tympanic membrane and middle ear
The malleus, incus and stapes transmit tympanic motion to the oval window. A central function is impedance matching between air and cochlear fluid. The effective tympanic area is roughly 55 mm² and the stapes footplate about 3.2 mm², giving a pressure-concentration ratio of about 17:1. Together with an ossicular lever ratio around 1.3:1, the middle ear can provide a pressure ratio on the order of 20-22, roughly 26-27 dB. This greatly improves energy transfer between air (characteristic acoustic impedance roughly 415 Pa·s/m) and fluid (order of 1.5×10⁶ Pa·s/m).
Inner ear - cochlea
The fluid-filled cochlea supports a travelling wave along the basilar membrane. Mechanical properties vary with position: higher frequencies peak near the base, lower frequencies towards the apex. This tonotopic organisation is the biological basis of frequency analysis. Hair-cell stereocilia convert mechanical motion into electrochemical signalling that reaches the auditory nerve.
The cochlea is better thought of as an active, nonlinear, frequency-selective analyser than a passive microphone. Outer-hair-cell feedback increases sensitivity and sharpens tuning at low and moderate levels.
6. We ultimately hear in the brain
Neural coding from the cochlea is processed through the auditory nerve, brainstem and higher auditory pathways. What we perceive as a stable object - a voice, instrument or location - is a constructed percept combining spectral, temporal and binaural information. Context, attention and adaptation therefore matter alongside the physical waveform.
7. How do we know where a sound comes from?
Horizontal localisation relies strongly on two binaural cues. Interaural time difference (ITD) is dominant at lower frequencies where the wavelength is large compared with head size. Interaural level difference (ILD) becomes more useful at higher frequencies where the head casts an acoustic shadow. In Rayleigh’s duplex framework, ITD tends to dominate below roughly 1.5 kHz, while ILD becomes progressively stronger above roughly 1.5-2 kHz. The exact crossover is not a hard boundary.
The pinna, torso and head also create direction-dependent spectral filtering described by the HRTF. This helps resolve elevation and front/back ambiguity.
8. Why does the same SPL sound different at different frequencies?
Human sensitivity varies strongly with frequency. Equal-loudness contours describe the SPL required at each frequency to produce approximately equal perceived loudness. Sensitivity is generally greatest through the speech-relevant midrange and decreases toward the extremes.
Phon and sone
The phon scale expresses loudness level relative to a 1 kHz reference. The sone scale attempts to represent perceived loudness more directly. Neither should be confused with SPL: they describe perception, not merely physical pressure.
9. Critical bands and masking
The ear does not resolve frequency with infinitely narrow filters. Auditory filters have finite bandwidths. Critical-band models, the Bark scale and ERB (Equivalent Rectangular Bandwidth) describe this frequency selectivity. Bandwidth is roughly near 100 Hz at low frequencies and becomes approximately proportional to centre frequency above several hundred hertz, giving a broadly logarithmic character.
A strong component can make a nearby weaker component inaudible - masking. Modern perceptual codecs exploit this effect, but masking is equally relevant when balancing instruments, noise, speech and PA systems.
10. How large is the dynamic range of hearing?
The auditory system covers an enormous physical range, from pressure variations near the threshold of hearing to levels that are uncomfortable or damaging. The useful range depends on frequency, duration, individual sensitivity and hearing status. It is misleading to compress the entire concept into a single “0-120 dB” number.
Cochlear compression is one reason why our perceptual system can operate across such a large range. Loss of outer-hair-cell function can reduce this compression, producing loudness recruitment and a narrower comfortable range.
11. How does hearing change with age?
Age-related hearing change - presbycusis - is statistical, not deterministic. Population standards such as ISO 7029 describe distributions of hearing thresholds as a function of age and sex, but they do not predict an individual. Genetics, lifetime noise exposure, disease, medication, cardiovascular factors and many other variables contribute.
High-frequency thresholds tend to deteriorate earlier and more strongly, but ageing can also affect temporal processing, frequency selectivity, binaural processing and speech understanding in noise.
There is no meaningful “age → maximum audible frequency” table
A single upper-frequency number from an online sweep is not a diagnosis. Two people of the same age can differ by many kilohertz, and equipment limitations can dominate the result.
12. Noise-induced hearing damage - level is not the whole story
Risk depends on both level and exposure duration, as well as spectral content, impulsiveness, recovery time and individual susceptibility. In an energy-based 3 dB exchange-rate model, every +3 dB halves the allowable duration:
Symbols: t₁ is the permitted duration at level L₁ and t₂ the duration at level L₂. Under a 3 dB exchange-rate model, every 3 dB increase halves the permitted duration; this is an energy model, not an individual medical limit.
This is the physical-energy basis used by NIOSH/ISO/EU approaches. The traditional US OSHA system uses a 5 dB exchange rate in several contexts; that is an administrative rule and must not be mixed with the 3 dB energy model.
Temporary and permanent threshold shifts
After intense exposure, temporary muffling, fullness or tinnitus may occur. Thresholds can partly or fully recover, but recovery does not prove that no microscopic damage occurred. Cochlear synaptopathy and so-called “hidden hearing loss” remain active research areas in humans.
The classic 4 kHz noise notch
Noise-induced hearing loss often shows a 3-6 kHz depression, frequently near 4 kHz, but this pattern alone does not prove noise causation. Audiograms must be interpreted in clinical and exposure context.
13. Tinnitus - sound perception without an external source
Tinnitus can be ringing, whistling, buzzing or another auditory sensation without an external acoustic source. It often accompanies hearing loss or noise exposure, but its mechanisms can involve both peripheral and central nervous-system processes.
Persistent or one-sided tinnitus, sudden hearing loss, dizziness, ear pain, persistent muffling after noise exposure or a new decline in speech understanding warrants professional audiological/ENT assessment. Sudden hearing loss may require urgent evaluation.
14. Why does our own voice sound different? - bone conduction
When we speak, our own voice reaches the cochlea not only through air conduction but also through vibration of the skull. A recording removes this bone-conducted component, which is why our recorded voice often sounds thinner or unfamiliar. Clinical audiometry compares air and bone conduction to help distinguish conductive and sensorineural components of hearing loss.
15. Infrasound and ultrasound
Frequencies below roughly 20 Hz are conventionally called infrasound and those above roughly 20 kHz ultrasound. These boundaries are not absolute biological switches. Very low frequencies can become audible or bodily perceptible at sufficiently high levels, while the upper hearing limit varies greatly between people and with age.
“Inaudible” does not mean “unmeasurable”. The bandwidth of an engineering system and the bandwidth of human hearing are separate questions.
16. What does all this mean for sound engineering?
A sound system is not successful merely because a datasheet claims “20 Hz-20 kHz”. Within the useful band, frequency response, phase, distortion, dynamics, temporal behaviour, directivity, noise floor and spatial uniformity all matter. The listening environment matters just as much.
Objective measurement and controlled listening are complementary: measurement tells us what physically changed; listening tells us whether the change is perceptually meaningful.
Source → electrical/digital signal → processing and gain → loudspeaker → acoustic field → outer and middle ear → cochlea → auditory nerve → brain → auditory perception.
17. Inner and outer hair cells do different jobs
Inner hair cells provide most of the afferent information transmitted to the auditory nerve. Outer hair cells actively feed mechanical energy back into basilar-membrane motion. This “cochlear amplifier” raises sensitivity and sharpens frequency selectivity at low and moderate levels.
Hearing damage is therefore not a simple linear loss of gain. Reduced outer-hair-cell function can elevate threshold, broaden auditory filters, alter compression and cause loudness to grow abnormally rapidly.
18. Temporal resolution - the ear is not only a spectrum analyser
Music and speech contain important time-domain information: attacks, envelopes, periodicity and micro-timing. The auditory system follows fine structure and amplitude envelope differently across frequency regions. Two systems with similar smoothed magnitude response can still differ in transient behaviour, decay or nonlinear distortion. “Frequency response is everything” is as misleading as saying frequency response does not matter.
19. Speech in noise - why can it be difficult with a normal audiogram?
A pure-tone audiogram measures detection thresholds for calibrated tones. Real-world speech understanding additionally depends on auditory-filter selectivity, temporal processing, binaural cues, attention and language processing. Two listeners with similar audiograms can therefore perform very differently in noise, especially with ageing.
20. Otoacoustic emissions - the ear can emit measurable sound
Active outer-hair-cell mechanics can return tiny amounts of acoustic energy into the ear canal. Sensitive microphones can measure these otoacoustic emissions (OAEs), which are used clinically as an objective indicator of cochlear - particularly outer-hair-cell - function. Their presence is not a “sound quality score”; rather, it demonstrates that the inner ear is an active system.
21. Audiograms: dB HL is not dB SPL
Clinical audiograms typically use dB HL, referenced to frequency-specific normal hearing thresholds. Consequently 0 dB HL at different frequencies does not mean the same physical dB SPL. Treating an audiogram as though it were a loudspeaker response plot is a category error: the scales have different references and purposes.
22. Hearing protection - attenuation is not simple subtraction
Earplug and earmuff attenuation is frequency-dependent and strongly influenced by fit. A poorly fitted plug can provide dramatically less protection than its laboratory rating. Musician-style earplugs aim for a flatter spectral attenuation so level is reduced without excessively changing tonal balance.
At very high occupational noise levels, dual protection can be appropriate, but the combined attenuation is not the arithmetic sum of the two nominal ratings.
23. Why YouTube and phone “frequency sweeps” are not hearing tests
The entire reproduction chain is unknown: codec, DAC, amplifier, transducer response, fit, playback level and distortion. At high frequencies a poor transducer can generate lower-frequency intermodulation or harmonic products that a listener mistakes for “hearing 18 kHz”. A proper test requires calibrated equipment and a defined method; standard audiometry and extended high-frequency audiometry are different procedures.
24. The central lesson shared by hearing science and sound engineering
Human hearing is extraordinarily good at extracting musical meaning and spatial patterns, but it is not a precise linear instrument. It adapts, masks, uses context and changes with age. Two opposite errors should therefore be avoided: claiming that measurement alone determines whether something “sounds good”, and claiming from uncontrolled listening alone that a physical difference must exist.
A disciplined method is simple: measure what changed first; then use controlled listening to determine whether the change has perceptual significance.
Sources and professional background
- NIDCD - How Do We Hear?
- World Health Organization - Deafness and hearing loss: Safe listening
- CDC/NIOSH - Understand Noise Exposure
- ISO 226:2023 - Normal equal-loudness-level contours
- ISO 7029:2017 + Amd 1:2024 - Statistical distribution of hearing thresholds related to age and gender
- Moore, B. C. J. - An Introduction to the Psychology of Hearing
- Pickles, J. O. - An Introduction to the Physiology of Hearing
Professional and legal notice
I prepare the technical descriptions, calculations, examples, diagrams and other information published in SWORD LAB for educational and informational purposes. When compiling the material I aim for technical accuracy, correct presentation of the underlying relationships and careful use of the available professional knowledge.
Nevertheless, the information may contain inaccuracies, errors or simplifications that cannot be applied unchanged to a specific system or environment. The calculations and engineering examples are generally based on stated or implicit assumptions. Real systems are also affected by the actual parameters of the equipment, system topology, environmental conditions, measurement method, installation practice, applicable standards, legislation and manufacturer requirements.
The material I publish does not constitute design documentation, an expert opinion, an installation instruction, a safety instruction or individual professional advice. It does not replace manufacturer documentation, current regulations and standards, or - where required - the examination, measurement or design work of a suitably qualified and authorised professional.
Technical standards, product data and technologies change over time. Before design, installation, measurement, operation, repair or equipment selection, I therefore recommend checking the current primary and authoritative sources.
I make every reasonable effort to prepare the material carefully; however, to the extent permitted by applicable law, I accept no liability for direct or indirect damage, loss, malfunction or interruption resulting from the use, misinterpretation or incorrect application of information published on this site, or from interventions carried out on that basis.
The purpose of SWORD LAB is to help explain engineering relationships and the physical and technical processes behind sound reinforcement, electroacoustics, digital audio and DJ technology. It is not intended to replace on-site investigation, measurement or engineering design of a specific system.
