← Back to Learn
Soundb Learn · Fundamentals

How Hearing Actually Works

The mechanical chain from eardrum to nerve — and the built-in boost that explains why 2 to 5 kHz is both the most intelligible and the most exhausting band you will ever work in.

Topic
Fundamentals
Level
Beginner
Format
Lesson
Time
14 min

Your ears are the only measurement instrument you will never be able to replace, recalibrate or send for repair. Every other tool in the studio exists to serve them. So it is worth knowing what they physically are — not out of curiosity, but because several of the most confusing things about mixing turn out to be straightforward consequences of the hardware.

The mechanics of hearing are well understood and genuinely elegant. What the brain does with the result afterwards is still argued over, and that half gets its own lessons. This one is the machine.

The problem the ear had to solve

Start with the difficulty, because the design only makes sense once you see it.

Sound arrives as a pressure wave in air. The part of you that actually detects it is suspended in fluid. And when a wave travelling through a light medium meets a much denser one, almost all of it reflects rather than crosses — which is why you hear so little of what is going on above the surface when your head is underwater. Left to itself, the overwhelming majority of the sound energy arriving at a fluid-filled sensor would simply bounce off.

So the outer and middle ear are not decoration. They are an impedance matcher: a mechanical system whose entire job is to take a large, weak movement of air and convert it into a small, strong movement of fluid with as little loss as possible. Everything in the next two sections is that conversion.

The outer ear: a funnel that is also a filter

The visible part — the pinna — collects sound over a larger area than the ear canal opening alone would, which is a modest free gain. But its more interesting job is that its folds and ridges are deliberately irregular. Sound arriving from different directions reflects around those folds differently, arriving at the canal with slightly different frequency colouration depending on where it came from.

That colouration is a direction cue. It is a large part of how you tell whether something is above you, below you, in front or behind — distinctions that pure left-right timing cannot make, because a sound directly in front and directly behind reaches both ears identically. Your brain has learned the specific filtering signature of your own two ears, which is why localization is subtly personal and why binaural recordings made on someone else's head never quite work as well as your own.

The ear canal is a tuned tube

The canal is roughly 2.4 cm long and closed at one end by the eardrum, which makes it a closed-tube resonator — the same physics as blowing across a bottle. A tube of that length resonates at around 3,700 Hz at body temperature, with a further sensitivity peak somewhere around 13 kHz.

Combine the pinna's collection, the canal's resonance and diffraction around the head, and the outer and middle ear together deliver something on the order of 20 dB of acoustic amplification, concentrated in the 2 to 5 kHz region.

Take that seriously, because it explains a great deal:

  • It is why that band carries speech intelligibility — the consonants that distinguish one word from another live there, in the region we are most sensitive to.
  • It is why a small presence boost around 3 kHz makes a vocal leap forward so dramatically, and why a large one becomes unbearable so quickly. You are boosting a band that is already boosted.
  • It is why that region causes listening fatigue faster than any other, and why long sessions leave you reaching for the monitor volume.
  • It is why almost every microphone marketed for vocals has a designed-in presence lift in roughly the same place — and why stacking a presence-peaked mic, a presence EQ boost and your own ear canal produces something genuinely harsh.

The middle ear: a lever and a piston

The eardrum (tympanic membrane) closes the end of the canal and moves in sympathy with the arriving pressure changes. Behind it sit the three smallest bones in the body — the hammer, anvil and stirrup, known collectively as the ossicles — which carry that movement to the oval window, the entrance to the fluid-filled inner ear.

The impedance matching happens in two ways at once.

  1. Area. The eardrum is roughly fifteen times larger than the oval window. Collecting force over a large surface and delivering it through a small one concentrates the pressure, in the same way a drawing pin concentrates the push of your thumb.
  2. Leverage. The ossicles are arranged as a compound lever, adding further mechanical advantage — modest, but multiplied by the area effect.

Together they deliver enough gain to push fluid effectively with nothing more than moving air. It is worth noting that the published figures for each stage vary between sources and are best treated as representative rather than exact; what is not in doubt is the principle, or the roughly 20 dB total.

The middle ear also has a protective function, covered further down, and it is connected to the throat by the Eustachian tube — which is why swallowing equalises the pressure in your ears on a descending flight, and why a blocked ear during a cold genuinely changes what you hear.

The inner ear: where sound becomes frequency

The cochlea is a fluid-filled spiral, a little over two turns, about 3.2 cm long if you could unroll it. It is divided lengthways into three channels. Two carry pressure; the third holds the sensor. The fluids in them differ chemically — and they have to stay separate, since mixing them impairs hearing.

Running along the middle channel is the basilar membrane, and sitting on it is the organ of Corti, the actual sensor. It carries rows of hair cells — somewhere between 16,000 and 20,000 of them, each topped with around a hundred fine projections called stereocilia.

When the stirrup pushes on the oval window, it sends a travelling wave down the fluid. That wave displaces the basilar membrane, the displacement bends the stereocilia, and the bending triggers an electrical impulse that leaves along the auditory nerve. That is the moment sound becomes information.

Place Theory: your ear is a spectrum analyser

Here is the part with the most direct consequences for your work.

The basilar membrane is not uniform. It is stiff and narrow at the end nearest the oval window and progressively floppier and wider toward the far end. Because of that gradient, different frequencies peak at different physical positions along it:

  • High frequencies peak early, close to the entrance.
  • Low frequencies travel much further along the membrane before they peak.

Which position gets excited most is what determines the pitch you perceive. This is Place Theory, and the implication is worth stating plainly: frequency analysis in your hearing is physical, not computational. Your ear separates a complex sound into its component frequencies by literally spreading them out along a strip of tissue, before your brain is involved at all.

Almost everything in later psychoacoustics follows from that layout. Masking happens because two frequencies close together compete for the same stretch of membrane. Critical bands are measured in millimetres of membrane. The reason a boomy low end obscures a vocal is that low-frequency energy displaces a wide region of the membrane on its way through, and that displacement is physical territory the vocal now has to share.

Place Theory is a good model, not a complete one — it accounts well for relative pitch and does not explain absolute pitch at all. But as a working mental picture for an engineer it is hard to improve on.

The resolution problem

A healthy ear can distinguish 440 Hz from 441 Hz — a difference of about 0.2 per cent. Across the whole membrane that works out to somewhere around 1,500 separately distinguishable pitches.

Do the arithmetic and something does not add up. With roughly 16,000 to 20,000 hair cells spread over 1,500 distinguishable pitches, only about a dozen cells are available per pitch. No purely mechanical resonance in a strip of tissue could possibly be that sharp — a floppy membrane in fluid should smear far more than that.

So there must be an active sharpening mechanism: the ear is not a passive microphone but an actively tuned system that narrows its own response, with the outer hair cells doing work rather than simply reporting. This matters practically for one reason above all — an active system can be exhausted and can be damaged, and it is the outer hair cells, the ones doing the sharpening and handling quiet detail, that go first. The next lesson deals with what that costs you.

The ear protects itself, badly

Exposed to sustained loud sound, small muscles in the middle ear tighten, repositioning the ossicles so that less force reaches the oval window. It is a genuine built-in limiter and it broadens the range your ears can survive.

It has two serious limitations, and both matter in a studio.

It is slow. It takes tens of milliseconds to engage, which is useless against exactly the sounds most likely to hurt you — a snare hit at close range, a feedback squeal, a dropped microphone, a click through headphones. The transient is over before the protection arrives.

And it fatigues. It cannot hold indefinitely, so it offers little defence against the real danger, which is not one loud moment but hours of merely loud.

Why two ears, and what happens after the nerve

Everything so far describes one ear. You have two, and the difference between what they report is where a large part of hearing actually happens.

Signals leave each cochlea along the auditory nerve, and crucially they do not stay on their own side. Information from both ears is combined at multiple relay points on the way up, and both hemispheres receive input from both ears. Your brain is not comparing two recordings after the fact; the comparison is built into the wiring at a low level.

That comparison buys you three things you would not otherwise have:

  • The ability to locate a sound source, from differences in arrival time and level between the ears.
  • The ability to separate several sources that are playing at once, by placing them in different directions.
  • The ability to pull one source out of a noisy background — the reason a conversation in a crowded room is possible with two ears and close to impossible through a single mono recording of the same room.

That last one has a direct studio consequence. A mono recording of a busy space sounds far more cluttered than the same space did in person, and nothing is wrong with the microphone. You have simply removed the mechanism your brain was using to unpick it. The localization lessons take this apart properly.

Bone conduction, and why your recorded voice sounds wrong

Sound reaches your cochlea by two routes. The one described above — air, eardrum, ossicles — and a second one: vibration conducted directly through the bones of your skull.

For every sound in the world except your own voice, the second route contributes almost nothing. But when you speak, your vocal folds vibrate your skull directly, and a substantial amount of low-frequency energy arrives at your cochlea through bone rather than through air.

So the voice you have heard your whole life is the air-conducted version plus a bass-heavy bone-conducted version that nobody else receives. A recording captures only the air-conducted half. It is not that the microphone is unflattering or that the recording is thin — it is genuinely accurate, and the version you are used to is the one nobody else has ever heard.

Worth knowing for two practical reasons. It explains why singers so often dislike their own recorded sound, which is a moment worth handling well in a session. And it is why a vocalist wearing only one headphone — with one ear open to the room and their own skull — hears something quite different from what is being recorded, and often pitches accordingly.

What changes with age

Hearing loss from ageing, called presbycusis, is close to universal and follows a predictable pattern: it starts at the top of the range and works downward. A typical forty-year-old has lost a good deal above 15 kHz; a typical sixty-year-old may hear little above 10 to 12 kHz.

This is not a reason for despair, and the reason is worth stating plainly, because young engineers worry about it and experienced ones mostly do not. Almost all musical information — every fundamental, the great majority of harmonics that define timbre, all of speech intelligibility — lives well below 10 kHz. What is lost at the very top is air and sheen rather than substance.

What it does mean is two working habits. Know roughly where your own hearing ends, because you cannot mix what you cannot hear, and a spectrum analyser is a reasonable prosthetic for the region above it — it is entirely normal to make decisions about 16 kHz partly by measurement. And take the damage lesson that follows seriously, because age-related loss is unavoidable and noise-induced loss on top of it is not. The two add up, and only one of them is under your control.

What the machine actually gives you

Worth keeping somewhere you will see it:

  • Frequency range roughly 20 Hz to 20 kHz, narrowing with age from the top down.
  • Response to frequency is logarithmic — which is why octaves, not hertz, are the musically meaningful unit.
  • Response is not flat — there is a large built-in boost around 2 to 5 kHz.
  • Dynamic range around 120 dB at its most sensitive frequencies.
  • Frequency discrimination of about 2 Hz at 1 kHz.
  • Level discrimination of roughly 1 to 3 dB, depending on frequency and level.
  • Very little sensitivity to the absolute phase of a single sound — but, as the phase lesson showed, enormous sensitivity to phase differences between two arrivals.

That last pairing is worth sitting with, because it resolves an apparent contradiction. You cannot hear the absolute phase of an isolated sound. You can hear, immediately and unmistakably, what happens when two copies of it arrive at slightly different times.

What this leaves out

This lesson has covered the mechanism — how sound gets from the air to the auditory nerve. It has deliberately not covered what your brain does with the result, and that is where most of the surprises live: why the same sound seems louder at some frequencies than others, why a mix that sounds balanced loud falls apart quiet, and what a decade of loud monitoring actually does to the hair cells described above.

That is the next lesson, and of the two it is the one more likely to change how you work tomorrow.

Studio Rule

Your ears have a built-in boost of roughly 20 dB centred where vocals live. Every decision you make between 2 and 5 kHz is being made through it.

What to practice

  • Cup your hands behind your ears and turn slowly toward a sound source. You have just changed your own directional filtering, and it is dramatic.
  • Play a 3 kHz tone and a 300 Hz tone at the same measured level. Note how much louder the 3 kHz one seems, then remember that difference every time you boost presence.
  • Sweep a narrow EQ boost through a mix and find the point where it stops sounding bright and starts sounding painful. That is your ear canal resonance, not the mix.
  • After a long session, play a reference track you know well. If it sounds dull, your ears have shifted, not the track. Stop for the day.
Paid Mentor Access
Ask About This Lesson

Students pay for getting unstuck: ask a concept question, routing issue, DAW confusion, or mix decision tied to this lesson.

0
Credits
Unlock direct answers

Free readers can learn from the public Q&A archive. Paid students can ask their own lesson-specific questions and get mentor replies.

Uses 1 credit.
Question saved — a mentor will post a reply here once it's answered.
Answered Questions