What Sound Actually Is
Frequency, amplitude and wavelength — the physical reality underneath every EQ move, every mic position and every mix decision you will ever make.
Frequency, amplitude and wavelength — the physical reality underneath every EQ move, every mic position and every mix decision you will ever make.
Every fader move, every EQ boost and every microphone position you will ever choose is really an argument with one thing: a pattern of pressure travelling through air. You cannot see it, you cannot pause it, and it is gone a fraction of a second after it arrives. Everything else in this curriculum — compressors, room treatment, monitor placement, mastering — exists to manage that one physical event. So it is worth understanding properly, once, in a way that stays useful for the next twenty years.
The good news is that it is genuinely simple physics. There is no part of this lesson that requires anything beyond division.
Any time information travels from one place to another, three things have to exist. Something has to disturb the world — a string, a drum head, a vocal fold, a speaker cone. Something has to carry that disturbance — air, usually, but water and steel work far better. And something has to be built to pick it back up at the other end — an ear, or a microphone, which is the same job done with a different material.
The middle one is the part beginners skip, and it has a consequence that sounds like a joke until you think about it. Sound is not a thing that exists on its own; it is something that happens to a substance. Remove the substance and there is nothing left to happen. In a vacuum there are no molecules to push against each other, so there is no sound — not quiet sound, not faint sound, none. That is the literal, physical reason nobody can hear you scream in space, and it is also why a mic inside a sealed vacuum chamber picks up silence no matter how loud the thing inside it is vibrating.
Air feels like emptiness, but it is packed with molecules — nitrogen, oxygen, a bit of everything else — sitting at roughly even distances from each other and jittering about at random. That even spacing is what we call normal air pressure, and it is the baseline everything else is measured against.
Now vibrate something in the middle of it. A speaker cone punches forward and shoves the molecules directly in front of it closer together. Those molecules, now crowded, push against their neighbours, which crowd and push against theirs, and so on outward in every direction. When the cone pulls back it leaves a slightly thinner region behind it, and molecules rush in to even things out. Push, pull, push, pull. What travels outward is a repeating pattern of slightly-crowded and slightly-thin — regions of higher pressure called compressions and regions of lower pressure called rarefactions.
Here is the part worth sitting with: no air actually travels from the speaker to your ear. Each molecule shuffles back and forth a tiny distance about its resting position and stays roughly where it was. What crosses the room is the pattern — the disturbance itself, handed from molecule to molecule.
Think about a traffic jam on a busy road. A driver at the front brakes; the car behind brakes; the one behind that brakes. A dense clump of stopped cars forms, and that clump travels backwards down the road — you can watch it move — even though every single car in it is facing forward and moving forward. The clump is not made of any particular cars. Cars enter it at one end and leave it at the other. The jam is a pattern moving through the traffic, not an object moving along the road. A sound wave is the same trick performed by air molecules, several hundred times a second.
Waves come in two shapes, and it matters which one sound is.
Throw a stone into a still tank of water and ripples spread outward. The water surface moves up and down while the ripple travels sideways — the movement is at right angles to the direction of travel. That is a transverse wave. It is the shape almost everyone pictures when they hear the word wave.
Sound is the other kind. The molecules move back and forth along exactly the same line the wave is travelling — forward and backward, crowding and thinning, in line with the direction of the sound. That is a longitudinal wave, and it is the correct picture for sound in air.
Which brings us to the picture you have definitely seen. Sound is almost always drawn as a smooth up-and-down curve, and that drawing is not wrong — but it is not a picture of air. It is a graph. The horizontal axis is time; the vertical axis is pressure. When the curve is above the centre line, pressure at that instant is above normal — a compression. When it dips below, pressure is below normal — a rarefaction. The centre line is ordinary undisturbed air.
So the crests are not air moving up. They are moments of higher pressure. Once you read that curve as a pressure meter rather than as a photograph, every waveform display in your DAW starts telling you something real, and a surprising amount of confusion about phase and polarity later on simply evaporates.
Sound moves at a fixed speed through a given medium. That speed does not depend on how loud the sound is or how high or low it is — a whisper and an explosion, a sub-bass note and a cymbal, all travel at exactly the same rate through the same air. What the speed does depend on is the air itself, and mostly its temperature.
In dry air at sea level and 0 °C, sound travels at 331.4 metres per second. Every degree warmer adds about 0.6 m/s; every degree cooler takes about 0.6 m/s away. That gives you a formula you can do in your head:
speed = 331.4 + (0.6 × temperature in °C)
So in a comfortable air-conditioned control room at 20 °C:
331.4 + (0.6 × 20) = 331.4 + 12 = 343.4 m/s
And in an un-airconditioned room in Chennai in May, at 35 °C:
331.4 + (0.6 × 35) = 331.4 + 21 = 352.4 m/s
That is a difference of 9 m/s, or about 2.6 per cent. It sounds trivial, and for most purposes it is. But it means the physical length of every wave in the room changes with the temperature, which means the frequencies at which your room resonates drift slightly over the course of a hot afternoon, and the delay between a direct sound and its reflection off the back wall drifts with them. It is one of several quiet reasons a mix that sounded settled at midnight can sound faintly different at 3 p.m.
Two more facts worth carrying. Sound travels faster in denser media, not slower — roughly 1,500 m/s in water and around 5,000 m/s in steel, which is why sound reaches you through a building's structure before it reaches you through the air, and why isolating a room from footsteps is a different problem from isolating it from voices. And at altitude, where the air is thinner, it travels more slowly.
Turn the speed of sound around and you get the single most-used conversion in a studio. If sound covers 343 metres in a second, then in one thousandth of a second it covers 343 millimetres — call it 34 cm per millisecond, or near enough a foot per millisecond if you think in feet.
Keep that in your head and a whole class of problems becomes arithmetic you can do while standing in the room:
That conversion also puts a floor under how tightly a band can play together. Musicians spread across a large stage are physically unable to hear each other in real time — at 10 metres apart there is a genuine 29 ms gap, which is more than a semiquaver at 120 bpm. This is not a discipline problem, and no amount of rehearsal removes it. It is why in-ear monitoring changed live performance so completely: it replaces a delayed acoustic path with an effectively instantaneous electrical one.
Count how many complete push-and-pull cycles pass a fixed point in one second and you have the frequency, measured in hertz (Hz). A wave completing 440 cycles per second is a 440 Hz wave. That is a measurable, physical, arguable-with-instruments fact about the wave.
Pitch is what your brain makes of that frequency, and it is not quite the same thing. Most of the time the two track each other so closely that the distinction seems academic — until it does not.
Here is the clearest example. Take a tone, do not change its frequency at all, and simply make it louder. Above roughly 2 kHz, the tone will seem to creep upward in pitch. Below roughly 2 kHz, it will seem to sink downward. The physical frequency has not moved by a single hertz; only the level changed. The effect is real and has been measured repeatedly — a 6 kHz tone taken from 60 dB to 90 dB has been found to rise by more than 30 cents in perceived pitch, while a 200 Hz tone dropped by about 20 cents across the same change. A cent is a hundredth of a semitone, so these are small shifts, but they are large enough to matter when you are tuning by ear at a different monitoring level from the one you will be judged at.
Humans hear roughly 20 Hz to 20 kHz, though the top of that range narrows with age and with years of loud listening, and almost nobody over thirty still has a genuine 20 kHz. Sound does not stop existing beyond those limits — it just stops being audible and picks up a different name. Above 20 kHz is ultrasound, which bats navigate with and which cleans your spectacles. Below 20 Hz is infrasound — too low to register as a note at all, but physically real, used by elephants over long distances, and a large part of what you feel in your chest rather than hear during a very loud low-end passage.
Frequency tells you how often the pressure swings. Amplitude tells you how far it swings — how much the pressure departs from normal at the peak of each cycle. Your ear reads amplitude as loudness.
These are independent properties, and blurring them is one of the most common beginner errors. A bass note is not automatically louder than a cymbal. A quiet kick drum is still a kick drum, just a small one. You can have a very high frequency at a very low amplitude and vice versa, in any combination. Frequency is the shape of the wave; amplitude is the size of it.
This distinction is the entire reason EQ and compression are different tools rather than two flavours of the same knob. An equaliser changes the balance of frequencies without caring about the overall size. A compressor manages amplitude over time without touching the frequency content at all. When you can look at a problem and say that is a frequency problem or that is an amplitude problem, you have already done most of the diagnostic work.
There is more than one way to measure amplitude, and the difference will matter constantly from here on.
Peak amplitude is the highest instantaneous value the wave reaches — the single furthest excursion from the centre line. It is the number that tells you whether something is about to clip, because clipping is an instantaneous event.
RMS (root mean square) is a kind of meaningful average across the wave, and it corresponds far more closely to how loud the sound actually seems, because your ear responds to sustained energy rather than to momentary spikes. For a pure sine wave there is a fixed relationship between the two:
So a sine wave peaking at 1.0 has an RMS value of 0.707. Those two numbers are reciprocals of each other and they come from the geometry of a sine — you will meet them again in metering, in gain structure and in every compressor that offers you a choice between peak and RMS detection.
The crucial caveat: that 0.707 relationship holds for a sine wave. Real music is not a sine wave, and the gap between its peak and its average is much larger and constantly changing. That gap has a name — crest factor — and managing it is essentially what dynamics processing is for.
A wave travels at a fixed speed and repeats at a fixed rate, so each complete cycle occupies a specific physical distance in the air. That distance is the wavelength, and it is simply speed divided by frequency:
wavelength = speed of sound ÷ frequency
At 343 m/s, that gives you a set of numbers worth internalising, because they explain an enormous amount of otherwise mysterious behaviour:
Look at the range. The things you can hear span physical sizes from longer than a bus to smaller than your fingernail, and they all have to coexist in the same room and be captured by the same microphone. Almost everything difficult about acoustics comes from that fact.
The single most useful consequence is this: a wave interacts with an object based on the object's size relative to the wavelength. If an obstacle is much smaller than the wavelength, the wave simply bends around it as though it were not there. If the obstacle is much larger, the wave is reflected or blocked.
You have two musicians in one room and you put a 1.2 metre acoustic screen — a gobo — between them, hoping to stop one bleeding into the other's microphone. Does it work?
Take a cymbal at 4 kHz. Its wavelength is 343 ÷ 4,000 = 8.6 cm. The screen is 1.2 m — roughly fourteen times bigger than the wave. To that wave, the screen is a wall. It is reflected and blocked, and very little gets past.
Now take the bass guitar's low E at about 41 Hz. Its wavelength is 343 ÷ 41 = 8.4 metres. The screen is 1.2 m — about a seventh of the wave's length. To that wave the screen is barely an obstruction at all; it bends straight around it and carries on.
So the gobo works, but only on the top half of the spectrum. What arrives at the neighbouring microphone is not less bleed — it is bleed with the highs stripped off, which sounds muddy and indistinct, and is often harder to live with in a mix than the full-range version would have been. That is not a flaw in your gobo; it is 8.4 metres versus 1.2 metres, and no amount of money spent on a better screen changes the arithmetic.
The same reasoning, run in different directions, explains why a small bedroom struggles specifically with low end and not with treble, why a kick drum needs more physical space to develop than a hi-hat does, and why the thump from a party three houses away arrives without any of the vocals attached.
“Bass travels further than treble.” Not quite — the useful version is that bass gets through obstacles better and is absorbed less by air over long distances. In open space with no obstructions, all frequencies obey the same inverse-square spreading. The everyday experience that produces the myth is real; the explanation usually given for it is not.
“Turning it up adds bass.” The amplifier is not adding low frequencies. Your hearing's frequency response genuinely changes with level, and it gains low-end sensitivity as things get louder. The change is in you, not the signal — which is a large enough deal to have its own lesson at the end of this course.
“A higher sample rate means better frequency response you can hear.” Sample rate sets the highest frequency a digital system can represent, and the ceiling at standard rates already sits above human hearing. There are real arguments for higher rates, but they are about filter behaviour and processing headroom, not about capturing audible frequencies you were previously missing. The digital audio course takes this apart properly.
One thing has been deliberately left out. Almost nothing you actually hear is a single pure frequency — a real note is a stack of many frequencies sounding at once, and the particular recipe of that stack is why a guitar and a flute playing the identical note are instantly distinguishable. That is a big enough subject to have its own lesson, and it is the next one: harmonics and timbre.
Before that, though, the wave has to get across the room, and the room has opinions. That is the lesson immediately after this one.
Every tool you will ever learn is a specific way of manipulating one of the properties in this lesson. An equaliser reshapes which frequencies are loud. A compressor reshapes amplitude over time. Reverb and delay exploit travel time and distance. Microphone placement is a decision about wavelength and reflection. Room treatment is an argument with the numbers in that wavelength table.
None of that requires you to do physics while you work. It requires you to be able to look at a problem and say what kind of problem it is. That sounds bad is not a diagnosis. That is a cancellation caused by a reflection arriving a millisecond late is — and the fix follows from it immediately, which is the entire point.
Frequency is the shape of the wave; amplitude is its size. Nearly every mix problem is one or the other, and naming which one is most of the fix.
Students pay for getting unstuck: ask a concept question, routing issue, DAW confusion, or mix decision tied to this lesson.
Free readers can learn from the public Q&A archive. Paid students can ask their own lesson-specific questions and get mentor replies.