Editorial review
Pitch, timbre, resonance and formants explained
The source-filter model in plain terms, why two voices at the same hertz sound different, what formants actually are, and which of these a pitch app can and cannot measure.
If you have ever measured a low pitch and still not sounded the way you wanted, the explanation is in this guide. Pitch is one property of a voice. Most of what makes a voice recognisable, and most of what listeners react to, lives somewhere else.
Four words that get used interchangeably
| Term | What it is | Where it comes from |
|---|---|---|
| Pitch | The perception of how high or low a voice is | Tracks fundamental frequency, F0 |
| Timbre | The overall quality that tells two voices apart on the same note | Everything except F0 and loudness |
| Resonance | How the airway above the folds shapes the sound | The size and shape of the vocal tract |
| Formants | Measurable peaks of acoustic energy | Produced by that shaping |
Only the first is measured in hertz on a pitch display. The other three are why the number never tells the whole story.
Source and filter
The standard model of voice production has two stages.
The source is the vocal folds. They vibrate and produce a buzzy sound that is rich in harmonics, and the rate of that vibration is F0.
The filter is everything above them: throat, mouth, and sometimes the nasal cavities. This tube does not create sound, it shapes it, boosting some frequency regions and damping others.
The crucial property is that the two are largely independent. Moving your tongue, changing your lip shape, or letting the larynx sit lower changes the filter without changing F0 at all. That is the whole answer to why two people measuring the same hertz can sound nothing alike, and why a low measurement does not guarantee a low- sounding voice.
What formants actually are
A formant is a frequency region the tract boosts. They are numbered from the bottom up: F1, F2, F3, and so on. Do not confuse F1 with F0. F0 is the fold vibration rate, and F1 is the first resonance of the tube above it.
Two things are worth knowing.
F1 and F2 mainly determine which vowel you hear. Change the shape of your mouth and you change the vowel, without touching your pitch.
The overall spacing of formants relates to the length of the tract. A longer tract pushes formants down, which reads as bigger and darker. This is why lowering the larynx and opening the throat changes perceived depth even when the hertz stay put.
There is measured support for this mattering more than pitch does. In the meta-analysis by Pisanski and colleagues, formant-based estimates of vocal tract length explained up to about 10% of the variation in body height, while fundamental frequency explained under about 2%. If acoustics carry an impression of size, formants carry more of it than pitch does.
What an app can measure, and what it reports
Worth being precise about, because it determines what feedback is trustworthy.
F0 is measured well. It is a rate, it is robust, and it is the number Basaltone tracks over time.
F1 is estimated, coarsely. Basaltone computes it continuously and turns it into a three-way band rather than a raw number: low larynx, neutral, or high larynx. That is deliberate. Formant estimation from a phone microphone is far less reliable than pitch detection, and a precise-looking number would imply an accuracy that is not there. A band you can act on beats a decimal you cannot trust.
The fine measures are live only. Higher formants, harmonics-to-noise ratio and spectral tilt are computed while you speak, to drive live feedback, and are not stored. They are useful as an in-the-moment indicator and too noisy to make a progress curve from.
Timbre as a whole is not measured. No consumer tool reduces it to a score.
What this means for training
The practical ranking is counterintuitive, and it is why pitch-only training stalls.
Pitch is the easiest thing to measure and often the least decisive perceptually. Resonance is harder to measure and frequently the bigger lever, particularly if hormones are not part of the picture, because it does not depend on fold anatomy at all. Speech patterns, meaning intonation and pacing, are not acoustic properties in this sense at all and are fully learnable.
So track pitch when it serves a defined exercise, and watch ease, stability, resonance and how the voice behaves in connected speech alongside it. Voice masculinization and the limits of pitch works through what that means when the goal is being heard differently, and understanding your vocal frequency in Hz covers how to read the one number you do get reliably.
What none of these are
No single acoustic marker defines gender, quality, health or identity. Formants do not either, despite being the better size cue: "better" here means about 10% of the variance, not a verdict.
A training plan that chases any one number at the expense of comfort has the priorities backwards, and vocal fatigue, tension and safe practice covers what that costs.