Phonetics
Before a recognizer can read speech it has to know what speech is. This first part covers the linguistic substrate: phones and their transcription in the IPA and ARPAbet; articulatory phonetics — how the vocal tract shapes airflow into consonants and vowels; and prosody — stress, tune, and the F0 contour.
╌╌╌╌
A written character like p or a is already a scientific model of speech: a
claim that the continuous, gliding sound of a spoken word decomposes into a small
set of discrete, reusable units. That claim is old — the earliest writing systems
already used some symbols for sounds rather than whole words — and it is the same
claim that makes speech recognition
and text-to-speech possible.1 Phonetics is the study of those units:
what they are, how the human vocal tract produces them, and what their acoustic
signal looks like once it reaches a microphone.
This lesson is the linguistic and physical groundwork for the recognizer that follows. That recognizer never sees a phone directly; it sees a waveform, turns it into a spectrogram, and learns to read the acoustic traces that phones leave there. Everything here — the vocal tract as a filter, formants as resonances, F0 as the vibration of the vocal folds — explains where those traces come from. The log-mel front end of a modern ASR system is a mechanized, learned version of exactly the acoustic analysis this chapter does by hand.
Speech sounds and phonetic transcription
We represent the pronunciation of a word as a string of phones — speech
sounds, each written with a symbol adapted from the Roman alphabet.1 A
phone is a physical event, an actual sound produced by a vocal tract, and a symbol
in square brackets, like [p], denotes a specific articulation.
The standard cross-linguistic system is the International Phonetic Alphabet
(IPA), first developed in 1888 and still evolving, which has a symbol for every
sound in every documented language. For English work it is common to use the
ARPAbet instead: a plain-ASCII phonetic alphabet covering the American-English
subset of the IPA, which is convenient precisely because it needs no special
glyphs. The two are in one-to-one correspondence over their shared inventory; the
ARPAbet [sh] is the IPA ʃ, [iy] is i, [aa] is ɑ.
The reason we transcribe at all — rather than just spelling — is that English orthography is a poor guide to sound. The mapping from letters to phones is opaque: one letter spells many sounds, and one sound is spelled many ways.
Many languages — Spanish, for one — are far more transparent, but for English, talking about sound requires a notation for sound. The table below is a working subset of the ARPAbet, with the IPA equivalent and an example word.
| ARPAbet | IPA | Example | ARPAbet | IPA | Example | |
|---|---|---|---|---|---|---|
[p] | p | parsley | [iy] | i | lily | |
[t] | t | tea | [ih] | ɪ | lily | |
[k] | k | cook | [eh] | ɛ | pen | |
[b] | b | bay | [ae] | æ | aster | |
[d] | d | dill | [aa] | ɑ | poppy | |
[g] | g | garlic | [ao] | ɔ | orchid | |
[m] | m | mint | [uh] | ʊ | wood | |
[n] | n | nutmeg | [uw] | u | tulip | |
[ng] | ŋ | baking | [ah] | ʌ | butter | |
[f] | f | flour | [er] | ɝ | bird | |
[s] | s | soup | [ax] | ə | lotus | |
[sh] | ʃ | squash | [ay] | aɪ | iris | |
[ch] | tʃ | cherry | [aw] | aʊ | flower | |
[l] | l | licorice | [oy] | ɔɪ | soil | |
[r] | r | rice | [ow] | oʊ | lotus |
The left half of the table is consonants, the right half vowels; the rest of the
lesson explains what that division means physically. A companion notion is the
phoneme, the abstract sound category a language treats as one unit, versus the
allophone, the concrete variant that actually surfaces in context. English
treats the aspirated [p] of pin and the unaspirated [p] of spin as the
same phoneme even though they are physically distinct phones; the two are
allophones of one phoneme. Phonetics studies the phones; phonology studies which
distinctions a language treats as meaningful.
Articulatory phonetics
Articulatory phonetics studies how phones are produced as the organs of the mouth, throat, and nose modify airflow from the lungs.2 Sound starts as air expelled from the lungs up the trachea (windpipe). At the top of the trachea sits the larynx — the Adam's apple — which holds two folds of muscle, the vocal folds. The gap between them is the glottis.
If the folds are held close together, airflow makes them vibrate, and the sound is
voiced; if they are apart, no vibration occurs and the sound is unvoiced
(voiceless). All English vowels are voiced, as are [b] [d] [g] [v] [z]; [p] [t] [k] [f] [s] are their voiceless counterparts. Above the larynx is the vocal
tract, split into the oral tract (out the mouth) and the nasal tract
(out the nose). Sounds routed partly through the nose — [m] [n] [ng] — are
nasal.
Phones split into two classes. Consonants restrict or block the airflow;
vowels leave it comparatively open, are usually voiced, and are louder and
longer. Semivowels like [y] and [w] sit between: voiced like vowels but short
and non-syllabic like consonants.
Consonants: place of articulation
Because a consonant is made by pinching the airflow somewhere, we can sort consonants by where that pinch happens — the place of articulation.2
Working front to back: labial sounds close the two lips (bilabial [p] [b] [m]) or press the lower lip to the upper teeth (labiodental [f] [v]).
Dental sounds [th] [dh] place the tongue at or between the teeth. Alveolar
sounds [s] [z] [t] [d] press the tongue tip to the alveolar ridge, the bump just
behind the upper teeth. Palatal (strictly, palato-alveolar) sounds [sh] [ch] [zh] [jh] and the palatal [y] raise the tongue toward the hard palate. Velar
sounds [k] [g] [ng] press the back of the tongue against the velum, the soft
palate at the very back. A glottal stop closes the glottis itself.
Consonants: manner of articulation
The place says where; the manner of articulation says how the constriction is made — a full stop, a narrow hiss, or a gentle approach.2 Place and manner together, plus voicing, almost always identify a consonant uniquely.
A stop (plosive) blocks the air completely, then releases it in a burst; the
silent block is the closure and the burst is the release. English has
voiced stops [b] [d] [g] and voiceless [p] [t] [k]. Nasals [n] [m] [ng]
lower the velum so air escapes through the nose. In a fricative the channel is
narrowed but not closed, and the turbulent air hisses: [f] [v] [th] [dh] [s] [z] [sh] [zh], of which the loud, high ones [s] [z] [sh] [zh] are sibilants. A
stop released straight into a fricative is an affricate, [ch] and [jh]. In
an approximant the articulators come close but not close enough for turbulence:
[y] [w] [r] [l], with [l] a lateral because air flows over the lowered
sides of the tongue. A tap or flap [dx] is a single quick brush of the
tongue against the alveolar ridge — the middle consonant of lotus [l ow dx ax s]
in most American dialects.
Vowels and the vowel space
Vowels have no real constriction, so they are described not by place and manner but
by the position of the tongue body and the shape of the lips.2 Three
parameters do the work: height (how high the tongue's high point sits),
frontness/backness (whether that high point is toward the front or back of the
mouth), and rounding (whether the lips are pursed). For [iy] (beet) the
tongue is high and front; for [uw] (boot) high and back; for [ae] (bat) low
and front.
The space is drawn as a quadrilateral because that roughly matches both tongue
positions and, more accurately, the acoustics we get to below. A vowel with a
fixed tongue position is a plain vowel; one in which the tongue moves markedly
during the vowel is a diphthong, drawn here as an arrow from start to end —
English is rich in them ([ay] [aw] [oy]). Certain back vowels ([uw] [ao] [ow])
are pronounced with the lips rounded.
Syllables
Consonants and vowels group into syllables. The vowel at the core is the
nucleus; the consonants before it are the onset; the consonants after it
are the coda; and the nucleus plus coda together form the rime. Dog[d aa g] has onset [d], nucleus [aa], coda [g].
Which phone sequences a language permits in an onset or coda is its
phonotactics — English allows the onset [str] of strike but forbids
[zdr]. These constraints can be captured by a finite-state or language model over
phone strings, and breaking a word into syllables is syllabification.
A worked transcription
Put the pieces together on one word. Take greasy, from the TIMIT sentence she
had your dark suit in greasy wash water. Its ARPAbet transcription is [g r iy s iy] — the IPA [g r i s i]. Read left to right against the machinery above: [g]
is a voiced velar stop (back of tongue on the velum, folds vibrating); [r] is a
voiced alveolar approximant; [iy] is a high front tense vowel; [s] is a
voiceless alveolar sibilant fricative; [iy] again. Every phone is now a triple of
articulatory facts, not a letter. Syllabify it and the two vowels each anchor a
syllable: the first is [g r] onset [iy] nucleus, the second [s] onset [iy] nucleus, with no coda on either. The onset [g r] is legal English
phonotactics; [r g] would not be.
Prosody
Everything so far concerns the identity of sounds. Prosody concerns their melody and rhythm: the use of F0, energy, and duration to carry meaning that the phone string alone does not.3 Prosody marks the difference between a statement and a question, signals which word matters, conveys affect (happiness, surprise, anger), and manages turn-taking in conversation.
Prominence, stress, and schwa
Some words, and some syllables within them, sound more prominent — perceptually salient — than others. A speaker makes a syllable prominent by making it louder, longer, or by varying its F0. We mark prominence with a pitch accent; a word bearing an accent is accented. But not every syllable can host the accent: it must land on the syllable that carries lexical stress, a fixed property of the word's dictionary pronunciation. Surprised is stressed on its second syllable, so an accent on surprised strengthens that syllable, never the first.
Lexical stress is marked in dictionaries — the CMU dictionary tags each vowel 0
(unstressed) or 1 (stressed): table is [T EY1 B AH0 L]. Stress can even
distinguish words: the noun content is [K AA1 N T EH0 N T], the adjective
[K AA0 N T EH1 N T]. An unstressed vowel is often reduced to a weaker,
centralized quality, most commonly the schwa [ax] (the IPA ə) — the second
vowel of parakeet [p ae r ax k iy t]. The result is a continuum of prominence:
accented, then stressed, then full-vowel, then reduced.
Tune and intonation
Two utterances with identical words and stress can still differ in tune — the rise and fall of F0 over the utterance.3 The classic example is the contrast between a statement and a yes-no question on the same words: a final F0 rise (a question rise) versus a final drop (a final fall).
English uses tune widely: besides the question rise, a list of comma-separated
nouns often carries a small continuation rise after each item, and there are
characteristic contours for contradiction and surprise. Utterances also have
prosodic structure — words group into intonation phrases with audible
breaks between them, often at commas and often aligned with syntactic constituents.
One influential labeling scheme, ToBI (Tone and Break Indices), tags each
accented word with one of a handful of pitch-accent types (H*, L*, L*+H, …)
and each phrase with a boundary tone (L-L% for the declarative fall, H-H% for
the question rise). Predicting these boundaries automatically matters for
text-to-speech, and models are trained on corpora such as the Boston University
Radio News Corpus.
From articulation to acoustics
Everything so far concerns how a phone is made and what a listener perceives — the articulatory and prosodic side of speech. But a microphone never sees a tongue or a glottis; it sees only the pressure wave those movements produce. The second part follows the sound out of the mouth: the waveform, its decomposition into a spectrum, the source-filter model that explains why each vowel has its own formants, and the spectrogram that a recognizer actually reads. This continues in Acoustic Phonetics.
Footnotes
- Jurafsky & Martin, Speech and Language Processing (3rd ed.), Ch. 25 — Phonetics; §25.1 Speech Sounds and Phonetic Transcription: phones as the units of pronunciation, the IPA and the ASCII ARPAbet, and the opaque letter-to-sound mapping of English. ↩ ↩2
- Jurafsky & Martin, §25.2 — Articulatory Phonetics: the vocal organs, voicing at the glottis, place and manner of articulation for consonants, the height/backness/rounding parameters for vowels, and syllable structure. ↩ ↩2 ↩3 ↩4
- Jurafsky & Martin, §25.3 — Prosody: prominence, pitch accent, lexical stress and vowel reduction to schwa, prosodic phrasing, tune (question rise vs. final fall), and the ToBI labeling system. ↩ ↩2
╌╌ END ╌╌