Speech/Phonetics

Lesson 8.11,896 words

Phonetics

Before a recognizer can read speech it has to know what speech is. This first part covers the linguistic substrate: phones and their transcription in the IPA and ARPAbet; articulatory phonetics — how the vocal tract shapes airflow into consonants and vowels; and prosody — stress, tune, and the F0 contour.

╌╌╌╌

A written character like p or a is already a scientific model of speech: a claim that the continuous, gliding sound of a spoken word decomposes into a small set of discrete, reusable units. That claim is old — the earliest writing systems already used some symbols for sounds rather than whole words — and it is the same claim that makes speech recognition and text-to-speech possible.1 Phonetics is the study of those units: what they are, how the human vocal tract produces them, and what their acoustic signal looks like once it reaches a microphone.

This lesson is the linguistic and physical groundwork for the recognizer that follows. That recognizer never sees a phone directly; it sees a waveform, turns it into a spectrogram, and learns to read the acoustic traces that phones leave there. Everything here — the vocal tract as a filter, formants as resonances, F0 as the vibration of the vocal folds — explains where those traces come from. The log-mel front end of a modern ASR system is a mechanized, learned version of exactly the acoustic analysis this chapter does by hand.

The speech chain. A speaker's articulators shape airflow into a pressure wave; a microphone records it as a waveform; acoustic analysis turns the waveform into a spectrogram, which is what a recognizer reads. This lesson walks the chain left to right.

Speech sounds and phonetic transcription

We represent the pronunciation of a word as a string of phones — speech sounds, each written with a symbol adapted from the Roman alphabet.1 A phone is a physical event, an actual sound produced by a vocal tract, and a symbol in square brackets, like [p], denotes a specific articulation.

The standard cross-linguistic system is the International Phonetic Alphabet (IPA), first developed in 1888 and still evolving, which has a symbol for every sound in every documented language. For English work it is common to use the ARPAbet instead: a plain-ASCII phonetic alphabet covering the American-English subset of the IPA, which is convenient precisely because it needs no special glyphs. The two are in one-to-one correspondence over their shared inventory; the ARPAbet [sh] is the IPA ʃ, [iy] is i, [aa] is ɑ.

The reason we transcribe at all — rather than just spelling — is that English orthography is a poor guide to sound. The mapping from letters to phones is opaque: one letter spells many sounds, and one sound is spelled many ways.

English spelling hides the phones. The single letter c is the phone [k] in cougar but [s] in cell; conversely the phone [k] surfaces as c, k, ck, cc, and the x of fox. Transcription records the sound, not the spelling.

Many languages — Spanish, for one — are far more transparent, but for English, talking about sound requires a notation for sound. The table below is a working subset of the ARPAbet, with the IPA equivalent and an example word.

ARPAbetIPAExampleARPAbetIPAExample
[p]pparsley[iy]ilily
[t]ttea[ih]ɪlily
[k]kcook[eh]ɛpen
[b]bbay[ae]æaster
[d]ddill[aa]ɑpoppy
[g]ggarlic[ao]ɔorchid
[m]mmint[uh]ʊwood
[n]nnutmeg[uw]utulip
[ng]ŋbaking[ah]ʌbutter
[f]fflour[er]ɝbird
[s]ssoup[ax]əlotus
[sh]ʃsquash[ay]iris
[ch]cherry[aw]flower
[l]llicorice[oy]ɔɪsoil
[r]rrice[ow]lotus

The left half of the table is consonants, the right half vowels; the rest of the lesson explains what that division means physically. A companion notion is the phoneme, the abstract sound category a language treats as one unit, versus the allophone, the concrete variant that actually surfaces in context. English treats the aspirated [p] of pin and the unaspirated [p] of spin as the same phoneme even though they are physically distinct phones; the two are allophones of one phoneme. Phonetics studies the phones; phonology studies which distinctions a language treats as meaningful.

Articulatory phonetics

Articulatory phonetics studies how phones are produced as the organs of the mouth, throat, and nose modify airflow from the lungs.2 Sound starts as air expelled from the lungs up the trachea (windpipe). At the top of the trachea sits the larynx — the Adam's apple — which holds two folds of muscle, the vocal folds. The gap between them is the glottis.

The vocal organs in side view. Air from the lungs passes the vocal folds at the glottis, then through the oral and nasal tracts, which act as the resonating cavities that give each phone its acoustic shape.

If the folds are held close together, airflow makes them vibrate, and the sound is voiced; if they are apart, no vibration occurs and the sound is unvoiced (voiceless). All English vowels are voiced, as are [b] [d] [g] [v] [z]; [p] [t] [k] [f] [s] are their voiceless counterparts. Above the larynx is the vocal tract, split into the oral tract (out the mouth) and the nasal tract (out the nose). Sounds routed partly through the nose — [m] [n] [ng] — are nasal.

Phones split into two classes. Consonants restrict or block the airflow; vowels leave it comparatively open, are usually voiced, and are louder and longer. Semivowels like [y] and [w] sit between: voiced like vowels but short and non-syllabic like consonants.

Consonants: place of articulation

Because a consonant is made by pinching the airflow somewhere, we can sort consonants by where that pinch happens — the place of articulation.2

The major English places of articulation, ordered front (lips) to back (glottis) along the vocal tract. The place is the point of maximum constriction.

Working front to back: labial sounds close the two lips (bilabial [p] [b] [m]) or press the lower lip to the upper teeth (labiodental [f] [v]). Dental sounds [th] [dh] place the tongue at or between the teeth. Alveolar sounds [s] [z] [t] [d] press the tongue tip to the alveolar ridge, the bump just behind the upper teeth. Palatal (strictly, palato-alveolar) sounds [sh] [ch] [zh] [jh] and the palatal [y] raise the tongue toward the hard palate. Velar sounds [k] [g] [ng] press the back of the tongue against the velum, the soft palate at the very back. A glottal stop closes the glottis itself.

Consonants: manner of articulation

The place says where; the manner of articulation says how the constriction is made — a full stop, a narrow hiss, or a gentle approach.2 Place and manner together, plus voicing, almost always identify a consonant uniquely.

Manner of articulation ordered by how completely the airflow is obstructed, from a total stop through turbulent fricatives to open approximants. Each manner is illustrated by the airstream through the constriction.

A stop (plosive) blocks the air completely, then releases it in a burst; the silent block is the closure and the burst is the release. English has voiced stops [b] [d] [g] and voiceless [p] [t] [k]. Nasals [n] [m] [ng] lower the velum so air escapes through the nose. In a fricative the channel is narrowed but not closed, and the turbulent air hisses: [f] [v] [th] [dh] [s] [z] [sh] [zh], of which the loud, high ones [s] [z] [sh] [zh] are sibilants. A stop released straight into a fricative is an affricate, [ch] and [jh]. In an approximant the articulators come close but not close enough for turbulence: [y] [w] [r] [l], with [l] a lateral because air flows over the lowered sides of the tongue. A tap or flap [dx] is a single quick brush of the tongue against the alveolar ridge — the middle consonant of lotus [l ow dx ax s] in most American dialects.

Vowels and the vowel space

Vowels have no real constriction, so they are described not by place and manner but by the position of the tongue body and the shape of the lips.2 Three parameters do the work: height (how high the tongue's high point sits), frontness/backness (whether that high point is toward the front or back of the mouth), and rounding (whether the lips are pursed). For [iy] (beet) the tongue is high and front; for [uw] (boot) high and back; for [ae] (bat) low and front.

The English vowel space, a schematic quadrilateral with tongue height on the vertical axis and frontness on the horizontal. Point vowels hold a fixed tongue position; diphthongs (drawn as arrows) glide from one position to another.

The space is drawn as a quadrilateral because that roughly matches both tongue positions and, more accurately, the acoustics we get to below. A vowel with a fixed tongue position is a plain vowel; one in which the tongue moves markedly during the vowel is a diphthong, drawn here as an arrow from start to end — English is rich in them ([ay] [aw] [oy]). Certain back vowels ([uw] [ao] [ow]) are pronounced with the lips rounded.

Syllables

Consonants and vowels group into syllables. The vowel at the core is the nucleus; the consonants before it are the onset; the consonants after it are the coda; and the nucleus plus coda together form the rime. Dog[d aa g] has onset [d], nucleus [aa], coda [g].

Syllable structure for ham, drawn as a tree. The onset precedes the nucleus; the nucleus plus coda form the rime; sigma marks the whole syllable.

Which phone sequences a language permits in an onset or coda is its phonotactics — English allows the onset [str] of strike but forbids [zdr]. These constraints can be captured by a finite-state or language model over phone strings, and breaking a word into syllables is syllabification.

A worked transcription

Put the pieces together on one word. Take greasy, from the TIMIT sentence she had your dark suit in greasy wash water. Its ARPAbet transcription is [g r iy s iy] — the IPA [g r i s i]. Read left to right against the machinery above: [g] is a voiced velar stop (back of tongue on the velum, folds vibrating); [r] is a voiced alveolar approximant; [iy] is a high front tense vowel; [s] is a voiceless alveolar sibilant fricative; [iy] again. Every phone is now a triple of articulatory facts, not a letter. Syllabify it and the two vowels each anchor a syllable: the first is [g r] onset [iy] nucleus, the second [s] onset [iy] nucleus, with no coda on either. The onset [g r] is legal English phonotactics; [r g] would not be.

The word greasy, [g r iy s iy], split into two syllables. Each vowel is a nucleus; the consonants before it are the onset. Sigma marks a syllable. The word has no coda, so each rime is just its nucleus.

Prosody

Everything so far concerns the identity of sounds. Prosody concerns their melody and rhythm: the use of F0, energy, and duration to carry meaning that the phone string alone does not.3 Prosody marks the difference between a statement and a question, signals which word matters, conveys affect (happiness, surprise, anger), and manages turn-taking in conversation.

Prominence, stress, and schwa

Some words, and some syllables within them, sound more prominent — perceptually salient — than others. A speaker makes a syllable prominent by making it louder, longer, or by varying its F0. We mark prominence with a pitch accent; a word bearing an accent is accented. But not every syllable can host the accent: it must land on the syllable that carries lexical stress, a fixed property of the word's dictionary pronunciation. Surprised is stressed on its second syllable, so an accent on surprised strengthens that syllable, never the first.

Lexical stress is marked in dictionaries — the CMU dictionary tags each vowel 0 (unstressed) or 1 (stressed): table is [T EY1 B AH0 L]. Stress can even distinguish words: the noun content is [K AA1 N T EH0 N T], the adjective [K AA0 N T EH1 N T]. An unstressed vowel is often reduced to a weaker, centralized quality, most commonly the schwa [ax] (the IPA ə) — the second vowel of parakeet [p ae r ax k iy t]. The result is a continuum of prominence: accented, then stressed, then full-vowel, then reduced.

Tune and intonation

Two utterances with identical words and stress can still differ in tune — the rise and fall of F0 over the utterance.3 The classic example is the contrast between a statement and a yes-no question on the same words: a final F0 rise (a question rise) versus a final drop (a final fall).

The same words carry two tunes. A final F0 rise (blue) signals a yes-no question; a final fall (red) signals a declarative statement. Only the pitch contour over the last words differs.

English uses tune widely: besides the question rise, a list of comma-separated nouns often carries a small continuation rise after each item, and there are characteristic contours for contradiction and surprise. Utterances also have prosodic structure — words group into intonation phrases with audible breaks between them, often at commas and often aligned with syntactic constituents. One influential labeling scheme, ToBI (Tone and Break Indices), tags each accented word with one of a handful of pitch-accent types (H*, L*, L*+H, …) and each phrase with a boundary tone (L-L% for the declarative fall, H-H% for the question rise). Predicting these boundaries automatically matters for text-to-speech, and models are trained on corpora such as the Boston University Radio News Corpus.

From articulation to acoustics

Everything so far concerns how a phone is made and what a listener perceives — the articulatory and prosodic side of speech. But a microphone never sees a tongue or a glottis; it sees only the pressure wave those movements produce. The second part follows the sound out of the mouth: the waveform, its decomposition into a spectrum, the source-filter model that explains why each vowel has its own formants, and the spectrogram that a recognizer actually reads. This continues in Acoustic Phonetics.

Footnotes

  1. Jurafsky & Martin, Speech and Language Processing (3rd ed.), Ch. 25 — Phonetics; §25.1 Speech Sounds and Phonetic Transcription: phones as the units of pronunciation, the IPA and the ASCII ARPAbet, and the opaque letter-to-sound mapping of English. 2
  2. Jurafsky & Martin, §25.2 — Articulatory Phonetics: the vocal organs, voicing at the glottis, place and manner of articulation for consonants, the height/backness/rounding parameters for vowels, and syllable structure. 2 3 4
  3. Jurafsky & Martin, §25.3 — Prosody: prominence, pitch accent, lexical stress and vowel reduction to schwa, prosodic phrasing, tune (question rise vs. final fall), and the ToBI labeling system. 2

╌╌ END ╌╌