Guides · 9 min read

Connected speech in English: linking, elision and assimilation, and why fast English sounds like one long word

Native English runs words together by rule, not by carelessness. Here are the four processes — linking, elision, assimilation and weak forms — with the patterns behind each and a drill you can record and score.

There is a particular kind of frustration that comes from reading an English transcript after failing to understand the audio. Every word is familiar. You know all of them. And yet what you heard was "wʤədu" rather than "what did you do".

That gap is not vocabulary, speed or hearing. It is that written English shows words with spaces between them and spoken English does not. When words touch, their edges change — and they change according to a small number of rules that can be learned in an afternoon and then practised for a month.

Learning these rules does two things at once, which is why they are worth more time than most learners give them. They make you easier to understand, because your speech starts to follow the shape listeners expect. And they transform your listening, because you finally know what to listen for.

The four processes

Everything that happens between words in English falls into four categories.

ProcessWhat happensExample
Weak formsFunction words reduce to a schwa"fish and chips" to "fish ən chips"
LinkingA final consonant joins the next vowel"an apple" to "a napple"
ElisionA sound disappears"next day" to "nex day"
AssimilationA sound changes to match its neighbour"ten boys" to "tem boys"

They happen simultaneously, which is why fast English feels overwhelming. Taken one at a time they are all quite simple.

1. Weak forms

The foundation. English has two pronunciations for most of its small grammatical words: a strong form used in isolation or under emphasis, and a weak form — almost always with a schwa — used everywhere else.

WordStrongWeak
and/ænd//ən/ or /n/
to/tuː//tə/
of/ɒv//əv/ or /ə/
for/fɔː//fə/
can/kæn//kən/
was/wɒz//wəz/
them/ðem//ðəm/ or /əm/
you/juː//jə/

If you use strong forms everywhere, two things go wrong. You sound slow and effortful, and — more seriously — you signal emphasis you did not mean, because a full vowel on a function word is exactly how English marks contrast. "I CAN do it" answers an accusation. "I can do it" answers a question. The whole layer is covered in the schwa.

2. Linking

When one word ends in a consonant and the next begins with a vowel, English moves the consonant across the boundary. It genuinely resyllabifies: the consonant becomes the first sound of the next syllable.

  • an apple → "a napple"
  • turn it off → "tur ni toff"
  • pick it up → "pi ki tup"
  • far away → "fa raway"

This is why word boundaries disappear in fast speech: they are actually gone. The syllables have been rebuilt across them.

There are two special cases worth knowing.

Linking /r/. In non-rhotic accents (most British, Australian, and some American), a written r at the end of a word is silent — except when the next word starts with a vowel, where it comes back. Far is /fɑː/, but far away is /fɑːr əˈweɪ/, with the r reappearing to bridge the two words. Rhotic speakers pronounce the r either way, so this one is accent-dependent.

Vowel-to-vowel glides. When one word ends in a vowel and the next starts with one, English inserts a tiny /j/ or /w/ to bridge the gap: I always becomes "I-y-always", go out becomes "go-w-out", the end becomes "thee-y-end". You do not need to force this — it appears on its own once you stop putting a glottal stop between the words.

3. Elision

Sounds get dropped, most often when three consonants collide.

/t/ and /d/ between consonants is the big one. In next day, the /t/ sits between /s/ and /d/ and disappears: "nex day". Same in:

  • last night → "las night"
  • must be → "mus be"
  • old man → "ol man"
  • sandwich → "sanwich"
  • friendship → "frienship"

/h/ in unstressed pronouns. He, him, her, his lose their /h/ when unstressed and not at the start of a phrase: "tell him" becomes "tell im", "ask her" becomes "ask er", "give his book" becomes "give iz book". This one is enormously common and almost never taught, and it accounts for a large share of the pronouns learners fail to hear.

Whole syllables. Frequent long words lose an unstressed syllable entirely: comfortable is normally two-and-a-bit syllables ("comf-tuh-bul"), interesting is three ("in-truh-sting"), vegetable is three ("vej-tuh-bul"), every is two ("ev-ree"). Saying all the written syllables is a reliable non-native marker.

4. Assimilation

A sound changes to become more like its neighbour, because that is less work for the mouth.

Place assimilation is the most frequent. A final /n/, /t/ or /d/ takes on the position of the following consonant:

  • ten boys → "tem boys" (n becomes m before b)
  • ten girls → "teng girls" (n becomes ŋ before g)
  • that boy → "thap boy"
  • good girl → "googgirl"

Coalescence — two sounds fusing into a third — is the one that produces the most famous unrecognisable phrases:

WrittenSpoken
did you"dijoo" /dɪdʒə/
would you"wudjoo"
don't you"downchoo"
what you"watchoo"
miss you"mishoo"
as usual"azhoozhual"

Did you is /d/ plus /j/ becoming /dʒ/, the sound at the start of jam. Don't you is /t/ plus /j/ becoming /tʃ/, the sound at the start of chair. Once you know those two equations, a whole family of previously incomprehensible phrases becomes readable.

Should you produce all of this, or just recognise it?

Honest answer: the priorities are different for listening and speaking.

For listening, learn all four. Recognition is where connected speech pays off fastest. Knowing that "wʤədu" is what did you do is the difference between following a conversation and drowning in it, and it costs you nothing but attention.

For speaking, prioritise weak forms and linking. These two produce most of the rhythmic improvement and are low risk. Elision and assimilation are natural consequences of speaking at a normal rate — they emerge on their own once the first two are in place, and deliberately forcing them tends to produce a strange, over-casual imitation.

The order that works: reduce function words, link consonant to vowel, then let speed do the rest.

One caution, and it is the reason this guide keeps mentioning weak forms first. Connected speech is not licence to blur everything. Elision applies to specific sounds in specific environments — it does not mean word endings are optional. If you drop the /t/ in walked because you read that English elides /t/, you have deleted the past tense. The relevant distinction is in word endings for Mandarin and Vietnamese speakers.

The drill

Block 1 — weak forms in phrases. Say each one as a single unit, fast, with only one strong beat.

  • a cup of tea
  • fish and chips
  • I have to go
  • more than ever
  • some of them

Block 2 — linking. These should come out as one continuous run with no gaps.

  • turn it off
  • pick it up
  • an hour and a half
  • far away from here

Block 3 — sentences. Record these and listen for whether your words have joined up.

The first sentence contains coalescence (did you), linking (at the end) and a weak form (of) in nine words. If it comes out as nine separate words, that is your starting point.

What it looks like on a recording

Connected speech is easy to fake and hard to self-assess, because "fast and blurry" sounds like the target while actually being a different failure. The useful check is whether the right things reduced.

Recording the drill sentences in sayit gives you three things worth looking at. Words per minute against a target tells you whether you are actually connecting or just reading quickly — the two feel similar and are not. The stress score tells you whether there are clear strong beats for the weak syllables to sit between; a low stress score with a high word rate usually means everything got blurred equally, which is the failure mode. And the per-word verdicts tell you which words survived: content words should still be clear. If asked and night got flagged while him and to reduced, that is connected speech working. If it is the other way round, you have not connected your speech — you have just dropped consonants.

The same logic explains why practising this into a dictation app teaches you nothing: a transcript-based model reconstructs "what did you do" from the language model whether or not you produced anything like it. The mechanism is in the autocorrect problem.

Questions people actually ask

Is connected speech lazy or informal? No. It is the normal register of spoken English at every level of formality. News broadcasts, university lectures and job interviews are all full of weak forms and linking. What changes with formality is speed and the amount of elision, not whether the processes apply.

Do I need to sound this way to be understood? Not strictly. Careful, unconnected English is intelligible. But it is slower, harder to listen to for long stretches, and it signals effort. And you will still need to understand connected speech regardless, because everyone else uses it.

Which accent should I model? Pick one and be consistent. The linking-/r/ rule and some vowel reductions differ between British and American English; the four processes themselves are the same in both. Consistency matters more than the choice.

Why does listening feel harder than speaking? Because in speaking you choose the words and in listening you have to recover them from a signal that has been reshaped. Working on production is the fastest route into recognition: sounds you can make are much easier to hear.

Where to go next

Connected speech sits on top of rhythm, which sits on top of stress. If the layers below are shaky, start there: the schwa for vowel reduction, and word stress and intonation for the beat pattern that everything else attaches to. To practise by imitating a model rather than by rule, shadowing is the technique built for exactly this.

The quickest test of where you are: open sayit free — no install, no card — say "What did you do at the end of the day?" and check whether your pace and stress scores agree that you connected it, or whether you just read nine separate words quickly.

Free in your browser

Hear exactly which sounds to fix.

Say one sentence and get sound-by-sound feedback in seconds. No install, no card.