Connected Speech: Why You Can't Understand Fast English (and How to Speak It)

Connected Speech: Why You Can't Understand Fast English (and How to Speak It)

Native speakers are not speaking faster than you. They are joining words together, and if you learned English one word at a time, your ear has no map for the joined version. Connected speech, the linking of words in English, is that map.

The short answer

Connected speech is what happens when English words touch each other in a sentence. Consonants link onto the next vowel (an apple becomes a-napple), sounds disappear (next day becomes nex day), sounds change (did you becomes didja), and small words shrink (want to becomes wanna). You cannot understand fast English until you can hear these four things, and you will sound choppy until you can do them. The fix is not more vocabulary. It is 15 minutes a day of listening for the joins and then copying them out loud.

Why connected speech linking in English feels like a wall

When a student tells me "Americans talk too fast," I ask them to read one sentence aloud: What do you want to do tonight? They say seven clean words. Then I say it the way I would say it to a friend: whaddaya wanna do tonight. Same sentence, four beats. Nothing was fast. The words simply fused.

This is the core problem. Your brain stores English as separate dictionary entries: what, do, you. But native speech never delivers those entries one by one. It delivers chunks, and the boundaries between words are the first thing to disappear. So your ear hears whaddaya, searches the dictionary, finds nothing, and panics. Meanwhile you have missed the next three words.

In my lessons the turning point is almost always the same: the moment a student realizes they already know every word in a sentence they could not understand. Once you know the joining rules, "fast" English slows down by itself. Several students told me their first American meeting after we covered this was the first one where they did not need the transcript.

Rule 1: Consonant to vowel linking (the big one)

When a word ends in a consonant sound and the next word starts with a vowel sound, the consonant jumps forward and becomes the first sound of the next word. This is the most common join in English and the easiest to learn.

WrittenSpoken
an applea-napple
turn it offtur-ni-toff
pick it uppi-ki-dup
hold onhol-don
an hour agoa-nou-ra-go
this is itthi-si-zit

Notice pick it up. The T in it lands between two vowels and turns into the soft American flap T, so you get pi-ki-dup. Linking and the flap T work together constantly, and this is why get a sounds like gedda and a lot of sounds like a lodda.

The test I give students: say hold on and put your hand on your chin. If your jaw closes fully between the two words, you are separating them. Native speakers never close the gap. The D slides straight into the O.

Rule 2: Elision, the sounds that quietly leave

English drops sounds when three consonants pile up, especially T and D in the middle of a cluster. Nobody decided this. It is simply too much work for the tongue at conversational speed.

  • next day becomes nex day (the T between X and D disappears)
  • most people becomes mos people
  • I don't know becomes I dunno (the T and most of the vowel go)
  • friendship becomes frenship
  • first time becomes firs time

The H at the start of he, him, her and his also vanishes when the word is not stressed. Tell him I called comes out as tell-im I called, and Where is he comes out as where-zee. Students who listen for a clear H in every pronoun miss half the pronouns in a conversation.

Watch Ashley explain it: Why does TU sound like CH in English? · under a minute on YouTube

Rule 3: Assimilation, when two sounds make a new one

Some sounds change their identity to match a neighbor. The pattern that matters most for American English is T, D, S or Z followed by the Y sound at the start of you, your and yet.

CombinationResultExample
D + Yj as in judgedid you: didja, would you: wouldja
T + Ych as in churchdon't you: doncha, meet you: meetcha
S + Yshmiss you: mishoo, this year: thishear
Z + Yzh as in measureas you know: azhoo know

There is also a smaller change that happens before P, B and M: an N turns into an M. Ten minutes sounds like tem minutes, and in Boston sounds like im Boston. You do not need to produce this one on purpose. If your linking is smooth it happens by itself. But you do need to hear it, or in Boston will sound like a word you have never met.

Rule 4: Reductions, the small words that shrink

Function words carry grammar, not meaning, so English shrinks them to almost nothing. The vowel collapses into the schwa, the relaxed uh sound, and the word takes a fraction of the time. This is what produces the famous spellings you see in song lyrics.

  • want to becomes wanna: "I wanna go."
  • going to becomes gonna: "It's gonna rain." (Only for the future. "I'm going to Dallas" keeps its full form.)
  • got to becomes gotta, have to becomes hafta, has to becomes hasta
  • kind of becomes kinda, sort of becomes sorta, out of becomes outta
  • a cup of coffee becomes a cuppa coffee
  • let me becomes lemme, give me becomes gimme

One important distinction. Wanna and gonna are informal, and I would not write them in an email. But the vowel reductions inside to, for, and, of and can are not slang. They happen in a presidential speech and a job interview. Bread and butter is bread-n-butter for every American speaker, in every register. If you say a full and with a clear A each time, you sound like you are reading a list.

Reduced can deserves its own article because it causes real misunderstandings in meetings. I wrote one: can vs can't in American English. The short version is that can reduces to kn and can't never does.

How to hear it: the transcript method

You cannot copy what you cannot hear, so listening comes first. Passive listening will not do it. You need a short clip and its transcript, and you need to work with a pencil.

  1. Choose 30 to 60 seconds of natural American speech with an accurate transcript. Podcast interviews and my YouTube Shorts work well. Scripted news reads are too clean.
  2. Listen once without the transcript. Write down how much you caught, roughly, as a percentage.
  3. Read the transcript and mark every join you predict: draw an arc between a final consonant and the next vowel, cross out T and D in clusters, circle every to, and, of, you.
  4. Listen again with the transcript in front of you. Check your predictions. Every place you were wrong is a rule your ear has not learned yet.
  5. Listen a third time without the transcript. The percentage jumps. That jump is the whole method.

This is the part of my listening routine that students resist at first because it feels slow. It is slow for a week. Then the joins start jumping out at you on their own, in movies, in meetings, in the coffee line.

How to speak it: three drills that build the habit

Once you can hear the joins, you need to move them into your mouth. Knowing the rule for an apple and saying a-napple under pressure are two different skills, and the second one takes repetition.

Drill 1: the backchain. Take a phrase and build it from the end. For turn it off: say off, then toff, then ni-toff, then tur-ni-toff. Starting from the end forces the joins because you are never starting a word fresh. Five phrases, three times each, two minutes total.

Drill 2: the single breath. Pick a sentence of eight to ten words and say it on one continuous breath with no gaps, even if it sounds strange. I have got to get out of here in an hour becomes I've godda ge-doudda here i-na-nour. Record it. Play it back. If you hear tiny stops between words, do it again. Most of my students need three or four tries before the stops disappear.

Drill 3: shadowing. Play your marked-up clip and speak with the recording, half a second behind, matching the rhythm rather than the words. This is where connected speech becomes automatic. I have a full guide on the shadowing technique, including the three passes and how to choose audio. The important point here: shadow for the joins, not for the sounds. Let the individual vowels be imperfect. Keep the flow.

What students cannot do alone is hear which joins they are still missing in their own speech. You will be convinced you said a-napple when the recording says an. apple. This is the first thing I listen for in a lesson. I stop the student mid-sentence, play back the four words, and show them the gap. Once they hear it, they can fix it. Until then, they cannot.

The mistakes that make connected speech sound wrong

Learners who discover linking sometimes overdo it, and the result is worse than choppy English. Three things to watch.

Do not reduce content words. Nouns, main verbs, adjectives and adverbs keep their full vowels. I want to go to the store reduces want to and to the, but go and store stay full and clear. If you shrink everything, nobody can find the meaning.

Do not link across a pause. Linking happens inside a thought group, not across commas or breaths. When I arrived, everyone had left links when-I-arrived but does not link arrived to everyone, because there is a pause between them.

Do not write the reductions. Gonna and wanna belong in your mouth, not in your work email. I see this often with students who learn connected speech from song lyrics and then type I wanna confirm to a client.

For a wider view of how all of this fits with contractions, stress and rhythm, my article on understanding fast American speech covers the same ground from the listening side.

What to do this week

Do not try to learn every rule at once. Pick consonant-to-vowel linking, because it is the most frequent and the most audible. Each day for seven days: one 45-second clip, the transcript method (10 minutes), the backchain drill on five phrases from that clip (3 minutes), and one recording of the single-breath sentence (2 minutes). Fifteen minutes. On day seven, listen to the day-one clip again without the transcript and compare your percentage.

If you want someone to catch the gaps you cannot hear, book a first lesson. It is 50 minutes, $25, and I spend the first part of it listening to you speak naturally so I can tell you exactly which joins you are missing. Everything about how I work, and my prices for ongoing lessons, is on the pricing section.

Questions students ask

Is connected speech linking in English the same as slang?

No. Slang is vocabulary. Connected speech is how any vocabulary sounds when words are said together at normal speed. Linking, elision and vowel reduction happen in formal speeches, news broadcasts and job interviews. Only the written forms like "gonna" and "wanna" are informal.

Will I sound lazy or unprofessional if I link words?

The opposite. Separating every word is what sounds unnatural to an American ear, and it makes listeners work harder. Smooth linking with clear content words is exactly what professional native speakers do. Keep your nouns and verbs full and let the small words shrink.

Should I learn to hear connected speech before I try to produce it?

Yes, by about a week. Producing a join you cannot hear leads to guessing. Do the transcript method first, and start the speaking drills once you can predict most of the joins in a clip before you listen. Most of my students reach that point in five to seven days.

Why can I understand my colleagues but not American TV?

Your colleagues are probably slowing down for you and cutting reductions without knowing it. TV dialogue is written to sound like natural speech between friends, so it uses every join at full speed. Working with a transcript on 30-second TV clips closes this gap faster than anything else.

How long until connected speech becomes automatic in my own speaking?

Hearing it improves within one to two weeks of daily transcript work. Producing it without thinking takes longer, usually a couple of months of shadowing and feedback, because you are replacing a habit, not learning a fact. Students who record themselves progress noticeably faster than those who only practice out loud.

Ready to be understood the first time?

Private American English coaching with Ashley Curry: pronunciation, accent and confidence, built around your goals.