Trial Lesson With an English Tutor: What to Expect and How to Judge It
What a real trial lesson contains, why mine costs $25 instead of nothing, and an eight-point scorecard for judging any tutor after one hour.

Most English listening practice fails for one reason: you listen to understand the message, and understanding the message teaches you nothing about how the sounds were made. Listening only improves your speaking when you listen for the mechanics, and that takes 20 focused minutes, not three hours of podcasts in the background.
Passive listening builds comfort with English, not skill. To improve speaking, take a clip of 60 to 90 seconds, listen once without help, then with the transcript, and mark every place where what you heard did not match what was written. Those gaps are your pronunciation targets: flap Ts, dropped sounds, linked words, unstressed syllables. Then say the clip back, recording yourself, until your version has the same rhythm. Done daily for 20 minutes, this changes both what you can hear and what you can say within a few weeks, because the two skills share the same map.
I hear this in almost every first lesson: "I watch everything in English, I listen to podcasts on my commute, and I still cannot understand Americans on the phone." Then the student speaks, and I hear the same thing every time. Their English is full of words they have read a thousand times and never actually heard as sounds.
Here is why. When you listen for meaning, your brain does what it is built to do: it grabs the content words, predicts the rest, and throws the sound away. You finish a 40-minute episode knowing what the guest thinks about remote work and having zero memory of how she said "a lot of" (it was "a lodda"), "did you" ("didja") or "comfortable" (three syllables, "KUMF-ter-bul", not four). The meaning arrived. The mechanics never did.
That is fine if your goal is to follow a podcast. It is useless if your goal is to be understood in a meeting, because speaking requires the mechanics. You cannot produce a rhythm you have never consciously noticed, and you will not notice it by listening for content.
Perception and production are linked more tightly than most learners expect. If your first language does not have the vowel in ship versus sheep, you do not just struggle to say it. You struggle to hear it, because your ear files both sounds into the same box. The same goes for the flap T in water, the reduced vowel in the second syllable of problem, and the way "want to" collapses into "wanna".
This has a practical consequence. Every time you train your mouth to make a distinction, your ear gets better at catching it in fast speech. And every time you train your ear to catch a reduction, saying it yourself becomes easier. Listening practice that includes your mouth is twice as efficient as listening alone, which is why every routine in this article ends with you speaking.
It also explains why fast American speech feels fast. It is not faster in words per minute than most languages. It is compressed: unstressed syllables shrink to a schwa, consonants link across word boundaries, and whole sounds vanish. I break down the specific patterns in why Americans talk so fast (they don't) and in connected speech and linking. Once you can name those patterns, you start hearing them everywhere, and your listening jumps.
This is the routine I give students between lessons. It uses one short clip, and it takes 20 minutes only if you actually stop the audio and do the steps. Skipping the recording step at the end cuts the benefit in half.
| Minutes | Step | What you are training |
|---|---|---|
| 0-2 | Play the 60-90 second clip once. No transcript. Write down what you understood in one sentence. | Gist, and an honest baseline |
| 2-6 | Play it again in 10-second pieces. After each piece, write what you heard, word for word, even if it is nonsense. | Sound-level attention |
| 6-10 | Open the transcript. Compare it with your notes. Circle every word you heard wrong or missed completely. | Finding your personal blind spots |
| 10-13 | For each circled word, replay it three times and answer: what actually happened to the sound? A flap? A dropped T? Two words linked? | Naming the pattern |
| 13-18 | Shadow the clip: speak along with the audio, half a second behind, transcript in front of you. Two passes. | Rhythm, stress, linking in your own mouth |
| 18-20 | Record yourself saying two of the circled sentences without the audio. Listen back once. | Production, and proof of change |
The circled words are gold. After a week you will see the same categories repeating: maybe every missed word had a flap T, or every missed phrase was a preposition swallowed between two content words. That list tells you exactly what to practice, and it is the first thing I ask to see when a student brings their week's notes to a lesson.
The biggest mistake is choosing material by topic instead of by usefulness. Interesting is not the same as trainable. Use these filters.
Short. Sixty to ninety seconds is enough. A clip you can repeat six times teaches more than an episode you hear once. YouTube Shorts and the first two minutes of an interview are ideal. My own channel is built around this: each video isolates one sound or phrase so you can loop it.
Unscripted, or scripted to sound unscripted. News anchors are too clean. You want real speech: podcast conversations, interviews, vlogs, sitcom dialogue. That is where "gonna", "didja" and "whaddaya" live, and that is what you will face on a call.
With a transcript you trust. Auto-generated captions are good enough for most content now, but check them against the audio. If the caption says "want to" and the speaker clearly said "wanna", the caption is telling you what was meant, not what was said. Your job is to notice the difference.
One voice for a week, then change. Stay with one speaker long enough to learn their patterns, then switch. Men and women, fast and slow, Texas and California. If you only ever train on one clear voice, the first mumbling colleague will undo your confidence.
Slightly too hard. If you understand 100 percent on the first pass, there is nothing to circle. If you understand 30 percent, you will quit. Aim for a clip where the first blind pass gives you around 70 to 80 percent, and the transcript reveals the rest.
A transcript is a diagnostic tool, not a reading exercise. Students often open it first, read along, feel that they understood everything, and learn nothing. Reading while listening is reading. The order matters: ears first, eyes second, then ears again with a specific question.
When you compare your notes with the transcript, sort each miss into one of four buckets. I use these in lessons because they map directly onto what we fix next.
| Bucket | Example | What it usually means |
|---|---|---|
| Word I did not know | "ballpark figure" | Vocabulary. Add it as a phrase, not a word. |
| Word I know but did not recognize | "comfortable" heard as two unknown words | Your stored pronunciation is wrong. Highest priority. |
| Words merged together | "What are you doing" heard as "whacha doin" | Connected speech. Learn the reduction as one unit. |
| Small word vanished | "a cup of coffee" heard as "a cuppa coffee" | Function word reduction. Normal, and you should do it too. |
The second bucket is the one that changes speaking fastest. Every word there is a word you have been mispronouncing, probably for years, and nobody told you. When students find vegetable, colleague, chocolate or interesting in that bucket, the fix is usually one syllable fewer and a different stressed syllable. If stress is new to you, the word stress rule that fixes 100 words is the place to start.
Subtitles in your own language. This is reading a translation while sound plays in the background. Nothing about English is being learned. If you need subtitles to enjoy a show, enjoy the show, but do not count it as practice.
Only clear, slow, teacher voices. Graded listening has its place at B1. At B2 and above, it becomes a comfort zone. Real speech is where the misses are, and the misses are the lesson.
Never producing. A student can do active listening perfectly for a month and still not change their speaking if they never say the clip back. The mouth has to move. If you feel self-conscious doing this alone, shadowing step by step covers how to set it up so it feels normal after a week.
Trusting your own ear on your own voice. This is the hard one. When you listen to your recording, you hear what you intended, not what you produced. That filter is exactly why I exist as a teacher: in a lesson, I hear that your "better" came out with a hard British T, or that your "can't" was indistinguishable from "can", and I can show you on the spot. Students like So Yeon and Luca have told me the most useful moment of a lesson was hearing me repeat their sentence back exactly as they said it, next to the American version. That comparison is not something you can do for yourself, at least not at first.
You can run this routine alone, and you should. What you cannot do alone is verify. Once a week, bring three things to a lesson: your circled-word list, one recording, and one question about a sound you keep missing. In 50 minutes we confirm which patterns are real, fix the ones you cannot hear, and choose next week's clip based on what your ear needs, not what the algorithm suggests. Students who do this typically arrive at the fourth or fifth lesson saying that phone calls have become noticeably easier, and their own speech has started to pick up the reductions they were circling. You can see how lessons are priced on the pricing section.
Choose one 60 to 90 second clip from a real conversation with a transcript. Run the 20-minute routine five days in a row on the same clip. Yes, the same one. By day three the blind pass will be nearly perfect and the shadowing will start sounding like the speaker. Keep your circled-word list in one place. On day five, record the full clip without the audio and compare it with your day-one recording. If you hear a difference, you have found the method that works. If you cannot tell, book a first lesson and bring both recordings; telling you exactly what changed, and what did not, is what I do in the first fifteen minutes.
Twenty focused minutes with a transcript and a recording step beats hours of background audio. If you have more time, add passive listening on top for comfort and vocabulary, but do not swap it for the active session. The active twenty minutes is where speaking improves.
English subtitles, sometimes. Use them on the second viewing to check what you missed, not on the first. Subtitles in your own language do nothing for your listening; count that time as entertainment, not practice.
Podcast hosts speak to a microphone, with good audio and a consistent pace. Real people mumble, overlap, turn away and use more reductions. Train on unscripted clips with imperfect audio, such as interviews and vlogs, so the gap between practice and reality shrinks.
One speaker for about a week, then switch. Staying with one voice lets you learn their patterns and see real progress on the same clip. Switching regularly stops you from tuning to a single accent and getting lost when a new colleague speaks.
Only when you say the clip back. Listening alone improves recognition. Listening plus shadowing plus recording moves the patterns into your mouth, because hearing a reduction and producing it use the same mental map. The recording step is not optional.
That is normal and it is exactly where a tutor helps. I will exaggerate the two sounds, show you the mouth position, and have you produce them until your ear catches up with your mouth. Perception usually follows production within a few sessions.
What a real trial lesson contains, why mine costs $25 instead of nothing, and an eight-point scorecard for judging any tutor after one hour.
The rule, the exceptions, the spelling traps and a drill list for the sound that makes water, better and city sound American instead of textbook.
Americans do not say the T in button or mountain. They stop the air in the throat. Here is the rule, the exceptions and a five-minute routine.
Private American English coaching with Ashley Curry: pronunciation, accent and confidence, built around your goals.