What Actually Separates IELTS Speaking Band 6 from Band 7?
Band 7 is band 6 plus some sustained features of band 8. Here's that gap in plain terms, with a self-check for which side your own speaking currently sits on.
The jump from band 6 to band 7 in IELTS Speaking feels vague from the outside — "more fluent," "better pronunciation" — but the actual descriptor gap is narrower and more specific than that framing suggests. Across every criterion, band 6 is "mixed control, sometimes effective" and band 7 is "generally effective, sustained more of the time." That's a consistency gap, not a knowledge gap, and it shows up as one or two recurring habits rather than a general shortfall.
30-second version: Band 6 answers work sometimes; band 7 answers work most of the time. Find the one habit that breaks down under pressure — a recurring mispronunciation, or fluency that collapses in Part 3 — and fix that specifically rather than practising everything equally.
What does "sustained" actually mean in practice?
It means the examiner is listening across all fourteen-ish minutes of the test, not just your best thirty seconds. A candidate who gives one beautifully fluent, clearly pronounced Part 2 answer and then goes flat and hesitant in Part 3 is, by definition, describing band 6 — "not sustained" is close to the exact wording of the gap. This is why band 6-to-7 practice should weight the harder, later parts of the test more heavily than the easier warm-up, even though the warm-up is where most candidates feel most comfortable rehearsing.
Where does the gap usually actually live?
In practice, almost always in one of two places: a single recurring pronunciation habit that shows up in high-frequency words, or intonation that flattens once the answer gets longer than thirty seconds. Rarely is it a broad deficit across every feature at once — that pattern is more typical of band 5 and below. Finding which of the two applies to you (often both) is more useful than generic "practise more" advice, because it tells you exactly where to spend your limited prep time.
A self-check to find your specific gap
| Question | If mostly yes | If mostly no |
|---|---|---|
| Does the same sound substitution show up across several unrelated words? | Likely a recurring individual-sound issue — drill that one contrast specifically | Individual sounds are probably not your ceiling |
| Does your pace and pitch stay similar in Part 3 as it was in Part 1? | You're likely sustaining — this isn't your gap | Fluency probably isn't sustained — this is worth targeted work |
| Do pauses happen between ideas, or mid-word while searching? | Chunking is working in your favour | Practise structuring answers before speaking, not mid-sentence |
How do I fix a recurring pronunciation habit specifically?
Isolate the sound, then rebuild it in increasingly natural contexts: the sound alone, then in a word, then a minimal pair against the sound you're substituting it for, then a full sentence at speed. sayit's phoneme pages cover the individual articulation for the most commonly flagged contrasts, and a recorded IELTS-style answer will show you directly whether the fix is holding up once you're focused on content again, not just on the isolated sound.
How do I fix flattening fluency in the longer parts?
Practise the length that's actually breaking down, not the length you're comfortable with. If Part 3 is where things flatten, don't spend your practice time on more Part 1 warm-up questions — record unscripted two- and three-minute discussion answers specifically, and listen back for where the pace or pitch range noticeably drops. This is uncomfortable to practise precisely because it's the part that's already hard, which is exactly why most candidates under-practise it relative to how much it's actually costing them.
Try it
Record a Part 1 answer and a Part 3 answer back to back in sayit and compare the two — that gap, more than either one alone, is the band 6-to-7 signal worth watching. Full timed practice across all three parts is at /ielts.
Hear exactly which sounds to fix.
Say one sentence and get sound-by-sound feedback in seconds. No install, no card.