Two words, one sound apart.
Can you hear it?
- Free
- No account
- Listening test on every pair
Ear first, then mouth
Find the contrast your first language does not make — each group below says which ones typically merge it. Open a pair, run the five-round listening test, and only then record the two words. If your score on the listening test is low, that is the real diagnosis: your ear is filing both words in one drawer, and no amount of tongue placement will fix that until the ear separates them.
You do not need an account for any of it. The listening test is free and unlimited; the scored recording is one free take per device, and it shows you the full analysis rather than a preview of one.
A note on accents
A few of these contrasts have merged for some native speakers — cot and caught sound identical in most of North America, poor and pour in much of England. Where that is true the group says so. It is still worth knowing the contrast exists, because most of the English-speaking world keeps it.Short I vs long E
Long E is tense — lips spread in a smile, tongue pushed high and forward and held; short I is the same region with everything relaxed and over in an instant.
Short OO vs long OO
Long OO pushes the lips forward into a tight ring and holds it; short OO keeps the rounding loose and lets go almost immediately.
Short A vs short U
Short A drops the jaw and pushes the tongue forward and low; short U leaves everything slack in the middle of the mouth and barely opens the jaw.
Short A vs short E
Both are front vowels; short E sits at mid height with the jaw barely open, while short A drops the jaw a full finger's width lower.
F vs V
The mouth is identical — top teeth on the bottom lip — and only the voice changes: put a hand on your throat and V buzzes while F is silent.
B vs V
B closes both lips completely and pops them open; V never closes anything — the top teeth rest on the bottom lip and the air keeps moving.
S vs Z
Same narrow groove behind the teeth, same hiss position: S is pure air, Z turns the voice on so the hiss becomes a buzz you can feel in your throat.
S vs SH
S uses a narrow groove at the tip of the tongue with the lips spread; SH slides the whole tongue back a centimetre, widens the channel and pushes the lips forward.
Voiceless TH vs S
For TH the tongue tip comes forward to the teeth — you can see it — while for S it stays behind them near the ridge; TH is soft and spread, S is a thin, sharp hiss.
Voiceless TH vs T
T stops the air completely with the tongue on the ridge; TH never stops it — the tongue touches the teeth lightly and the air keeps hissing through.
Voiced TH vs D
D taps the ridge and blocks the air for a moment; voiced TH puts the tongue on the teeth and keeps a soft buzz running without ever closing.
Voiced TH vs Z
Both buzz, but TH buzzes at the teeth with a wide, soft tongue while Z buzzes behind them through a narrow groove.
R vs L
L presses the tongue tip firmly on the ridge and lets the voice spill round the sides; R touches nothing at all — the tongue curls back or bunches up and floats.
W vs V
W rounds the lips into a tight ring and glides, with the teeth nowhere near the lip; V rests the top teeth on the bottom lip and buzzes there without moving.
CH vs SH
CH starts with the tongue stopping the air against the ridge and then releases into the hush; SH never stops anything — it is the hush on its own.
J vs Y
J stops the air with the tongue blade on the ridge before it buzzes; Y never touches — the tongue rises toward the palate and glides straight on.
CH vs J
Identical mouth shape, identical release — only the voice differs, so J hums behind the stop and CH does not.
N vs NG
Both hum through the nose; N puts the tongue TIP on the ridge at the front, NG lifts the BACK of the tongue to the soft palate and leaves the tip lying down.
Short I vs short E
Short E opens the jaw about a finger's width more than short I and drops the tongue to mid height; both are quick, so the difference is entirely in how open you are.
Short O vs AW
Short O is quick with the jaw low and only a little rounding; AW is long, with the lips pushed firmly forward and the tongue a step higher.
Long O vs AW
AW is one steady rounded vowel held to the end; long O starts lower and slides — the lips tighten during the sound, which is what makes it a glide.
ER vs AW
ER is central and unrounded with the tongue bunched in the middle; AW pulls the tongue back and rounds the lips hard — if your lips are rounded on 'work' you have said 'walk'.
EAR vs AIR
Both glide back to the same relaxed centre; they differ only in where they START — EAR begins high on the 'sit' vowel, AIR begins lower on the 'bed' vowel.
P vs B
Identical lips: P releases with a puff of air and no voice, B releases gently with the voice already humming.
T vs D
Same tongue tip on the same ridge: T releases with a puff and no voice, D keeps the hum going straight into the vowel.
K vs G
Same seal at the back of the tongue: K bursts with a puff of air and no voice, G releases quietly with the voice already on.
Long A vs short E
Both start in the same place; long A keeps moving, gliding up toward 'ee' and closing the jaw, while short E stops dead on the first sound.
Long I vs long A
Both glide up to the same 'ee' ending; long I starts with the jaw wide open on 'ah', long A starts halfway up on 'eh'.
OW vs long O
Both end with the lips rounding; OW starts with the jaw wide open on 'ah', long O starts already half-rounded on 'oh'.
OY vs long I
Both glide up to 'ee'; OY starts with the lips firmly rounded and the tongue back, long I starts unrounded with the jaw wide open.
M vs N
Both hum through the nose; M closes the lips and does nothing with the tongue, N leaves the lips open and presses the tongue tip on the ridge.
AH vs short A
Both drop the jaw; AH pulls the tongue back and holds the sound long, short A pushes the tongue forward and keeps it bright and quick.
URE vs AW
URE starts high on the 'book' vowel and glides down to a relaxed centre; AW is one long rounded vowel that never moves.
ZH vs J
J stops the air with the tongue on the ridge before it buzzes; ZH is that same buzz with no stop in front of it.
ZH vs SH
Identical tongue and lips; only the voice changes, so ZH hums where SH is pure air.
Stress pairs and the schwa
The two words are spelled identically: the noun stresses the first syllable and reduces the second to a schwa, the verb stresses the second and reduces the first.
More on sounds and contrasts
The 44 sounds of English
The full chart, with a page on how to make every consonant, vowel and diphthong.
A minimal-pair practice plan
Which contrasts matter for your first language, and a four-step drill protocol.
The hardest English sounds, ranked
Why /θ/, the English R and the schwa cause the trouble they do.
Straight answers
What is a minimal pair?
Two words that differ by exactly one sound — ship and sheep, right and light, thin and tin. Because everything else is identical, that one sound carries the whole meaning, which makes the pair a precise test of whether you can hear and produce the contrast.
Why practise pairs instead of single sounds?
Because a sound you can make in isolation is not yet a sound a listener can identify. What matters is whether it is distinct from its nearest neighbour, and only a pair tests that. It is also the fastest diagnostic: if you cannot hear the difference, drilling your mouth will not fix it, so the ear comes first.
Which minimal pairs should I practise?
The ones your first language does not separate. A Japanese or Korean speaker starts with right/light; a Hindi or Urdu speaker with wine/vine; a Spanish or Italian speaker with ship/sheep; a Mandarin speaker with think/sink. Each contrast page names the language groups that typically merge it.
Is it normal not to hear the difference at first?
Completely. Your ear was trained by your first language to ignore distinctions it does not use, and that filtering is efficient rather than broken. Recognition comes back with exposure, usually before production does — which is why every pair page starts with a listening test and only then asks you to speak.
Say both words and see which one you actually said
Record the pair in your browser. Reference mode lines every sound up against its target, so a single take comes back with both sides scored — free, no card.
Transcriptions follow the teaching set sayit uses throughout: British-style vowels with the American GOAT vowel, and non-rhotic endings for the centring vowels.