Why do 'cut' and 'cat' get confused with each other?
Cut and cat share the same consonants and differ only in the vowel — short u vs short a — two sounds many languages don't split into separate phonemes.
Because the two words are built on the exact same consonant frame — /k_t/ — and differ in nothing but the vowel in the middle, and that particular vowel, /ʌ/ (the STRUT vowel in "cut"), doesn't exist as an independent sound in a large number of languages.
30-second version: "Cut" is /kʌt/ and "cat" is /kæt/. Both vowels are short, both are made with the tongue roughly central-to-front, and the difference — tongue height and a touch of openness — is small enough that many language backgrounds simply don't carve out two separate phonemes there. Spanish, for instance, has one open central-ish vowel that does the job of both English targets, so the words genuinely can sound like the same word with a different consonant at the end, not two clearly different words.
Why this pair, specifically
English vowels are packed unusually tightly compared to most languages' vowel inventories — five or so vowel letters cover a dozen-plus distinct sounds. /ʌ/ and /æ/ sit close together in that crowded space: /æ/ is a fully open front vowel (jaw drops, mouth wide), /ʌ/ is a shorter, more central, less open vowel that many learners initially produce as a version of /æ/ or of the neutral schwa, because their ear hasn't learned to carve out a third category in that region yet.
This is a genuinely different kind of problem from a consonant substitution — there's no "wrong tongue placement" video that fixes it in one watch, because the issue is categorical: does your phonological system have a box for this sound at all yet? Building that box takes repeated, deliberate contrast practice more than it takes a single articulation correction.
How sayit separates the two
Because the consonants on either side of the vowel are identical in this pair, the entire signal the model has to work with is the vowel formants in that one syllable — which is exactly what an acoustic model is built to measure precisely, even in cases where the difference is genuinely subtle to an untrained ear. If you say "cat" and the vowel's acoustic profile sits closer to /ʌ/ than /æ/, the breakdown records a substitution at that one vowel slot, with the "sounded like" hint naming "cut" directly if that's what the produced sound spelled.
What the fix looks like
Say both vowels back to back with an exaggerated jaw drop on /æ/ — it should feel almost too open, like the start of a yawn. Pull the jaw back in for /ʌ/ — shorter, more relaxed, tongue more central. The two should feel clearly different in your mouth even before they sound clearly different to your ear; building that physical distinction first often makes the auditory one easier to notice.
A trick that helps some learners: /æ/ is close to the vowel in many languages' "ah" but with the mouth spread slightly wider and more forward; /ʌ/ is closer to a short, unstressed grunt — think of the sound you'd make bumping your shin, cut short.
Drill it
Cat vs cut is the direct pair. Related contrasts using the same two vowels: bad vs bud, hat vs hut, and ran vs run — working through several pairs matters more here than for most contrasts, because the goal is building a new phonemic category, not fixing one word.
Record: "The cat cut its paw on the fence." Both target vowels sit two words apart for an easy playback comparison.
The /phonemes/short-a and /phonemes/short-u pages have the individual articulation steps.
How do I know it's improving?
Random-order listening discrimination is the real test here, more than production accuracy — if a native or TTS voice says one of the pair at random, can you identify which one without seeing it written? Many learners can produce the contrast reasonably well in a rehearsed drill before they can reliably hear it in someone else's speech; both skills need practice, and the listening side often lags.
Try it
Open sayit free — no install, no card — and record the drill sentence above. If both vowels come back scored the same, that's the signal this contrast needs dedicated pair practice, not a general pronunciation review.
Hear exactly which sounds to fix.
Say one sentence and get sound-by-sound feedback in seconds. No install, no card.