Why does sayit show the target sound next to what you actually said?
A score tells you something was wrong. Target-vs-produced IPA, side by side, tells you what to change — which is why sayit always shows both.
sayit shows the target sound next to what you actually said because a number alone doesn't tell you what to do differently, and a word alone doesn't tell you which part of it went wrong. Put the two sounds side by side in IPA and the fix becomes visible: not "try harder," but "your tongue needs to go between your teeth for that one sound, right there."
30-second version: For any flagged word, sayit shows two IPA strings — the target pronunciation and the one it detected from your recording — lined up phoneme by phoneme, with the mismatch highlighted. A plain-English respelling (capitals on the stressed syllable) sits alongside for anyone who doesn't read IPA, plus a short instruction for what to physically do differently.
What does the comparison actually look like?
Take the word "pleasant." The target IPA is /ˈplɛznt/. If you tense the vowel, sayit might detect /ˈpleɪznt/ — an /ɛ/ that slipped toward /eɪ/. Both strings render on screen, atom by atom, with the atom that differs marked. You don't have to guess which of the six-odd sounds in the word was the problem; it's pointed at directly. The per-sound feedback article covers how that comparison gets color-coded across a whole sentence.
Why not just say "wrong" or show a percentage?
Because neither tells you what changed. A percentage answers "how far off," which is useful for tracking progress over time but useless in the moment you're trying to fix something. "Wrong" doesn't even answer that much. Target-versus-produced answers the only question that actually changes your next attempt: which specific sound moved, and in which direction. That's the same logic behind the one-fix recommendation sayit gives at the end of a take — a diagnosis, not a grade.
I don't read IPA. Is there another way to see it?
Yes. Every flagged word also gets a plain respelling with the stressed syllable in capitals — "pleasant" reads as PLEZ-uhnt — so you can act on the feedback immediately without learning the phonetic alphabet first, while the IPA sits right there if you want the precision later. The browser extension uses the same respelling on any page you're reading, for exactly this reason.
Does it also show me what to do with my mouth?
For the sounds where it helps, yes. Alongside the IPA, sayit pairs a short articulation cue with the specific phoneme that slipped — tongue position, lip shape, voicing — rather than a generic "practice more." A dropped final /ŋ/ gets "close the back of your tongue and hum"; a tensed /ɛ/ gets "relax the jaw and keep the vowel short." These are written per phoneme, so the tip always matches the actual sound that needs to change, not the word in general.
How target-vs-produced compares to a raw score
| A single accuracy number | Target vs produced IPA | |
|---|---|---|
| Tells you something is wrong | Yes | Yes |
| Tells you which sound | No | Yes, atom by atom |
| Tells you what to change physically | No | Yes, with an articulation cue |
| Useful for tracking trends | Yes | Yes, alongside the score |
Does this work on whole sentences, or just single words?
Every word in a take gets its own target-versus-produced comparison when it's flagged; clear words don't need one, since nothing there needs fixing. Across a full passage, sayit also rolls repeated substitutions into a focus-areas summary, so if the same target-to-produced pattern shows up five times in one read — say /v/ consistently coming out as /w/ — that pattern surfaces once as the thing worth drilling, rather than getting buried in five separate word flags.
Try it
Read any sentence in sayit and look at what happens the moment a word gets flagged — the target IPA and your IPA appear together, not one after the other. If you'd rather see it on a specific word first, look it up at /pronounce or check a tricky consonant like /r/ or /v/ on the sound chart.
Hear exactly which sounds to fix.
Say one sentence and get sound-by-sound feedback in seconds. No install, no card.