Guides · 4 min read

What is per-sound feedback, and how does sayit's IPA heatmap work?

sayit marks every word by how clearly you said it, then opens the exact phoneme that slipped in real IPA, side by side with the target sound.

Per-sound feedback means sayit scores your pronunciation one word at a time instead of handing you a single number for the whole sentence, and for any word that slipped, it opens the exact phoneme that caused it — shown in IPA, next to the sound you actually produced. That is the difference between being told "72%" and being told "your /θ/ came out as /s/."

30-second version: Read a line aloud and every word gets a clarity verdict — clear, unclear, or check. Tap a flagged word and it expands into IPA: the target sound and the sound sayit actually heard, side by side, with the one thing to change. Nothing is guessed from a transcript; it comes from a phoneme recognizer that hears the raw sound, not the word it expects.

What does a word-level verdict actually mean?

Every word you read gets one of three states. Clear means it landed cleanly. Unclear means sayit isn't confident enough to accuse you of anything specific — the signal is ambiguous, so it stays quiet rather than guessing. Check is the only "you mispronounced this" verdict, and it is kept deliberately narrow: it only fires when there is a real, serious error — a dropped consonant or a substitution that changes the sound, not a soft reduction — and the acoustic model is confident enough to say so. A near-merger like an R-colored vowel or a reduced final consonant scores as clear with a note, never as an accusation.

That caution is intentional. A tool that flags something wrong every time it is slightly unsure trains you to distrust it. sayit would rather stay silent than manufacture a false alarm.

Why IPA instead of a plain-language description?

Because "say the TH sound better" doesn't tell you what your mouth is currently doing wrong. Real IPA does. When a word is flagged, sayit shows the target phoneme and the phoneme it actually detected, in the alphabet linguists use — /θ/ next to /s/, /ɪ/ next to /iː/ — so you can see precisely which contrast collapsed. Every phoneme link opens its own articulation page: what the mouth does, example words, and a way to hear it. The TH sound and the short I versus long E contrast are two of the most common flags for learners of many first languages.

What's the "heatmap" part?

Inside a flagged word, each IPA atom is color-coded by whether it matched, was substituted, or was dropped — a small strip of green and red under the spelling that shows you which part of the word to fix, not just that something in it was wrong. On the word as a whole, the same three-color logic applies at the word level: green (clear), amber (unclear), red (check), so a whole paragraph reads at a glance as a map of where to focus.

Does it grade my accent?

No. The comparison is against the target sound, not against a "correct" voice — regional and first-language coloring inside a phoneme's own territory scores as clear. What gets flagged is a contrast that actually changed the sound, the kind that would cost you the word with a real listener. That's covered in more depth in how sayit handles accents.

How is this different from a transcript with red underlines?

A transcript-based tool runs speech-to-text and compares the words it guessed against the words you were supposed to say. The problem is that speech recognizers are built to guess the most likely sentence, so they quietly repair your pronunciation on the way to the page — say "sink" for "think" and a good transcription engine often still prints "think." Per-sound feedback works the other way: a phoneme recognizer with no language model attached listens for the raw sound, so it can catch the exact substitution a transcript would erase. The mechanics of that pipeline are in how the phoneme recognizer differs from speech-to-text.

A quick comparison

One sentence-level scorePer-sound feedback
What you seeA percentage for the whole lineA verdict on every word
What a mistake tells youSomething was offWhich phoneme, and what it should have been
Accent handlingOften penalizes any deviationIgnores regional variation, flags real contrasts
Next actionGuess and re-recordFix the specific sound, then re-record

Try it

Open sayit in the browser, read one sentence aloud, and tap any flagged word — you'll see your IPA next to the target IPA within seconds, no card and no install required. The full mechanism, including the acoustic model underneath, is on the features page.

Free in your browser

Hear exactly which sounds to fix.

Say one sentence and get sound-by-sound feedback in seconds. No install, no card.