Why does whispering or a noisy room mess up my pronunciation score?
The scoring engine reads acoustic evidence from your waveform. A whisper removes the vocal-cord data it needs, and background noise degrades the same signal.
Because both conditions genuinely degrade the same thing the scoring depends on — clean acoustic evidence of which sounds you made — and a whisper isn't simply a quiet version of normal speech to an acoustic model, it's missing an entire category of information that normal speech provides.
30-second version: Voiced sounds — most vowels, and consonants like /b, d, g, v, z/ — get their character partly from vocal cord vibration. Whispering removes vocal cord vibration entirely, so the model loses one of its clearest signals for telling voiced sounds apart from their voiceless counterparts (b vs p, v vs f, the whole /ð/ vs /θ/ contrast this series has covered repeatedly). Background noise is a different problem with a similar effect: it doesn't remove information, it buries it, lowering the model's confidence across the board rather than on any one specific sound.
Why a whisper specifically breaks voicing contrasts
Voicing — whether the vocal cords vibrate during a sound — is one of the primary things that separates paired consonants in English: /p/ vs /b/, /t/ vs /d/, /k/ vs /g/, /f/ vs /v/, /s/ vs /z/, /θ/ vs /ð/. In normal speech, that vibration is a strong, direct acoustic signal. In a whisper, there's no vibration at all — every one of those pairs loses its clearest distinguishing feature, and the model has to rely on secondary cues (timing, airflow noise) that are real but weaker. This is exactly why a whispered recording tends to produce more uncertain, low-confidence results specifically on voicing-based pairs, even when the rest of the pronunciation is genuinely fine.
Why background noise is a different kind of problem
Noise doesn't target a specific contrast the way whispering does — it degrades the whole signal roughly evenly, which is why a noisy recording tends to produce more unclear verdicts across many words rather than confident wrong flags on specific sounds. sayit's scoring is deliberately built to respond to this correctly: when the acoustic confidence for a word comes in low, a would-be strong "you got this wrong" verdict is softened to a more honest "not confident either way" rather than forcing a specific accusation the evidence doesn't actually support. That's the right behavior, but it also means a noisy recording gives you less useful feedback overall, not wrong feedback — you're not being penalized, you're just not getting a clear read.
What actually works around this
Speak at a normal conversational volume, not a whisper. If you're practicing somewhere you can't speak at full volume, a normal quiet speaking voice with full vocal cord engagement gives the model far more to work with than a whisper does, even at low volume — voicing information doesn't require loudness, it requires the vocal cords actually vibrating.
Reduce room noise before reducing your voice. Closing a door, turning off a fan, or moving away from a window does more for scoring quality than speaking more carefully does, because the issue is signal-to-noise, not effort.
Get closer to the microphone rather than raising your voice. This improves the ratio of your voice to the room noise without changing how you're actually pronouncing anything, which is the cleanest fix when environment can't be fully controlled. Microphone-specific tips go further into hardware and positioning.
How do I know it's actually my environment and not my pronunciation?
Re-record the identical sentence in a quieter space at normal volume and compare the two verdicts directly, the same check described in what to do when the app is wrong about a word. If the flags clear up substantially with better conditions, the original take was a signal problem, not a pronunciation one.
Try it
Record: "The vet went west in the wet weather." It's built on voicing contrasts, so it's a good stress test — open sayit free, no install, no card, and try it once whispered and once at normal volume in a quiet room, then compare.
Hear exactly which sounds to fix.
Say one sentence and get sound-by-sound feedback in seconds. No install, no card.