Technology · 3 min read

What's the difference between read-aloud and freeform mode in sayit?

Read-aloud scores you against a known passage, word for word. Freeform scores whatever you actually say, with no script to compare against.

Read-aloud and freeform are the two modes sayit scores speech in, and they exist because they answer two different questions. Read-aloud (also called reference mode) asks "how closely did you match this exact text?" — you're given a passage, and every word you're expected to say is known in advance. Freeform asks "what did you say, and how well did you say it?" — there's no script, so the words themselves have to be figured out from your recording before anything can be scored.

30-second version: In read-aloud mode, sayit already knows the target text, so it can measure completeness (did you finish it) and often skips speech-to-text entirely, deriving word timing straight from the phoneme stream. In freeform mode — answering a question in your own words, the way IELTS or TOEFL speaking sections work — sayit transcribes your speech first with ASR, then scores the phonemes in whatever words you actually produced. Both modes score pronunciation the same way underneath; what differs is where the words come from.

How does sayit know which mode to use?

It's determined by whether you're given text to read. If a passage is loaded — from the reading library, an imported book, or a generated paragraph — you're in reference mode. If you're answering an open question with no script on screen, the way IELTS Part 2 and 3 or a conversation practice turn work, you're in freeform.

Why does read-aloud skip transcription?

Because it's redundant. sayit already knows exactly what you're supposed to say, so instead of running full speech-to-text and then matching the result back to the known text, it can derive per-word timing directly from where the phoneme recognizer placed each sound — the pronunciation scores come out byte-identical either way, for a real cut in processing time per take. Freeform has no such shortcut: since the words aren't known in advance, ASR has to run to find out what you said before phoneme scoring can even start.

What can read-aloud measure that freeform can't?

Completeness — the fraction of the passage you actually attempted. Reference mode has a denominator (the full passage) and a numerator (the words you got to), so it can tell you honestly that you read 60% of a paragraph rather than silently scoring you only on the part you finished. Freeform has no fixed target, so there's nothing to be "complete" against — read reference scoring and completeness for how that number is calculated and why a short take doesn't get penalized twice.

What can freeform catch that read-aloud can't?

Whether you can actually produce language on your own, under the pressure of coming up with the words in real time — which is the entire point of an IELTS or TOEFL speaking test, and the reason exam sections score freeform answers rather than read-aloud ones. Freeform is also where filler words ("um," "like," "you know") get detected and counted, since a read-aloud passage doesn't naturally contain them the way spontaneous speech does.

Side by side

Read-aloud (reference)Freeform
The words to scoreKnown in advanceTranscribed from your speech
Completeness scoreYesNot applicable
Typical usePassages, drills, booksIELTS/TOEFL answers, conversation
Off-script speechReported separately, not charged to the passageN/A — it's all in scope
ASR (speech-to-text)Usually skippedAlways runs

What if I go off-script while reading a passage?

sayit doesn't silently charge stray words to whatever passage word you were closest to. If your speech drifts off the practiced text — a stray comment, a restart, a correction — that extra speech is reported on its own, separately from the passage's own scoring, rather than getting mixed into your accuracy for words you never actually tried to read.

Try it

Read a passage aloud in sayit to see reference mode, then answer an open IELTS speaking prompt to see freeform. The scoring engine underneath both is the same — only where the target words come from changes.

Free in your browser

Hear exactly which sounds to fix.

Say one sentence and get sound-by-sound feedback in seconds. No install, no card.