Technology · 3 min read

What is the pitch contour sayit draws from my recording?

sayit tracks your fundamental frequency across a sentence and plots it as a line, so you can see — not just be told — whether your voice rose, fell, or stayed flat.

The pitch contour is a line graph of your fundamental frequency — how high or low your voice was, moment to moment — drawn across the length of a sentence you read. Where the intonation score gives you a single number for how well your pitch pattern matched what a sentence like that one should sound like, the contour is the raw evidence underneath it: the actual shape of your pitch, so you can see where it moved and where it didn't.

30-second version: sayit extracts your fundamental frequency (F0) throughout a sentence from the raw waveform and plots it as a curve. That curve reveals the shape a single score can't fully describe: whether your pitch fell at a statement's end, rose on a genuine question, or barely moved at all across the whole phrase. A pitch-range reading flags delivery that's unusually flat, separate from whether individual sounds or stress placement were correct.

What does the contour actually show?

A continuous trace of pitch over time, synced to the words you spoke. On a well-delivered statement, the line typically drifts downward toward the end. On a genuine yes/no question, it typically lifts. On a list, it often rises slightly on each item before a final fall. The shape is the point — a monotone reader and a naturally intonated one can both score well on individual phonemes, but their contours look completely different, and that difference is exactly what carries meaning like emphasis, contrast, and sentence type in English.

How is pitch actually extracted from the audio?

sayit estimates fundamental frequency directly from the recorded waveform at a fine time resolution, tracking it continuously through voiced portions of your speech. The result is a sequence of pitch values across the sentence, which is what gets classified into a rising, falling, or flat pattern and rendered as the visible curve.

What is "monotone detection," and why does it matter separately from intonation?

Monotone delivery is measured by how much your pitch varies across a take — the spread between your lowest and highest points, and how consistently it stays near one value rather than moving. A speaker can technically produce a rising contour on a question and still sound flat overall if the rise is barely perceptible; a low pitch-range reading catches that even when the directional pattern is technically correct. This matters because flat delivery is frequently a confidence artifact rather than a skill gap — many second-language speakers flatten their pitch specifically in higher-stakes situations, and it tends to loosen once the speaker relaxes, which is a useful thing to notice about your own recordings over time.

Does the contour compare my voice to a "correct" one?

Not to a single fixed voice — there's no one right pitch to imitate, since natural pitch range varies enormously between speakers. What sayit compares is shape: the expected pattern for the sentence type you produced (statement, question, list) against the pattern your recording actually shows, not the absolute pitch height. A naturally low or high voice isn't penalized; a contour that doesn't move the way its sentence type calls for is what gets flagged.

Reading the contour at a glance

Sentence typeExpected shapeWhat flat delivery loses
StatementFalls toward the endSounds unfinished or hesitant
Yes/no questionRises toward the endCan be heard as a statement, not a question
ListSmall rises per item, fall on the lastItems blur together

Try it

Read a question aloud in sayit — something ending in a question mark — and look at the pitch curve in your results. If the line stays nearly flat where it should lift, that's the same gap the intonation score is measuring, just drawn out so you can see exactly where it happens.

Free in your browser

Hear exactly which sounds to fix.

Say one sentence and get sound-by-sound feedback in seconds. No install, no card.