Guided reading: text highlighted word by word, who it helps and what it takes
The word being read lighting up on screen is not an effect, it is real help for anyone who reads with difficulty. But it cannot be improvised: it rests on one precise piece of data that not every engine returns. What it takes for highlighting to exist, said plainly.
There is one feature readers ask for by name: the text lighting up in time with the voice, word after word, while they listen. Call it guided reading, or synced highlighting. It looks like a display detail. It actually rests on one precise piece of data, and understanding which one explains both who it helps and why not every tool offers it.
What guided reading actually does
The principle is simple to describe: the article stays on screen, and the word the voice is speaking right now lights up. As the voice moves forward, the highlight moves with it. The reader follows with their eyes what they hear, without hunting for where they are.
The nuance that matters is the grain. Sentence-level highlighting lights a whole block for four to six seconds: that does not feel like following, it feels like a lamp left on. Word-level highlighting lights one word at a time, and that is the whole distance between an interface that follows the reading and one that approximates it. Guided reading worth the name is at the word, not the sentence.
Who it really helps
You have to be honest about the audience, because that is where the feature earns its meaning or loses it.
It helps, first, people who read with difficulty. For a dyslexic reader, following a line of text without losing the thread is a constant effort; the lit word gives an anchor that moves on its own, and listening supports reading instead of replacing it. The same logic holds for a child learning to read, an adult learning French, or anyone who tires quickly on a screen. The voice alone gives them the content; the voice plus the highlighted word makes the reading itself more accessible.
It helps less, it must be said, the rushed listener at the wheel or on foot: that person is not looking at the screen, and the highlight gives them nothing. Guided reading is not a feature for everyone, it is a decisive feature for part of the audience, and that is exactly what makes it real help rather than a gimmick. That demand to read French well, right down to what helps follow it on screen, is the line our French text to speech page defends.
What it takes for highlighting to exist
Here is the technical point, said plainly, because it is what separates the tools. To light the right word at the right instant, you have to know, for every word in the article, the precise moment the voice starts it and the moment it ends it in the audio file. Not for the paragraph, not for the sentence: for the word. This is called word-level timing, and a three-minute article produces a table of six to seven hundred entries, each saying "this word runs from this second to that second."
Without that table, highlighting is impossible, or it cheats: it estimates a word's position from its length, and it slips the moment there is a slightly long pause or a number read aloud in full. With the table, it lands right, because it guesses nothing.
Not every speech engine returns this timing. The reliable way to get it is for the engine to produce it at the same time it makes the sound: it knows what it spoke, so it knows exactly where. That is the precise path, and it costs nothing extra at generation. Reconstructing the timing after the fact, with a second tool that listens back to the audio, is possible but approximate, and that is where the slips come from. We laid out this mechanism, and the features that quietly depend on it, in word-level timestamps and what they are.
What it costs to produce, and for whom
Good news on cost: when the timing is returned by the engine at synthesis, it adds nothing to the bill. You do not pay for guided reading on top, it is a by-product of generating the audio. Highlighting is not a paid option to switch on, it is what the data already present lets you display.
The real cost is elsewhere, and it is historical: audio produced before word-level timing arrived does not carry this table. With us, guided reading works on articles voiced since mid-July 2026; for older audio, a simple regeneration makes it eligible. That is the only real spend, and it is one-off.
What highlighting drags along with it
One last point, because it changes how you judge the feature. The timing table that makes highlighting possible also serves, exactly, three other things: searching inside speech (jumping to the moment a word is spoken), subtitles that land right on a breath rather than an estimate, and correcting three mispronounced words without redoing the whole article. Guided reading is therefore not an isolated feature you bolt on: it is the visible part of a piece of data that makes audio steerable to the word. Asking for it is asking for much more than a display effect. For readers who want the reading experience itself, we also offer read-along audio with synced transcripts.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo