← Blog · 28 September 2026 · Lire en français

More than one voice in an article: why only quotes switch voice

Telling two speakers apart by ear keeps a long article listenable to the end, especially for a blind or low-vision listener. How a quote is spotted, and when the tool holds back.

A long article read in one flat voice eventually blends into background noise. That is true for everyone, and more so for a blind or low-vision listener, whose attention has only the ear to hold on to. Telling two speakers apart, the reporter's voice and the voice of what they quote, is exactly what makes it possible to stay to the end. This came from a publisher, from articles that quote diaries, wire reports, statements. The answer fits in one sentence: a second voice for reported speech. What follows is how the tool spots it, and above all when it refuses to guess.

What does not work: markup

The obvious first idea was to read the article's structure, its quote tags, and switch voice on whatever is marked as quoted. It does not hold, for a mechanical reason. When the article arrives through the API or an RSS feed, it arrives as text. When it comes from a web page, it goes through a preparation step that strips out what should not be read aloud, tags included. By the time synthesis runs, the text no longer carries a single quote tag. There is nothing to read. Relying on tags meant building on what has already vanished.

What we read instead: the quotation marks

Quotation marks, on the other hand, are in the reporter's own text. They survive both ways of preparing an article, and above all they ask nothing of anyone: no tag to add, no template to change, no extra field in a CMS. The author writes their quote inside quotation marks, as they always have, and the tool leans on what is already there. Our page on French text to speech comes back to this stance: work the real language rather than require the customer to tag their text for the machine.

One detail matters: only quotation marks whose opening and closing signs differ are used (French guillemets, English curly quotes). The straight quote, which opens and closes with the same character, is left out, because nothing says where the quote ends. A leftover from typing would be enough to tip half an article into the wrong voice.

The risk is not symmetrical, so the tool leans the safe way

This is the heart of it. Missing a quote costs almost nothing: it is read in the article's voice, as before, and nobody notices. Inventing one, by contrast, is expensive: a passage read in the wrong voice is heard at once and makes the product sound broken. Any doubtful decision therefore leans toward "this is not a quote". Three guards hold that line.

First, a quote has to begin a sentence. A term in quotation marks mid-sentence ("blank cheque") is an aside: switching voice for three words and switching back leaves two crumbs on either side and sounds like a fault. Second, a length floor: below about eighty characters, it is an expression, not a spoken turn. Third, the quote has to be closed within its paragraph. A quote mark never closed, or closed three paragraphs later, does not flip the rest of the article: that is the most expensive fault possible here, and a single typo is enough to trigger it.

A case where the tool holds back

Take an article that writes: the mayor spoke of a "blank cheque" before wrapping up. The term is in quotation marks, but it is short, mid-sentence, and opens nothing. The tool does not switch voice: it reads it all in the reporter's voice. That is the right call. Switching on two words, then switching back, would produce an audible hiccup far more noticeable than the nuance you thought you were gaining. Here, restraint is a quality, not a limitation.

An invisible but decisive property rounds this out: the voice split loses not a single character. Concatenating its pieces gives back exactly the original text. That is essential, because word-by-word highlighting and subtitles are aligned on that text: a character that vanished here would shift everything else in the article.

Beyond two voices

The same mechanics extend to written dialogue, the kind in a play or a field diary, where several people speak. There, the quotation marks do not say who is speaking: the shape of the dialogue does, a line that starts with a name followed by a colon. The account declares its characters and their voices, and an attributed line is read in the character's voice to the end of the line. The name itself stays in the body voice: it is the attribution, like "the director summed up". A name that has not been declared changes nothing, and case and accents do not count. Here too, the same guards, the same way: better to switch nothing than to switch wrongly.

Give your articles a voice with WeDispatch

This blog is itself voiced by WeDispatch. Curious how it sounds on your content?

Book a demo

Read next