Synthetic voices and emotion: what they convey, and what they don't act out
A neural voice conveys a question and an exclamation convincingly. Irony, reported anger, grief, far less. Why an article is not performed like a play, and exactly where the line falls, said without dressing it up.
There is a fair expectation behind the question "does the voice put feeling into it?". It comes from the fact that synthetic voices have improved so much that they are often taken for human ones. But getting more natural is not the same as learning to act. A neural voice reads accurately; it does not interpret. The difference is easy to miss on neutral text, and it jumps out the moment a passage calls for emotion. Here is where it falls, with three sentences that show it.
What the voice conveys well
Start with what works, because it is real. Grammatical intonation is well rendered. A question rises at the end: "Do you really think this budget will pass?" is read with the expected upward curve, with nothing to adjust. The handling of interrogatives is reliable enough to lean on in a news article, as we detail in our piece on question intonation in French. Exclamation works too: "What a day!" gets a believable emphasis. The voice reads punctuation as a cue for the melodic line, and on that ground it is at ease.
This base is why you can no longer spot a synthetic voice in the first word, a subject we cover in why you can't tell anymore. But recognising a good reading and recognising a good performance are two different judgements.
The three sentences where the line shows
Take irony first. "Well done, that is exactly what you should have done." In writing, context tells you whether it is praise or reproach. Spoken, an actor would flip the meaning with a dry inflection on "well done". The synthetic voice reads the sentence at face value: the punctuation is neutral, so the curve is neutral, and the irony is lost. This is not a flaw in the voice; irony lives in the gap between what is said and what is meant, and nothing in the text signals that gap.
Reported anger next. "He shouted: Get out of here." A human would drop a tone to set the narration, then rise sharply on the quotation. The synthetic voice reads both parts in the same register: you understand there was a shout because the text says so, not because you hear it. Switching voice on quotations helps tell two speakers apart, but it does not put anger into the quote; it sets it off, it does not perform it.
Grief last. "She left on a Tuesday, without telling anyone." Spoken by a human voice, that sentence slows, holds back, leaves a silence. The synthetic voice reads it at its usual pace, with the pause the comma calls for, no more. The weight is not there, because weight cannot be summoned by punctuation: it is acted, and acting is not reading.
Why we made this choice, and why it is a choice
One could want to "add emotion", and some tools offer it through intensity settings. We did not, and it is not only a matter of technical limits: it is a stance. A news article read with added emotion rings false, because pasted-on emotion is worse than no emotion. A news item delivered in a dramatic tone becomes a bad TV movie; an economic analysis coloured with enthusiasm becomes suspect. The neutrality of a good reading is a virtue for information, not a shortfall.
Where emotion really matters, for a story, a work of fiction, a poem, the right tool is not a synthetic voice asked to perform: it is a human voice. We say so plainly, because claiming otherwise would mean overselling what we do. Our job is to read French well, not to stage it.
What to take away before you choose
If you are evaluating a voice tool, test it on an ironic sentence and on a solemn one, not just on a neutral paragraph: that is where the real differences appear, and it is what our method for evaluating a neural voice recommends. And ask yourself the right question: does your content need to be performed, or to be read clearly? For news, documentary, professional content, an accurate, plain reading serves better than any simulated emotion. For a novel read aloud, no synthesis yet replaces an actor, and it is honest to say so.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo