Dialogue and quotations: what French text to speech does with the marks
A quote, a line of dialogue, a word set apart in quotation marks: none of them read like ordinary text. Here is what a voice engine actually does with them, where it stumbles, and how to fix it.
The moment a text reports someone's words, it shifts gear. A quotation, a line of dialogue, a word placed in quotation marks to hold it at arm's length: none of these read aloud the way a plain sentence does. The ear expects an inflection, a pause, sometimes a change of pace. This is exactly where a text to speech engine gets judged, because nothing in the written text tells it in so many words, "someone else is talking now." This article looks at what really happens, on cases you can reproduce yourself. It focuses on French, where the marks themselves differ from English.
The test paragraph
Take a passage that stacks the difficulties:
> The minister was blunt: « Nous ne reviendrons pas sur ce calendrier. » The mayor, for his part, spoke of a text « inapplicable en l'état », before adding, more quietly, that he would do « tout ce qui est en son pouvoir ».
French uses guillemets, the angle marks « and », where English uses straight or curly quotes. Three uses appear here. The first wraps a whole quoted sentence that stands on its own. The second isolates two words in the middle of a sentence, to flag that they are the mayor's and not the reporter's. The third does the same, at the end. A human reader handles the three differently without thinking. An engine has to decide.
What the engine does with the marks
First useful fact: the guillemets are not spoken. A good engine never says "open quote." It treats them as a punctuation signal, not as a character to read out. That sounds obvious, yet some tools, especially those tuned first for English, trip on the angle marks because they do not recognise them as quotation marks. They then either ignore them in the wrong place or insert a pause where none belongs. The first test of an engine on French is to hand it a quotation inside guillemets and listen for whether it swallows them cleanly.
Second, and subtler: the pause. After "The minister was blunt:", the colon calls for a short silence, and the opening of the quotation reinforces it. The engine should breathe there. Without that breath, the quote glues itself to the lead-in and the listener loses the boundary between reporter and minister. Punctuation carries this information, which is why a well punctuated text listens better than a careless one, a point we set out in our article on how text to speech breathes with punctuation.
Where nearly every current engine stops, though, is intonation. An actor would drop the voice a touch on a short embedded quote to show it is borrowed. Synthetic speech reads that short quote on the same melodic line as the rest. This is not a pronunciation flaw, it is a limit of meaning: the engine does not know those two words belong to someone else. The result stays perfectly clear, but the sense of quotation is lost. Better to know this before promising a newsroom an actor's delivery.
Dialogue, and the dash that opens a line
Fiction and interviews add one more hurdle: the dialogue line opened by a dash at the start of the line. In French, a dash there announces a change of speaker rather than a word to read. An engine that ignores it runs the exchanges together into a single speech; an engine that treats it as strong punctuation marks a clear pause at each turn, which is almost always right for the ear. The classic problem is not the dash itself but the tag that follows it: in "Je refuse, dit-il en se levant," the "dit-il" should come out lower and quicker than the line, or it carries as much weight as the refusal. No consumer engine makes that shift automatically today. The remedy sits in the writing: a short tag, cleanly set off by commas, reads acceptably even without modulation.
Fixing without rewriting the whole text
When a quotation contains a term the engine mishandles (a proper name, an initialism, a foreign word), the move is not to touch the visible text, which must stay exact, but to go through the pronunciation lexicon. You declare the intended reading once, and it applies everywhere, including inside the quotation marks. That is the right granularity: the text stays what the reader sees, the correction lives beside it.
For the rest, the best preparation is editorial. Introduce your quotations with a verb and a colon rather than dropping them into the middle of a sentence, keep dialogue tags short, and reread aloud any passage that reports speech. A text written to be heard travels through synthesis better than one written for the eye alone.
What to take away
Text to speech handles the mechanics of quotation well: it ignores the angle marks, respects the pauses that punctuation announces, and separates the turns of a dialogue cleanly. What it does not yet do is perform the quotation, lower the voice to signal that another person is speaking. Knowing where that line falls avoids two mirror mistakes: thinking an engine replaces an actor, and thinking it stumbles on difficulties it in fact handles well. If you want to check all this yourself, our checklist for evaluating a neural voice includes a passage with quotations, and you can run it against any tool, ours included. That same standard drives our work on reading French aloud, described on the French text to speech page.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo