Test a French text to speech voice in ten minutes
A short, reproducible protocol to choose between two French synthetic voices, usable against any tool including ours. Six trials, ten minutes, one decision.
Choosing a synthetic voice on a hunch, from a well-picked demo, means deciding in thirty seconds what you will hear for months. The bill comes later, in production, when the voice that sounded lovely on three sentences mangles a town name or reads "quinze slash zéro trois" instead of a date.
Here is a ten-minute protocol, timed, to apply as is to any tool, ours included. It needs no technical skill and it is deliberately reproducible: you can run it against three competitors in the same half hour and compare on identical ground. Prepare one thing first: a real article from your site, not a demo text. Yours contains your proper nouns, your figures, your sentence length. It is the judge.
Minutes 1 to 2: your worst paragraph
Find in your article the passage you would struggle to read aloud yourself: a long sentence with two asides, or a paragraph stringing together three organisation names. Give it to the voice. A voice that holds on difficult text will hold anywhere; the reverse is never true. If it collapses here, you can stop the test.
Minute 3: numbers and dates
Find an amount, a date, a percentage, a time, a reference number. Listen only to those spots. In French, "1 200 €" must be said "mille deux cents euros", "15/03" must be said "quinze mars". A voice that reads "quinze slash zéro trois" has failed at something basic, and your articles are full of it. This thirty-second test rules out more candidates than any other. The detail of the forms that trip a voice up is in reading numbers aloud in French.
Minute 4: the proper nouns of your area
Have it read five names that recur for you: towns, districts, officials, local companies, trade acronyms. Count the errors. That number predicts your satisfaction at six months better than any impression of the timbre. And ask the real question straight away: can these errors be fixed without changing voice, through a pronunciation list that holds over time? With us, that is the job of the pronunciation lexicon.
Minute 5: liaison
Have it read "Les haricots sont bons" and "Ils en ont assez". The first forbids the liaison ("lay / aricot"), the second links everything up. A voice that says "lay-z-aricot" does not know the list of aspirated-h words; a voice that detaches "ils / en / ont" sounds robotic. The full test, with five trap phrases, is in liaison, the hardest test for machine-read French.
Minutes 6 to 8: ten minutes of listening, not thirty seconds
Listening fatigue cannot be measured on a clip. Start a whole article and listen to it while doing something else, as your audience will. Watch for the rhythm flaw: a cadence too regular, a breath in the same spot every sentence, unnoticed at first and grating after five minutes. If you want it to stop, so does your reader. This is the one point of the protocol that runs past the ten minutes on the label, and it earns it: start it first and let it play through the other trials.
Minute 9: the same text, twice
Regenerate the same passage and compare. A voice that produces two noticeably different versions will be a problem the day you fix one sentence in an already-published article: the remade segment will not join up with the rest. It is invisible in a demo and decisive in production.
Minute 10: what the voice returns besides sound
Ask whether the voice provides word-level timestamps, that is, the position of each word in the file. That is what enables highlighted text during playback, search inside the audio, and subtitles. The answer varies from one voice to another with the same vendor, which often surprises, and a voice that returns only a sound file closes doors you cannot see yet.
What the protocol tells you, and what it does not
In ten minutes you know whether a voice truly reads French or merely pronounces it, whether its errors are fixable, and whether it lets you build more than a plain sound file. What it does not tell you is the real cost at your volume and the generation time on your lengths: two points to measure next, developed in how to evaluate a neural voice.
This protocol favours no one: it is a grid, not a pitch. We publish it because our stance, set out on the French text to speech page, is to read one language well rather than thirty roughly. An honest grid serves that position: apply it to our voices as to the others, and judge on your own text.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo