Choosing a French text to speech tool: six questions to ask first
Not a brand comparison, a checklist. Six questions to put to any French speech synthesis tool, and our own honest answer to each, including where it is no.
Most people pick a speech synthesis tool by ear, on a thirty-second demo. That is the worst moment to decide, because every voice sounds good on text chosen to flatter it. The questions that matter are not about the timbre. They are about what the tool does with your real articles, your real proper nouns and your real figures, once the demo is over.
Here are six questions to ask any provider. This is not a ranking of products: it is a checklist you can apply yourself, to us or to a competitor. To stay honest, we answer each one for WeDispatch, including when the answer is no.
1. Does the tool read French, or merely pronounce it?
This is the core question, and it splits tools into two families. Pronouncing means turning letters into sounds. Reading means deciding first how to say "1,200 €", "3:30 p.m.", "SNCF" or "M. Charrier", and only then pronouncing it. The difficulty of machine-read French does not come from the voice, it comes from that preparation step.
Our answer: this is our whole bias, and the reason the French text to speech page exists. The engine prepares numbers, acronyms, times and punctuation before it speaks. Test it yourself by feeding the tool a paragraph that stacks several of these in a row: that is where most of them stumble.
2. How many languages, and to what standard?
A catalogue of thirty languages is reassuring on a spec sheet and tells you nothing about how well any single one is read. Reading a language well takes preparation work specific to its grammar, and that work does not replicate for free from one language to the next.
Our answer, and it is a deliberate no: we read five languages, French, English, Spanish, German and Italian. Not Portuguese, not Arabic. We would rather read a few languages well than read thirty badly, because a French-speaking publisher judges a tool on French, not on a count of flags. If you publish in twelve languages, we are not the right tool, and it is better to know that now.
3. Can I fix a pronunciation, and does the fix stick?
Every voice gets a town name, a surname or a brand wrong. The real question is not "does it make mistakes?" but "can I teach it, once, and does that apply to every article after?". A tool that makes you redo the correction on every text will cost you more time than it saves.
Our answer: yes, through a pronunciation lexicon. You declare a word's expected pronunciation once, and it applies everywhere after. This is worth testing explicitly with any provider, because the answer varies a great deal.
4. What does the tool return besides the sound file?
A tool that only hands back an MP3 quietly closes doors you cannot yet see: text that highlights as the voice reads, search inside the audio, subtitles, fixing one passage without redoing the whole file. All of them rest on a single piece of data, the exact position of each word in the file.
Our answer, with the trade-off attached: we kept a single reading engine, and we chose it because it timestamps at the word. Some better-sounding voices do not, and we unplugged them rather than promise a highlight they cannot deliver. That is a choice, not a free advantage: we preferred the feature to half a tone more in the voice.
5. Where does my data go, exactly?
Text sent to a voice tool passes through at least two stages: preparing the reading, which sees the whole article, and synthesising the voice. Ask where each one happens, not "are you compliant" in general.
Our answer, with its share of no: text preparation happens in the European Union. Voice synthesis does not, yet. We do not pretend otherwise, and the so-called sovereign tier stays greyed out until the voice chain is European end to end. A provider that answers "it is all in Europe" without separating the two stages is worth pushing on.
6. What does it cost per article, regenerations included?
Prices are shown per million characters, which helps no one decide. Nobody listens to a million characters: they listen to an article. Multiply by the average length of yours, and add the regenerations, because an article corrected after the fact is an article paid for twice.
Our answer: we think in articles and monthly volume, not per character, and a free trial lets you hear the result on your own texts before paying anything. The best way to answer this question is still to voice three of your articles and look at the result.
These six questions do not crown a winner. They give you a grid to judge with, and a way to spot a provider who dodges: one who answers "yes" to all six without ever naming a limit has not answered you, they have sold to you. To go further on judging a specific voice, our piece on how to evaluate a neural voice sets out the tests to put it through.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo