How text-to-speech actually works, in plain terms
Between the written text and the sound that comes out, more happens than you would think. An article's journey, step by step, and the exact place where the quality of a French voice is decided.
People often picture text-to-speech as a machine reading letters out loud, one after another, mechanically. That was true of the early robotic voices. It no longer is. A recent neural voice does not read letters; it produces a sound signal that resembles a human voice, from a text that has first been prepared. Understanding that journey helps explain why two tools, fed the same article, do not produce the same result, and where the real work lives.
The text is not read as it is written
First surprise: before it ever reaches the voice engine, the text is transformed. An article is full of forms that are written short and said long. "30°C" is said "thirty degrees", "SNCF" is spelled out, "wedispatch.fr" is sounded out letter by letter, "1st" becomes "first". A voice engine handed "30°C" raw would hesitate, or read "thirty C". This preparation step replaces each form with what it should become to the ear.
In our case, that preparation produces two texts at once, and this is a choice many tools skip. The first, the one sent to the engine, says "thirty degrees". The second, the one that stays displayed under the player, keeps "30°C", the writer's own spelling. We fixed a bug on exactly this point in September: the text prepared for the machine was sometimes shown instead of the real one, and a reader would see "ess enn cee eff" written where "SNCF" belonged. What the machine hears and what the reader sees are two different things, and they must stay apart.
The lexicon: what no rule can guess
After the general preparation comes a step nobody can fully automate: proper nouns. The Breton town "Ploërmel" is said "Plo-air-mel", and no French reading rule guesses that. You have to know the place. This is where the pronunciation lexicon comes in, where a publisher notes once and for all how the names the engine mangles should be said, and that setting then applies to all their later articles. This part of the work is the least imitable thing a voice tool has, and it is also why you can fix a word's pronunciation without regenerating everything.
An honest aside on this: we once listened to seventy-one hard reading cases, one by one. The raw engine handled twenty-four of them on its own, with no help at all. In other words, recent voices are already good on many cases, and the temptation is to conclude there is nothing left to do. That is wrong: it is precisely the remaining cases, place names, in-house acronyms, rare words, that decide whether an article is listened to all the way through or grates by the third paragraph.
The rules of reading French
Next comes the specifically French reading: liaisons, silent final consonants, stress falling at the end of a group. These rules cannot be deduced from spelling; they have to be known. This is the work that separates a voice that has learned French from a multilingual voice that guesses at it, and it explains why a French synthetic voice sometimes sounds English: a model trained mostly on English applies the wrong reflexes. The liaison remains the hardest test, because it depends on meaning and not only on letters.
This is exactly what our French text to speech page argues: reading a language is not stringing sounds together, it is applying its rules. The voice engine supplies the timbre and the naturalness; the preparation and the lexicon supply the accuracy. Both matter, but they are not the same job.
The engine, then the re-gluing
The prepared text is finally sent to the neural engine, which produces the audio file. If the article asked for word-level alignment, to highlight the text as it is read, the word positions are returned by the machine on the transformed text, then re-glued onto the displayed text. That is why a timestamp has to account for the substitutions made upstream: "thirty degrees" takes longer to say than "30°C" takes space on screen.
What to take away
Modern text-to-speech is therefore not a single box but a chain: prepare the text, apply the client's lexicon, apply the rules of French, produce the sound, re-glue what must be. The voice engine, the part everyone talks about, is just one link. When a tool disappoints, it is almost never the timbre's fault, which is excellent everywhere today: it is the preparation that is missing, or the lexicon that does not exist. A good way to judge a tool, then, is to give it not a pretty demonstration sentence but your own hard cases: a town name, an acronym, a date, an amount. That is where the chain shows itself, or breaks.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo