Homographs: when French text to speech has to guess how to read a word
Est, fils, president, couvent: the same written word reads two ways depending on meaning. How a French speech engine decides, and where it gets it wrong.
There is a class of traps that French text to speech cannot solve by looking at the letters, because the letters are identical. These are heterophonic homographs: words spelled the same but pronounced differently depending on meaning. In French, "il est là" (he is there) and "le soleil se lève à l'est" (the sun rises in the east) both contain "est", yet one is said "eh" and the other "est". No letter-to-sound rule separates them. Only context does, and that is exactly where an engine reveals whether it reads French or merely decodes it.
Why this is harder than liaison
A liaison is decided at the boundary between two words, with grammar rules you can code. A homograph is decided inside a single word, and the only clue available is the role that word plays in the sentence. The engine has to parse the sentence, recognise that here "couvent" is a verb (the embers smoulder) and not a noun (the convent), before it even picks a sound. That is a decision of understanding, not of pronunciation. A tool that treats each word in isolation has no chance: it will always play the same card.
Five cases where the same word reads two ways
Here are five classic French homographs, with the two expected readings and the clue that settles the choice.
| Written word | First reading | Second reading | What decides |
|--------------|---------------|----------------|--------------|
| est | "eh" (verb, to be: il est tard) | "est" (compass point: vers l'est) | A determiner or preposition standing before it |
| fils | "fees" (the child: mon fils) | "feese" (thread, wire: des fils électriques) | Meaning, often the plural and the domain |
| couvent | "coove" (verb: les braises couvent) | "coo-vahn" (noun: entrer au couvent) | The grammatical function in the sentence |
| president | "prezidahn" (noun: le président parle) | "prezident" (verb: ils président la séance) | A plural subject "ils" forcing the verb |
| portions | "por-syon" (noun: deux portions) | "por-tyon" (verb, to carry: nous portions des sacs) | The pronoun "nous" and the tense of the story |
You could line up dozens more: "affluent", "négligent", "content", "violent", "éditions". What they share is that the correct reading is never found inside the word, only around it.
What our engine does, without overselling
On the most frequent cases, the engine decides by analysing the structure of the sentence: a plural subject "ils" followed by "président" tips it towards the verb, a determiner before "est" marks it as a probable compass point. This analysis is what separates reading French from plain phonetic playback, and it is the heart of the work described on the French text to speech page: how a word is pronounced depends on the whole sentence, not on the word taken alone.
But let us be honest about the limit. A short or ambiguous sentence may not give enough clues. "Les fils" with no context can mean the sons or the wires, and no grammatical analysis will resolve it, because the ambiguity exists for a human reader too. An elliptical headline, a caption, a two-word table cell: these are the places where the engine can pick the wrong card. We would rather say so than pretend the machine "understands". It infers, and sometimes it infers badly.
The move that guarantees the right reading
When a homograph really matters (a proper noun, a trade term, a headline where the error would be heard at once), the guarantee comes not from the voice but from the pronunciation lexicon. There you declare, once and for all, how a specific word must be read in your content. If your outlet is named "L'Est Républicain", you fix the reading "est" and never touch it again: the engine stops guessing on that word and applies your decision. It is the same principle as for proper nouns and place names, where the spelling says nothing about the expected sound.
How to test a tool in two minutes
Take three minimal sentences and listen: "Ils président la réunion" (expected: verb), "Les poules couvent leurs œufs" (expected: verb), "Le vent d'est se lève" (expected: compass point). A tool that reads all three correctly genuinely parses the sentence. A tool that says "prezidahn" for a plural subject is working from isolated words, and you will know where you stand before you fit out a whole site. It is the same reflex as for reading numbers aloud: you do not judge an engine on a nice paragraph, you judge it on its traps.
Homographs are not a purist's detail. On a news article, a single wrong reading in a headline is enough to make the ear drop out and signal that the audio was bolted on without care. That is exactly the kind of finish that separates a tool which pronounces French from a tool which reads it.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo