The œ ligature: the little sign that gives a French speech engine away
Cœur, œuf, œsophage, poêle: the œ ligature does not read one single way, and 'oe' is not always an œ. What French text to speech has to know before pronouncing it.
There is a character that is not on the keyboard and yet French uses every day: the œ ligature, the "e inside the o". You find it in cœur (heart), sœur (sister), œuvre (work), bœuf (ox), vœu (wish), mœurs (mores). It looks harmless, but for a speech engine it bundles three traps into one, and a tool that treats it carelessly is caught out in the first sentence. It is exactly the kind of finish a narration's credibility rests on, which is why it deserves an article to itself.
First trap: œ does not read one single way
In words inherited from popular Latin, œ is pronounced in the "eu" family (a rounded front vowel with no close English equivalent): cœur, sœur, œuvre, vœu, nœud all take that sound. So far, an engine that ties œ to this vowel is rarely wrong.
But in learned words that came through Greek, the same ligature reads like a French "é" (close to the vowel in "day"). Œsophage (oesophagus) is said "é-zophage", œnologie (oenology) "é-nologie", œdème (oedema) "é-dème", fœtus "fé-tus", and the name Œdipe (Oedipus) is said "É-dipe". An engine that applies the "eu" reading to these words produces something like "eu-sophage", and the error is instant to any listener. The correct reading cannot be deduced from the sign: it depends on the word's origin, information the letters do not carry.
| Written word | Expected reading | Common wrong reading | Origin |
|--------------|------------------|----------------------|--------|
| cœur | "keur" | (rarely wrong) | popular Latin |
| œuvre | "euvre" | (rarely wrong) | popular Latin |
| œsophage | "é-zophage" | "eu-zophage" | Greek |
| œnologie | "é-nologie" | "eu-nologie" | Greek |
| Œdipe | "É-dipe" | "eu-dipe" | Greek |
| fœtus | "fé-tus" | "feu-tus" | learned Latin |
Second trap: "oe" is not always an œ
Because the ligature is not easy to type, many texts write "coeur", "oeuf" or "soeur" with a separate o and e. A correct engine has to recognise that this "oe" stands for the ligature and reads the same way. So far the rule looks simple: "oe" equals œ.
Except it does not, and this is the reverse trap. In a whole set of words, the o and the e belong to two different syllables and form no ligature at all. "Coexister" (to coexist) is said "ko-exister", "coefficient" is "ko-éfficient", "moelle" (marrow) is said "mwal", "poêle" (stove, or frying pan) is "pwal", "foehn" is "fène". An engine that mechanically turns every "oe" into "eu" reads coexister as "keur-xister": a disaster that nevertheless passes every test as long as you only feed it hearts and sisters. The decision is therefore not "oe reads œ", but "is this particular oe a ligature or two vowels", and that is settled word by word.
Third trap: the plural that changes everything
French keeps some of its best-known irregularities for this family, and they are heard, not just written. "Un œuf" (an egg) is said with an audible f; "des œufs" (eggs) drops the f and shifts the vowel, giving roughly "day-zeu". The same goes for bœuf: singular with an audible f, plural without it. An engine that reads the plural with the f still sounding makes a mistake every French speaker spots at once. And "œil" (eye) does not merely change sound in the plural: it changes word, since its plural is "yeux". No letter-to-sound rule predicts that; the engine has to know these forms, just as it has to handle the silent e at the end of a word.
What our engine does, without overselling
On common words, native and learned alike, the engine applies the right reading: cœur in "eu", œsophage in "é", the plural of œufs and bœufs with its shifted vowel and silent f. It also treats an "oe" typed without the ligature as a ligature when the word calls for it, and recognises the two-syllable cases such as coexister or poêle. This is the kind of invisible work that separates reading French from decoding it letter by letter, the work set out on the French text to speech page: how a word is pronounced is not read off its letters, it is known.
But let us be honest about the limit, as with homographs. A rare learned word, a niche scientific term, a coinage, a foreign proper noun carrying an œ or an æ (ex æquo, nævus, curriculum vitæ, or the first name Lætitia) can fall on the wrong side, because nothing in their spelling states their origin. The machine infers, and on these rare cases it can infer badly.
The move that guarantees the right reading
When one of these words matters, a brand name, a trade term, a surname, the guarantee comes not from the voice but from the pronunciation lexicon. You declare the expected reading once, and the engine stops guessing on that exact word: "Lætitia" will always be said "Létitia", "nævus" always "névus", whatever the context. It is the same logic as for a proper noun whose spelling says nothing about the sound.
To test a tool in thirty seconds, three words are enough: "œsophage" (expected: é), "des œufs" (expected: eu, silent f) and "coexister" (expected: ko-exister). An engine that reads all three correctly truly knows what an œ is. An engine that says "eu-zophage" or sounds the f in "des œufs" is running on a surface rule, and you will know where you stand before you fit out a whole site.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo