← Blog · 6 August 2026 · Lire en français

Acronyms in French audio: SNCF is spelled out, OTAN is said as a word

SNCF is spelled letter by letter, OTAN is said as a word, CAF depends on context. How a synthetic voice decides, where it gets it wrong, and the exact move to teach it the right reading for good.

There are two ways to read a string of capital letters aloud, and French picks one or the other without warning. SNCF (the national railway) is spelled out, letter by letter: "ess-en-cé-eff". OTAN (NATO) is said as a word: "otan", not "o-té-a-enne". Nothing in the spelling tells you which: usage decides, and usage follows no rule you can derive. For a synthetic voice this is a real breaking point, because a mishandled acronym is heard instantly and gives the impression that nobody listened to the result.

Three families, three behaviours

The first case is acronyms that are spelled out because they cannot be pronounced as a word: SNCF, RATP, HDMI, PDG, RGPD. The voice must say each letter. The classic error is to attempt a "French" pronunciation of an unpronounceable string, which produces a mumble.

The second case is acronyms that have become words in their own right: OTAN, SIDA (AIDS), OVNI (UFO), ONU (the UN). You say them, you no longer spell them. The opposite error happens here: a voice that spells "o-vé-en-i" where everyone says "ovni" sounds just as wrong.

The third case is the trickiest: the same acronym changes reading depending on context or speaker. CAF may be spelled ("cé-a-eff", the family benefits office) or said differently across regions and habits. There is no single truth: there is a choice, and that choice should be yours, not chance's.

How a machine decides

A neural voice predicts the reading of an acronym from what it saw in training. For very common acronyms it stands a good chance of getting it right, because spelled-out "SNCF" and said "OTAN" appear everywhere in the corpora. The trouble starts with less frequent acronyms, those specific to your field or region, and above all the rare ones the model has almost never met. There it makes a guess, and a guess is wrong roughly half the time between spelling and pronouncing.

That is why you do not fix this by changing voice. Another voice will make exactly the same guesses on the same words. The question to ask a vendor is therefore not "does your voice read this acronym correctly?" but "can I impose the right reading, and does that instruction hold across all my articles?". This is the same French-preparation logic set out on the French text to speech page: the foundation is solid, and you keep control over the special cases.

The exact move in WeDispatch

The fix goes through the pronunciation lexicon, a small list you keep yourself, one line per entry, in the form spelling = pronunciation. For an acronym that should be spelled out but the voice tries to pronounce, you write the letters as they should be said. For an acronym that should be pronounced but the voice spells, you write the expected phonetic word. For instance: URSSAF = Urssaf to force the one-word pronunciation, or a phonetic spelling to force a precise reading.

Three properties matter. First, it is a rule, not an artificial intelligence: it applies identically to every synthesis, across all following articles, including in brand-voice mode. Second, it lives beside the text, not inside it: you do not rewrite your articles, and you do not regenerate those already produced to add an entry that will apply to the next ones. Third, and this counts, the text shown on screen (highlighting, subtitles) follows the corrected pronunciation, so the word emphasised on screen is the one you hear. A sub-editor fills this lexicon between deadlines, with no technical skill.

What it changes in practice

For a regional outlet, a public body, a firm or a company that speaks its trade language, acronyms are not a detail: they are the words the audience knows best and notices most when they sound wrong. The right reflex is to build, once, the list of the twenty or thirty acronyms that recur in your content, check what the voice does with them, and correct in one pass the ones that catch. After that, the topic disappears.

This is exactly the spirit of the test we recommend before any commitment, described in how to evaluate a neural voice, and it echoes the work on proper nouns covered in pronouncing proper nouns and place names in news audio. A mishandled acronym and a mangled town name call for the same remedy: a clear rule, written once, that holds over time.

Give your articles a voice with WeDispatch

This blog is itself voiced by WeDispatch. Curious how it sounds on your content?

Book a demo

Read next