Abbreviations, units and symbols: the small print that trips up French text to speech
M., Mme, cf., n, km, %, EUR, square metres: each has an expected reading aloud and a wrong one that gives the machine away. A clear table, and the move to fix what stumbles, once and for all.
A French text is full of little forms you read without thinking on the page, and that turn into traps the moment a voice has to say them aloud. "M. Dupont habite au n 12, a 3 km du centre, dans un studio de 25 m2 loue 600 EUR par mois." You read that sentence correctly in one pass. A speech tool, though, has to decide how to say "M.", "n", "km", "m2" and the euro sign, and each decision can land right or wrong. This piece walks through that small print of abbreviations, units and symbols, separate from the numbers and the acronyms we have covered elsewhere.
Three families, one demand
These forms fall into three families. Word abbreviations (M. for Monsieur, cf. for confer, etc. for et cetera): the voice must restore the whole word, not the letters. Units of measure (km, kg, %, the euro sign): they have to be read out in full and agree with the number in front of them. Symbols (the number sign, the ampersand, the at sign): each has a spoken name that has nothing to do with its shape. In all three cases the demand is the same: what is written short must be said long, and said right.
The table of everyday cases
Here are the forms that come up most, with the expected reading and the error you hear when a tool gets it wrong. The examples are French, because that is what the voice is reading.
| Written form | Expected reading | Common error |
|---|---|---|
| M. Dupont | Monsieur Dupont | the letter "em", or "metre" |
| Mme, Mlle | Madame, Mademoiselle | spelled out, or skipped |
| Me Martin | Maitre Martin | "me", like the pronoun |
| Dr, Pr | Docteur, Professeur | "der", "peer" |
| cf. | confer (or "voir") | "cee-eff" |
| c.-a-d. | c'est-a-dire | letters spelled out |
| n 12 | numero douze | "n" then "degree" |
| p. 40 | page quarante | the letter "pee" |
| 3 km | trois kilometres | "trois ka-em" |
| 25 m2 | vingt-cinq metres carres | "metre deux" |
| 90 km/h | quatre-vingt-dix kilometres par heure | "ka-em slash h" |
| 20 % | vingt pour cent | "vingt pourcentage" |
| 600 EUR | six cents euros | "euro" singular, or "E" |
| 12 M EUR | douze millions d'euros | "douze M euros" |
| 18 C | dix-huit degres Celsius | "degre C" |
| & | et | the symbol's name read aloud |
| @ | arobase | "at" the English way |
The table is not exhaustive, but it shows the pattern: each short form has a precise spoken form, and the error is almost always either to spell out what should be said, or to read a symbol as a character instead of by its name.
How a tool copes, and where it stalls
A neural voice handles the most frequent forms well, because it has seen them thousands of times in training: "M." before a name almost always becomes "Monsieur", "%" becomes "pour cent", "km" becomes "kilometres". Our engine expands these common cases with no setup. Where it gets harder is on ambiguous and rare forms. "M." at the start of a sentence can be Monsieur or a first-name initial. "Me" is Maitre for a lawyer but reads "me" elsewhere. A unit stuck to a number with no space, or an abbreviation specific to your field, comes out wrong more often. It is the same phenomenon as with acronyms: on the rare and the ambiguous, the machine makes a guess, and a guess is sometimes wrong.
Switching voices fixes nothing, because another voice will make the same assumptions about the same words. The real question to put to a tool is therefore not "does it read my abbreviations well?" but "can I force the right reading where it goes wrong, and does that instruction hold across all my articles?". That is the line our French text to speech page defends: a solid base on the common cases, and the author left in control of the special ones.
The fix
At WeDispatch, these cases are fixed through the pronunciation lexicon, a short list you keep yourself, one line per entry. In it you write the form as it should be said: a trade abbreviation the voice spells out, say, or a symbol it names badly. The rule then applies identically to every synthesis, across all following articles, with no need to rewrite your texts or regenerate the ones already produced. The text shown on screen, when read-along is on, follows the corrected pronunciation, so the highlighted word is the one you hear.
What to take away
The right reflex comes in three steps: spot the five or six short forms that recur in your content (the courtesy title, the in-house unit, the recurring symbol), listen to what the voice does with them, and fix in one pass the ones that stumble. After that the subject disappears, exactly as with a proper noun set once. It is the spirit of the test we recommend before any commitment, described in how to evaluate a neural voice: you do not expect a tool to guess everything, you expect it to let you pin down what matters.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo