Reading numbers aloud in French: what a machine has to know
1 200 EUR, 15 h 30, 3.5 percent, 1er, XIXe: each written form has an expected reading and a wrong one. The table of forms that trip up French text to speech, and how we handle them.
A written figure does not tell you how it is read. In French, "1 200" reads as "mille deux cents" (one thousand two hundred), but "1200" run together can be taken for a year, "douze cents" (twelve hundred). "15/03" is a date to you and a division to a machine that has not grasped the context. French read aloud is full of these forms where spelling and pronunciation diverge, and that is exactly where a synthetic voice stumbles, often without warning.
This is not a marginal problem. A news article, a product sheet, an annual report are full of them. A voice that reads sentences perfectly and then mangles an amount or a time breaks trust at the precise moment the reader needed the exact figure. Here are the forms that actually cause trouble, what is expected of each, and the common error alongside.
The table of forms that trip a voice up
| Written form | Expected reading | Common error |
|---|---|---|
| 1 200 € | mille deux cents euros | "douze cents euros", or a pause on the space |
| 3,5 % | trois virgule cinq pour cent | "trois cinq pour cent", or reading the comma as a full stop |
| 15 h 30 | quinze heures trente | "quinze h trente", or "quinze heures trois zéro" |
| 15/03 | quinze mars | "quinze slash zéro trois", "quinze divisé par trois" |
| 1er | premier | "un e r", "un er" |
| XIXe | dix-neuvième | "X I X e", or read as a single word |
| n° 12 | numéro douze | "n degré douze", "n rond douze" |
| 2 M€ | deux millions d'euros | "deux M euros", "deux méga euros" |
| 06 12 34 56 78 | zéro six, douze, trente-quatre... | digits read one by one, no grouping |
| 1/4 | un quart | "un sur quatre", "un slash quatre" |
Each row is a test you can run against any tool, ours included, in two minutes. The right-hand column is not theoretical: these are the wrong readings you actually hear from generic voices that predict pronunciation with no model of French.
Why this is hard for a machine
A neural voice does not "understand" a number, it predicts its pronunciation from what it has learned. The same symbol changes reading depending on what surrounds it: "/" is a slash in a URL, "mars" (March) in a date, "sur" or "quart" in a fraction. "h" is an hour marker after a number and a letter elsewhere. The thin space that separates thousands, correct in French typography, looks to an ill-prepared engine like two separate numbers.
Good preparation means normalising the text before synthesis: recognising that a pattern is a date, an amount, a percentage or a phone number, and rewriting it in the form the voice will read correctly. This is rule work, not voice work. A beautiful voice fed badly will misread an amount; an average voice prepared well will read it right. That is the whole logic of the French text to speech page: what matters happens before any sound comes out.
What WeDispatch does, without overpromising
The forms in the table above (amounts, percentages, times, common numeric dates, everyday ordinals, reference numbers) are handled by the normalisation applied to every article, in both reading modes. That is the baseline, taken as a given rather than an option.
What remains are the cases that are ambiguous even to a human. "1200" with no separator can be a year or a plain number depending on the sense of the sentence, and no rule settles it with certainty. Two things there. First, well-typed text removes the ambiguity at the source: writing "1 200" with a non-breaking space and "1200" for a year is already half the work, and it is good writing practice regardless of audio. Second, for a particular reading that must hold across all your articles, the pronunciation lexicon lets you declare the expected form once, without regenerating anything.
We would rather put it that way than claim a machine always guesses. It does not guess: it applies solid rules to the large majority of cases, and hands you the rest.
The test to run before you choose
Take a real article from your site, not a demo text. Find the five or six places where a figure carries information: a price, a date, a statistic, a time, a reference number. Have the tool you are evaluating read the passage and listen only to those spots. It is faster and far more predictive than a general listen, because these forms are the ones that recur every day in your content.
The same principle applies to liaison, the other great tell of machine-read French, which we detailed in liaison, the hardest test for machine-read French. And if you want a full method for choosing between two voices, it is in how to evaluate a neural voice. Numbers alone already rule out more candidates than most people expect.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo