Glossary

Text normalization

The step that turns written text into speakable text: numbers become words, acronyms are expanded or spelled out, abbreviations are written in full.

Text normalization is the step that separates what is written from what must be said. The reader sees a date, an acronym, a score, a legal reference; the voice engine needs words. This conversion is the part of the work that decides the quality of a French reading, and it is the part almost nobody does properly.

Why the engine is not what decides

Today's speech engines are numerous and good. At equal timbre, what separates two products is the text they send to the engine. A gorgeous voice that mispronounces an ordinal, spells out an acronym letter by letter when it should be read as a word, or reads a football score as a subtraction is unusable in production, and no voice setting repairs that.

French is particularly demanding. The first day of the month is an ordinal and the others are not. A match score is spoken with a preposition. A legal reference has its own reading. An acronym is spelled out when it is not pronounceable as a word, and expanded on first use if the readership does not know it, but not afterwards, on pain of being unbearable.

Written rules or a model that guesses

Two approaches exist. A language model can rewrite the text to make it speakable; it is quick to set up and fundamentally unstable. The same article can be read two different ways two days apart, an error cannot be corrected because there is no rule to change, and the text has been touched, which an editorial AI charter often forbids.

Written rules do the opposite: they are tedious to produce, they cover one case at a time, and they are deterministic. The same sentence always yields the same reading, an error is fixed at its source, and the result can be checked. That is what makes a public record possible: we publish the test cases and what the product does with them, before you pay.

What normalization cannot guess

A town name, a surname, a local brand. No rule will say how they are pronounced. That is the job of the pronunciation lexicon, where the newsroom corrects once and the product remembers afterwards. Finally, normalization lengthens the text, which shifts every timestamp: they have to be stitched back onto the displayed text, or the highlighting and subtitles drift.

Related terms

  • Text-to-speech (TTS) : Text-to-speech converts written text into spoken audio. Recent neural models produce voices that are hard to tell apart from human speech.
  • Word-level timestamps : Word-level timestamps map every word of the text to its exact position in the audio. They power highlighting, quotes and chapters.
  • Read-along audio : Read-along audio highlights each word of the article as the voice speaks it, letting the user read and listen simultaneously on the page.

Give your articles a voice

WeDispatch automatically turns your articles into an audio version, the moment you publish. Try it free on your own articles: no credit card needed.

Book a demo