Glossary

Prosody

Everything in speech that is not the words themselves: rhythm, pauses, melody, intensity. It is what makes a reading sound human.

Prosody covers everything in speech that is not the choice of words: rhythm, pause length, sentence melody, variations in intensity, stress. Two people reading the same paragraph produce different sounds without changing a word; that difference is prosody.

What a human reader does without thinking

They slow down before an important piece of information. They take a breath after a quotation, and a shorter one between items in a list. They rise slightly at the end of a question, drop through a parenthetical, then pick up again. They give a date the room it deserves in the sentence. None of this is written in the text: these are decisions made while reading, based on MEANING.

Why older synthesis sounded mechanical

Concatenative systems assembled recorded fragments. The sequence of sounds was correct and the melody flat, because nothing in the process modelled the sentence as a whole. You recognised the machine within three words, and that is what durably damaged the reputation of automated reading: the voice of old satnavs and phone menus.

Neural models learned prosody by training on large volumes of natural speech. They no longer pronounce words one after another: they generate the signal taking the sentence context into account. That shift, more than timbre quality, explains why listeners often no longer wonder whether they are hearing a human.

What prosody does not repair

It does not fix a reading error. A beautifully modulated voice that says the wrong form of an ordinal is still unusable: the problem is not in the voice, it is in the text it was given. That is the job of text normalization, and that is where French reading is decided, not in the choice of engine. Prosody decides whether you listen; normalization decides whether what you hear is right.

Related terms

  • Neural voice : A neural voice is generated by a neural network trained on human speech. Its natural prosody makes it hard to distinguish from a real narrator.
  • Text normalization : The step that turns written text into speakable text: numbers become words, acronyms are expanded or spelled out, abbreviations are written in full.
  • Text-to-speech (TTS) : Text-to-speech converts written text into spoken audio. Recent neural models produce voices that are hard to tell apart from human speech.

Give your articles a voice

WeDispatch automatically turns your articles into an audio version, the moment you publish. Try it free on your own articles: no credit card needed.

Book a demo