Listening to a long text in French TTS: listening fatigue, and what actually helps
Past a quarter of an hour, the ear drops out. The thresholds where a long text is flagged before you launch, the spoken summary, and the split into episodes and chapters, described from what is actually implemented.
Reading a thousand-word article aloud is easy. Listening to a sixty-page report in one sitting is another matter. Past a certain point, the ear drops out: not because the voice is bad, but because a long listen with no landmarks, no breathing between parts, no sense of where you are, wears attention down. So the question is not only « is the voice good? » but « what, in the tool, makes a long text actually listenable? ». Here is what is in place, and what is not, said from the code.
A long text flags itself before you launch
The first useful move is not technical, it is informational. Three jobs in thirteen that were failing at one client had something in common: articles over thirty-eight thousand characters, which nothing on screen set apart from a thousand-word piece. The cost in credits was visible, but not the listening time or the risk. Two thresholds are now shown before you generate.
From around fifteen thousand characters, roughly a quarter of an hour of listening, the screen announces the duration and, if your account is entitled and has not switched it on, offers the spoken summary. The idea is simple: nobody listens for twenty minutes without knowing what it is about. Giving the gist in thirty seconds, at the head of the player, lets the listener decide whether to go on. That is what our piece on the audio summary before the full article sets out.
Above a hundred thousand characters, the message changes in kind: the text is too long for a single audio, because the maximum file size is near and the spoken text is longer than the written one (« 1914 » is said « mille neuf cent quatorze », so it inflates the count). The screen then offers to split it in two, or to remove passages. The button is greyed out exactly where generation would refuse, with the reason: no nasty surprise at launch.
Splitting into episodes, for genuinely long documents
A site article fits in one audio. A document you receive (a report, a dissertation, a sixty-page PDF) does not, and two and a half hours in one block will not be listened to on a commute anyway. For those cases, private listening splits the document into episodes, and the logic is worth knowing because it targets listening fatigue directly.
The target is about thirty minutes per episode, set against a real reading rate (around one thousand one hundred and fifty characters a minute). The split never falls in the middle of a word: it looks first for a paragraph end, failing that a sentence end, near the target. A document that fits in a single episode stays whole; as soon as there is more than one, each carries its position, « (2/5) », so you know where you are.
Chapters, when the document has them
Better than a cut by metronome: cut where the author cut. When a document carries chapters, the split follows them. It recognises a heading by its shape: a keyword followed by a number (« Chapitre 3 », « Partie II », « Section 1 »), a number at the start of the line (« 1. Introduction »), a Roman numeral, or a short line in capitals. A chapter that runs too long is re-cut in turn into half-hour slices; chapters that are too short are grouped, so you do not end up with a string of one-minute episodes. The result is a listen that breathes, with landmarks in the right places, just as you find a passage again thanks to chapters and the timestamped link.
This machinery is the one behind private document listening, the one that turns a pile of texts into a listening queue, as we explain about listening to your articles and PDFs like a radio. That is where it helps most: a long personal document becomes a series of episodes you listen to at your own pace.
What stays a matter of language
The split helps, but it does not do everything. A long listen also depends on what happens inside each sentence: a sustainable pace, pauses placed right on the punctuation, a reading that does not stumble on an acronym or a number every ten lines. That is the groundwork on French text to speech, and it counts double on a long text: a flaw you forgive over two minutes becomes unbearable over thirty. A long text well listened to is therefore two things together: a split that gives landmarks, and a reading that does not tire in itself. One without the other is not enough.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo