← Blog · 29 September 2026 · Lire en français

Reading a PDF aloud: what grates on the ear, and what we do with it

A PDF is built for the eye: repeated headers, page numbers, words split at line ends, columns, footnotes. What should not be read, and what our extraction removes or lets through, said without rounding off.

A PDF was never meant to be listened to. It was laid out for the eye: a running title repeated at the top of every page, a page number at the bottom, words split by a hyphen when the line runs out, sometimes two columns, footnote markers in superscript, a footer with the source. All of that is invisible when you read quickly, because the eye skips what repeats. But a voice skips nothing: it reads what it is given, in the order it receives it. Hence the honest question to ask before handing a PDF to a voice tool: what exactly does it do with it?

What our extraction does, and nothing more

When you drop in a PDF, WeDispatch pulls out the text, joins all the pages into one stream, then does a light tidy-up: it collapses runs of spaces into one, brings together blank lines that follow each other, and trims the leading and trailing whitespace. It also looks for a title, in the file's metadata if it carries one, otherwise in the first non-empty line. That is all. This step is purely mechanical, and it is deliberately cautious.

What it does not do deserves to be said plainly, because that is where the listening comfort is decided. It does not go hunting for running titles or page numbers. It does not rejoin a word split at the end of a line. It does not reorder columns. It does not delete footnote markers. Those moves look desirable, and sometimes they are, but each of them gets it wrong one time in ten: a page number mistaken for a number in the text, a column recomposed in the wrong order, an important note thrown out with the noise. Getting it wrong there means removing real content. Between reading a little noise and deleting a sentence from the document, we chose to remove nothing we are not sure we recognise.

What it changes when you listen

The consequence is simple: a well-built PDF listens beautifully, and a cluttered one carries part of its layout into the reading. A report that repeats its title at the top of every page will make you hear that title at regular intervals. Text set in two tight columns may come out in an order that is not the reading order. This is not a curse, it is a property of the file you supply, and knowing it spares you the bad surprise. The voice itself does its job well on what it receives: it is the same engine that voices your articles, and our approach to French text to speech holds here as everywhere else, on numbers, acronyms and proper nouns.

The rest of the chain is the same as for an article

Once the text is extracted, it joins the path common to everything we read. The French pronunciation rules apply, deterministic and written down, with no model: this is what spells out "SNCF" letter by letter, or corrects a name through your lexicon. Text preparation happens on our European infrastructure, and in charter mode that step is purely mechanical. In other words, what touches your document stays legible and checkable, and a name said wrong is corrected exactly as it is for a published article. For the part that really should not be heard (the typographic noise of a text), our piece on what should not be read aloud sets out what the reading handles by itself and what is left to you.

The cleanest way to hand over a text

If you want the tidiest result, the text you paste in yourself always beats the PDF, because you control its content: no repeated header, no line break, no column. The PDF stays the handiest option when the document only exists in that form (a report you received, an official note, a set of minutes), and for those cases, the extraction does what it knows how to do without ever mangling the substance. That is also what makes listening to personal documents useful day to day, as we explain in listening to your articles and PDFs like a radio: you drop in what you have, you listen to it, and you know what has been read.

Give your articles a voice with WeDispatch

This blog is itself voiced by WeDispatch. Curious how it sounds on your content?

Book a demo

Read next