Turning an article into audio without any AI rewriting the text
Many newsrooms have an editorial charter that bans AI-generated content. How to produce an audio version within that rule: text read exactly as written, prepared by rules rather than by a language model.
A newsroom put the question to us bluntly: "Is it exactly the same text, down to the comma, that gets read out? I have no idea." The remark was fair, and it lands on a point most audio tools would rather leave in the dark. Between the article you publish and the voice that reads it, there is usually an invisible step: a language model "prepares" the text. And a model, by its nature, can move a word, round a figure, smooth over a turn of phrase. For an editorial charter that forbids AI-generated content, that isn't a detail. It's a red line.
This matters well beyond France. Any publisher with a strict no-AI policy faces the same question the moment they consider audio, and the honest answer is usually buried.
What "preparing the text" means, and where the risk is
Before it's read aloud, an article goes through a cleaning step. Photo credits, "related reading" links, share prompts, everything written for the eye that means nothing to the ear, has to come out. Many solutions hand that job to a language model with an instruction along the lines of "don't rewrite, just prepare for reading". The problem is that the instruction is a wish, not a guarantee: the output is still produced by a model.
We measured it, on a ten-line control article run through a model. The result: "en un an" (in one year) came back as "en an" (a word dropped), a phone number lost its last pair of digits, and a unit was turned into an error that wasn't in the original. Nothing dramatic on its own, but impossible to predict and impossible to guarantee. For a publisher that puts a byline on its articles, that's the exact opposite of what it promises its readers.
The mode where no model touches the text
So WeDispatch offers a path where the text goes to no language model at all. There's only a voice engine at the end of the chain, and between the two, nothing but rules. The text that gets read is the journalist's, transformed only by readable, testable steps whose full list fits on one screen: strip the lines that make no sense out loud, restore a few safe accents, and correctly read numbers, dates, times, scores, percentages, units and currencies.
That last point is the heart of it, and it's a language job, not an artificial-intelligence one. Reading "1,200" as "one thousand two hundred", "3:30 pm" in its French form, "3.5%" as the words, each written form has an expected spoken form, and we produce it with deterministic rules rather than hoping a model guesses right. It's exactly the subject of our article on reading numbers aloud, and it's why we hold to our stance on what French quality means for a voice tool: read one language well rather than thirty badly.
What this mode gains you, and what it costs
The gain is clear: you can answer yes to the opening question. It's the same text, down to the comma, and you can prove it, because every transformation applied is logged. Nothing is rewritten, nothing is invented, nothing comes out of a model. For a charter that bans generative AI, the audio version finally becomes compatible with the house rule.
The cost, and it should be said plainly, is real. Expanding an acronym on first use, turning "the ARS" into "the Regional Health Agency", requires understanding the text, and that is beyond a rule. In this mode, acronyms are therefore not expanded automatically. That's the price of exactness, and it's yours to choose, not ours to choose for you. When an acronym recurs and must always be said the same way, the answer is still the pronunciation lexicon, which applies across the whole site and depends on no model.
See it before you spend
The right way to work with this kind of constraint isn't to trust a claim, it's to look. Before a credit is spent, you can see how an article will be read: which lines were removed, how each number and unit will be pronounced. A passage that bothers you gets fixed upstream, once, for the whole newsroom. It's the same spirit as our piece on acronyms and abbreviations: we don't hide where the reading stumbles, we show it and hand you the fix.
What we don't claim
We don't claim this mode reads everything better than a well-tuned model. A model can expand an acronym, infer context, patch a clumsy sentence; a rule cannot and never will. What this mode guarantees is something else, and it's what a charter demands: the assurance that no machine has rewritten the journalist's work. For many publishers, that assurance is worth more than an expanded acronym, because it's the condition for the audio to exist at all. The rest, a proper noun, a rare word, is handled by hand, and that control is precisely what they came looking for.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo