← Blog · 29 July 2026 · Lire en français

How to evaluate a neural voice before you commit

Every demo sounds good. Six tests to run on your own articles instead, plus the three questions that have nothing to do with how the voice sounds and decide everything anyway.

Every neural voice sounds good for thirty seconds. That is the whole problem. The demo you are listening to was recorded on text chosen to flatter it: short sentences, common vocabulary, no awkward proper nouns, no figures. On that ground, the gap between an excellent voice and a mediocre one is almost inaudible.

The gap shows up somewhere else, and it shows up on your content. Here is how to force it into the open before you sign, rather than three weeks into production.

If you need the definition first, it is in the glossary. This piece assumes you know what a neural voice is and now have to pick one.

Test 1: your worst paragraph, not your best

Take a real article from your site and find the passage you would struggle to read aloud yourself. It is usually a six-line sentence with two subordinate clauses, or a paragraph that stacks three institutional names in a row.

That is the paragraph to feed the voice, not your best-written column. A voice that holds up on difficult text will hold up everywhere. The reverse is never true.

Test 2: the proper nouns in your patch

This is the first breaking point, and the one your audience notices immediately. Town names, districts, councillors, sports clubs, local businesses: a generic voice will apply standard pronunciation rules, which sometimes lands correctly and sometimes produces something your readers will hear as an outright error.

Build a list of twenty proper nouns that recur in your coverage and have them read aloud. Count the mistakes. That number predicts your satisfaction six months from now better than any general impression of the timbre.

One point that is widely misunderstood: those errors are not fixed by switching voices. They are fixed with a pronunciation lexicon, a list of words and their intended pronunciation, applied to every article you publish. So the question to put to a vendor is not "does your voice say this name correctly?" but "can I teach it, and does the lesson stick?"

Test 3: numbers, dates, units

"15/03", "£1,200", "3.5%", "no. 12", "8.30pm". Each of those has one expected reading and several possible ones. A voice that reads "fifteen slash zero three" has failed at something basic, and your articles are full of it.

This test takes two minutes and eliminates more candidates than any other.

Test 4: ten minutes, not thirty seconds

Listening fatigue cannot be measured on a clip. Some voices have a rhythm defect (a cadence that is too regular, a breath in the same place in every sentence) that goes unnoticed at first and becomes wearing after five minutes.

Have a full article read. Listen to it while doing something else, the way your audience will. If you want it to stop after ten minutes, so will your reader.

Test 5: the same text, twice

Regenerate the same article the next day and compare. A voice that produces two noticeably different renderings of identical text will cause you a problem the day you need to fix one sentence in a published article: the corrected segment will not sit flush with the rest.

This is invisible in a demo and decisive in production. It determines whether fixing one passage without redoing the whole file is possible at all.

Test 6: what the voice returns besides sound

A voice that returns only an audio file quietly closes doors you have not thought about yet. Word-level timestamps, the exact position of every word in the file, are what make it possible to highlight text in step with the reading, to search inside the audio, and to generate subtitles.

Ask explicitly whether that data is available for the specific voice you are evaluating. The answer varies from voice to voice within the same catalogue, which catches people out.

The three questions that are not about sound

Once the tests are passed, three points decide how the thing actually gets used.

Cost per article, not per character. Pricing is quoted per million characters, which helps nobody decide anything. Multiply by your average article length and your monthly volume. The number that matters is the cost of one article, and you have to add regenerations, because a corrected article is an article paid for twice.

Turnaround. A piece that takes fifteen minutes to voice cannot ship alongside a breaking story. Measure the delay on text the length of yours, not on a single sentence.

What happens if the voice disappears. Catalogues change and voices get retired. Ask what becomes of your already-published articles, and whether there is a fallback path to a close-sounding voice. It is the least pleasant question to ask and the most useful to have asked.

What we do about it

At WeDispatch the evaluation happens on your own text from the start: you paste an article and listen to the result, with nothing to install. The pronunciation lexicon, the timestamps and targeted correction of a single segment are part of the offer rather than add-ons, precisely because those are the three points on which a voice choice pays for itself or costs you over time.

On choosing a voice that belongs to your title rather than one from a shared catalogue, see signature voice. And on why these voices became hard to tell apart from a human one, we wrote this.

Give your articles a voice with WeDispatch

This blog is itself voiced by WeDispatch. Curious how it sounds on your content?

Book a demo

Read next