Neural voice
A neural voice is generated by a neural network trained on human speech. Its natural prosody makes it hard to distinguish from a real narrator.
A neural voice is a synthetic voice produced by a deep learning model, as opposed to the older concatenative voices assembled from recorded fragments. The difference is audible within seconds, and it comes down to one word: prosody.
Prosody, or why you can no longer tell
Prosody covers everything in speech that is not the words themselves: rhythm, pauses, the melody of a sentence, shifts in emphasis. A human narrator slows down before a key fact, breathes after a quote, rises slightly at the end of a question. Older synthesis engines ignored all of this; neural models learned it by training on thousands of hours of natural speech.
The result is that the model does not pronounce words one after another, it performs a sentence. It can tell a list from an aside, a quote from the text around it. That is why a listener discovering a well-produced audio article often stops wondering whether they are hearing a human or a machine at all.
What matters for a newsroom
Neural voices are not interchangeable when the material is news copy. Three criteria separate them. Endurance first: a voice that convinces over two sentences can turn monotonous over a 700-word article. Language handling second: local place names, acronyms, figures and quotes are where weaker models stumble. Consistency last: the tone must hold from one article to the next, because the voice becomes part of the publication's sound identity.
That last point leads to voice cloning: a newsroom can turn one of its journalists' voices into a dedicated neural model, a signature voice, with the legal safeguards that requires. For publishers, the practical takeaway is simple: the technology has crossed the threshold where audio quality is no longer the obstacle to offering every article in audio.
Related terms
- Text-to-speech (TTS) : Text-to-speech converts written text into spoken audio. Recent neural models produce voices that are hard to tell apart from human speech.
- Signature voice : A signature voice is the cloned voice of a journalist or newsroom, used with written consent to narrate the publication's articles.
Give your articles a voice
WeDispatch automatically turns your articles into an audio version, the moment you publish. Try it free on your own articles: no credit card needed.
Book a demo