Writing a text-to-speech RFP: what to actually require
A public body or a media outlet putting its content out to tender often writes 'natural voice' and little else. Here are the requirements that can be verified, and our own answer to each, including the places where we answer no.
When a local authority, a media outlet or an institution puts the voicing of its content out to tender, the quality of what it gets back depends first on what it managed to ask for. A specification that says no more than "a natural, pleasant voice" commits nobody: it cannot be checked, and two vendors will both answer "yes" without lying and without resembling each other. The useful requirements are the ones a bidder can prove and an evaluator can test. Here are five, with the answer we give to each, limits included. A checklist is worth more when the person holding it agrees to go through it first.
1. An accessible player, not just an accessible voice
Adding audio does not make a site accessible if the player itself is not. The requirement to write down: the playback component must be usable with the keyboard alone, announced correctly by a screen reader, with a visible focus, enough contrast, and respect for the "reduced motion" setting. Ask the bidder to name the standard's criteria it checks, and how.
Our answer: the player is built around those criteria, and we spell out which ones in our piece on the accessible audio player and WCAG criteria. Where we promise nothing beyond what is real: an accessible player dropped onto a page that is not accessible does not make the page compliant. Audio is one brick, not a badge.
2. Hosting and the path the text takes
For a public actor, knowing where the text goes and for how long is not optional. The requirement: the bidder must describe the path a text takes through its service, where it is processed, and how long it is actually kept.
Our answer: we write that path down in plain terms, hosting included, in where your text goes, and it is exactly the kind of question we handle on our French text-to-speech page. We do not claim to be everyone's choice: an organisation under a very particular hosting constraint should hold our path up against its own and draw its own conclusion.
3. Reversibility, the clause people forget at signing
This is the most overlooked requirement, and the one that costs the most when it is missing. You have to write down what the vendor gives back when the contract ends: the audio files produced, the texts, and what becomes of a podcast feed that listeners are already subscribed to.
Our answer: a customer who leaves takes their files and their texts, and their feed stays a standard feed, readable by the platforms without going through us. We will not unpack here a case that deserves its own article, but the rule we follow is simple: what you produced is yours. A badly written reversibility clause turns a change of vendor into a fresh start, and that is the kind of trap you only see on the way out.
4. Price per character, and what it covers
Text-to-speech is billed by volume of text, not by number of articles nor by a flat monthly fee. The requirement: ask for a price per character or per sign, and ask what is counted (is a regeneration after a correction billed again? a partial re-read?). That is where offers really separate.
Our answer: we bill by volume, and the cost shows up before production, in listening time and in credits, never after. The detail of the calculation and the edge cases is in how much article audio costs. We do not quote a price per million characters in a blog post, because a rate card belongs on the pricing page, kept current, not in an article that will go stale.
5. Pronunciation, and the right to correct it
A voice that mangles a proper noun, a trade acronym or a local term drags down everything else. The requirement: the bidder must let you correct a pronunciation and store it, without regenerating the whole article by hand every time.
Our answer: a pronunciation lexicon lets you settle a word once and apply it from then on, and we flag words whose pronunciation is doubtful before reading rather than letting them slip through. What we do not do: we do not "perform" a text like an actor. Irony, reported anger, grief stay out of reach of a synthetic voice, and we would rather say so than promise it.
What to take away
The five requirements copy straight into a specification: accessible player, documented text path, written reversibility, price per character with its edge cases, correctable pronunciation. Each one can be verified, and that is the whole point: it forces the bidder to answer something other than "yes, of course." A good tender is not looking for the vendor who promises the most, it is looking for the one whose answers hold up under testing, including when the honest answer is "no, not yet."
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo