Reading an interview aloud: the questions, the answers, and two voices
An interview does not read like a running article. The rise and fall of questions in sequence, and telling the interviewer from the person answering: what French text to speech does with it, and where it stops.
The interview is one of the most common formats in the press, and one of the most distinctive to the ear. A running article has a single thread and a single implied voice, the journalist's. An interview has two voices answering each other, and it moves by its own mechanics: a question, an answer, a follow-up, another answer. Turned into audio without care, it becomes a block where you no longer know who is speaking or when a question begins. Two difficulties hide there, and they do not have the same answer: the melody of the questions, and telling the two speakers apart.
An interview is not an ordinary article
In a standard article the pitch rises and falls with the sentences, but the register stays level. In an interview the text alternates between two modes every few lines: a question, often short and leaning toward an answer, then a passage that can run a whole paragraph. A synthetic voice that reads everything on the same tone erases that alternation, and the listener loses the rhythm of the exchange. Making an interview listenable is therefore not only about pronouncing the words well: it is about conveying that a conversation is taking place.
The questions first: intonation in sequence
The first task is the questions. In French, a question is recognised by the voice rising at the end of the sentence, and the signal that triggers that rise is the question mark. In an interview, the three ways of asking a question sit side by side on the same page, and they do not call for the same curve.
| Written form | What marks the question | Expected intonation |
|--------------|-------------------------|---------------------|
| Pensez-vous revenir? | subject and verb inverted | a clear rise on the final syllable |
| Est-ce que le budget suffira? | the phrase "est-ce que" up front | a moderate rise, set from the start |
| Et après? | nothing but the question mark | a sharp rise, carried by the mark alone |
The third line is the most fragile, and the most frequent in a lively interview: short follow-ups ("And then?", "Really?") hold up only through their question mark. Our engine takes that mark as its main signal, so these follow-ups rise correctly, on one condition, which is the most useful reminder in this article: the question mark has to be written. A follow-up typed without it will be read as a flat statement. This is the same stance we set out on the French text to speech page: reading a language well means recovering what a human reader does without thinking, not decoding letters. The three forms and their limits are covered on their own in the piece on question intonation.
Two voices, on one condition
Then comes the question every editor asks: can you give one voice to the interviewer and another to the person answering? Yes, and the mechanism is the same as for written dialogue. An interview laid out line by line, where each turn begins with a name followed by a colon ("Dupont: we made the call"), is recognised as a dialogue. Provided the account has declared its characters and each one's voice, the attributed line is read in the character's voice to the end, and the leading name stays in the body voice, as a plain attribution. You set these once, and the exchange gains its relief.
When the interview is not laid out that way but written in running prose with the answers inside quotation marks, the logic of quotations applies instead: speech inside the French guillemets or curly quotes switches to a second voice, and the rest keeps the journalist's. Both paths are described in the article on multiple voices in an article.
What the tool does not guess
Honesty means saying where the machine stops. Without declared characters and without a recognisable layout, the tool does not guess who is speaking: it reads everything in a single voice, and that is the right choice, because inventing a wrong attribution is heard at once. The questions are still intoned correctly thanks to their punctuation, but telling the two speakers apart calls for the step described above.
Two shades still escape automatic reading. The rhetorical question ("Need we really say it again?"), which expects no answer, calls for a less rising intonation that the engine cannot tell from a real question. And the irony of an answer, the innuendo of a follow-up, are not performed: an interview is not acted out like a scene. What is rendered is the structure of the exchange and the melody of the questions; the rest is an approximation we would rather announce than promise.
The step that gives an interview relief
In practice, two settings are enough. Lay the interview out with a name and a colon at the head of each turn, then declare the interviewer's voice and the guest's. Short follow-ups keep their question mark, and the reading then restores what makes an interview alive: two people answering each other, each recognisable by ear. For a proper name whose pronunciation is not obvious, the pronunciation lexicon fixes the reading once and for all, so the guest is named correctly from one end of the interview to the other.
Give your articles a voice with WeDispatch
This blog is itself voiced by WeDispatch. Curious how it sounds on your content?
Book a demo