← Blog · 26 September 2026 · Lire en français

Image captions and photo credits: what French text to speech does with them

The caption under an image, the '© Photo: agency name' that follows it: these are text, so a voice engine can read them. The question is not whether it can, but whether it should, and how the copyright symbol, agency initialisms and credit lines sound to the ear.

An image in an article almost always comes with two pieces of text: a caption, which says what you are looking at, and a credit, which says who produced it. On the page they are set apart by size, position, often italics. To the ear that visual hierarchy vanishes: a text to speech engine sees only a run of words, and it reads them flat, along with the rest. So the real question is not whether an engine can read a caption, but which parts of these lines deserve to be heard, and how the parts you keep actually sound.

The caption, often the only description available

Start with what has value when listening. A caption is not decorative: it is frequently the only description of the image available to someone who does not have the screen in front of them. "Les manifestants devant la préfecture, mardi matin" places a scene the listener will not see. On that count a good caption is worth reading: it carries information, not just a label. It has something in common with the alt text of images, that description meant for screen readers which we covered in image descriptions and alt text: where the caption describes for everyone, the alt text describes for those who cannot see, and listening brings both needs together.

Still, a caption read as is sometimes arrives with no transition. "... a voté le budget. Les manifestants devant la préfecture, mardi matin. Le débat reprendra en septembre." The listener has no way of knowing the middle sentence sat under a photo. An engine does not flag that shift, because nothing in the raw text tells it to. If you want the caption to be heard as a caption, a short lead-in phrase in the text ("in the photo, ...") does the work the layout did for the eye.

The credit, typographic noise in its purest form

The photo credit is a different case. "© Photo : Agence France-Presse" performs a legal and editorial duty that is indispensable on the page, but when listening it adds nothing to the understanding of the article, and it can sound wrong. The "©" symbol first: depending on the engine it is ignored, which is what you want, or spoken as "copyright," or left aside with an awkward gap. Then the agency names, often initialisms: how they are said depends entirely on the engine's rule for acronyms, a subject we set out in acronyms and initialisms aloud. An initialism that should be spelled out but that the engine tries to pronounce yields an invented word; the reverse happens too. Finally the photographer's proper name, which usually has no reason to be spoken in the middle of a long-form piece.

The credit is therefore, for listening, almost pure typographic noise: a line that exists for perfectly valid reasons, none of which concern the listener. The point is not to fix it, it is to decide whether it should be read at all.

Deciding, without touching the visible text

The principle is the same as for any line that serves the eye and not the ear: the visible text stays exact, and it is the spoken version that adjusts. An image credit, in the vast majority of cases, is better left unread: it adds nothing to the argument and introduces noise. The caption, on the other hand, is most often worth reading, because it describes. That distinction, caption yes, credit no, is a simple rule that holds for almost every article, and it spares you having to decide image by image.

When a caption or a credit you want to keep contains a term the engine mishandles, an agency name to spell out, a foreign surname, a place name, you do not rewrite the displayed text. You go through the pronunciation lexicon, which declares the intended reading of the term once and applies it everywhere, captions included. The text stays what the reader sees, the correction lives beside it. More broadly, these lines belong to the class of elements to keep out of the reading, which we covered in what should not be read aloud.

What to take away

An image brings two pieces of text, and they do not share the same fate to the ear. The caption describes, it often deserves to be heard, and a brief lead-in gives it back the relief the layout gave it on the page. The credit serves rights and attribution, it adds nothing to the listening and is most often better dropped, all the more so because the copyright symbol and agency initialisms sound wrong. Deciding this once, at the level of the article rather than each photo, is enough. That kind of sorting, between what is seen and what is heard, is what separates a text read well from a text read mechanically, and it is the work we do on reading French aloud, set out on the French text to speech page.

Give your articles a voice with WeDispatch

This blog is itself voiced by WeDispatch. Curious how it sounds on your content?

Book a demo

Read next