Glossary

Word-level timestamps

Word-level timestamps map every word of the text to its exact position in the audio. They power highlighting, quotes and chapters.

Word-level timestamps are the data layer that maps every word of a text to the precise moment it is spoken in the corresponding audio file: the word 'council' starts at 42.3 seconds and ends at 42.9 seconds, for instance. Invisible to the user, this layer is what turns a plain MP3 into audio content you can actually work with.

Where timestamps come from

Two possible origins. With a human recording, text and audio must be aligned after the fact using speech recognition, with imperfect precision. With audio generated by text-to-speech, the system knows exactly when it utters each word, since it produces the signal itself: timestamps are exact by construction, with no extra processing. This is a structural advantage of generated audio over recorded audio.

What timestamps enable

Everything that connects a point in the text to a point in the sound depends on them. Read-along highlighting, which follows the voice word by word. Subtitles, where each line must appear at the right moment. Audio quote sharing: select a sentence in the article and get the exact matching sound bite, ready to circulate on social networks. Chapters, which let you link to a precise second of the listen (the start of each subheading, say) and share a URL that begins right there. And mid-article ad insertion, which must target the end of a paragraph rather than an arbitrary point in the file.

The lesson for a publisher comparing audio solutions: voice quality is what you hear, but the richness of the timing data determines everything you will be able to build afterwards. Audio without timestamps is a closed file; word-level timestamped audio is raw material. The WeDispatch audio player is built on this layer, which is why highlighting, quotes and chapters come as standard rather than as add-ons.

Related terms

  • Read-along audio : Read-along audio highlights each word of the article as the voice speaks it, letting the user read and listen simultaneously on the page.
  • Automatic subtitles : Automatic subtitles are SRT or VTT files generated from the audio, displaying the right text at the right time in videos and players.
  • Midroll : A midroll is an advertising message inserted partway through a listen, ideally at a paragraph boundary so it never interrupts a sentence.

Give your articles a voice

WeDispatch automatically turns your articles into an audio version, the moment you publish. Try it free on your own articles: no credit card needed.

Book a demo