1. Home
  2. Glossary
  3. Text-to-speech (TTS)
AI glossary · Media & modalities

Text-to-speech (TTS)

Text-to-speech (TTS): Text-to-speech (TTS) is AI that converts written text into natural-sounding spoken audio. Modern systems produce voices with realistic pacing and emotion, and some can clone a specific person's voice from a short sample.

Older text-to-speech sounded like a robot reading a list. Neural TTS models, trained on many hours of human speech, produce voices that pause, emphasize, and shift tone the way a person does. Services from ElevenLabs, OpenAI, Google, Amazon, and Microsoft offer dozens of stock voices and support for many languages.

Workplace uses include narrating training videos and e-learning modules, producing audio versions of reports and articles, voicing phone system menus and customer notifications, prototyping ad reads before hiring talent, and accessibility for colleagues who prefer or need audio.

Voice cloning raises the stakes. Cloning your own voice for internal narration can be a time-saver; cloning anyone else's without written permission is a legal and ethical problem, and cloned voices are a growing fraud vector. Set a rule that no employee's or customer's voice is cloned without explicit consent, and add a verification step to any process where a voice alone would authorize action.

Example at work

A training manager turns a 40-page onboarding guide into a series of short audio modules. She pastes each section into a TTS tool, picks a warm, neutral stock voice, adjusts a few pronunciations of product names, and publishes the audio alongside the written guide so field staff can listen on the drive to a site.

Why it matters

TTS lets you produce professional audio without a studio, a voice actor, or a re-record every time the content changes. Understanding what it can do, and where voice cloning crosses a line, keeps you productive and out of trouble.

Related terms