Voice cloning: Voice cloning is the use of AI to create a synthetic copy of a specific person's voice from a short audio sample, so that new text can be spoken in that voice. It powers useful narration tools and also some of the most convincing scams.
Modern text-to-speech systems can learn the timbre, pacing, and accent of a voice from anywhere between a few seconds and a few minutes of recorded audio. Feed the clone a script and it speaks it, including words the original person never said. Tools such as ElevenLabs offer this to ordinary users, and editors such as Descript build it into their workflow; higher-quality clones usually ask for more sample audio and an explicit consent step.
Legitimate uses are real: narrating a training course without booking a studio, fixing one flubbed word in a recorded podcast without re-recording the segment, publishing a Spanish version of your explainer video in your own voice, or preserving the voice of someone losing theirs to illness.
The risks are equally real. Cloned voices have been used to impersonate executives authorizing wire transfers and to fake distress calls from relatives. In the United States, the Federal Communications Commission ruled in 2024 that AI-generated voices in robocalls count as artificial voices under the Telephone Consumer Protection Act, making unsolicited cloned-voice calls illegal in most cases. Practical defenses: never approve a payment or credential change on the strength of a voice alone, agree on a callback procedure or a shared verification phrase, and treat urgency in a voice message as a warning sign. Only clone your own voice or one you have written permission to use; most vendors require it, and likeness rights vary by state.
Example at work
An HR manager records twenty minutes of sample audio, clones her voice, and produces narrated onboarding videos that she updates every quarter by editing the script instead of re-recording. The company also adds a line to its finance policy: any payment request received by phone or voicemail must be confirmed through a second channel.
Why it matters
A familiar voice used to be proof of identity. It no longer is. That changes both how you produce audio content and how your team verifies requests that arrive by phone.
Related terms
- DeepfakeA deepfake is synthetic audio, video, or imagery, generated or altered by AI, that convincingly depicts a real person saying or doing something they never did.
- Text-to-speech (TTS)Text-to-speech (TTS) is AI that converts written text into natural-sounding spoken audio. Modern systems produce voices with realistic pacing and emotion, and some can clone a specific person's voice from a short sample.
- Speech recognitionSpeech recognition, also called speech-to-text or automatic speech recognition (ASR), is AI that converts spoken audio into written text. It powers dictation, meeting transcription, voicemail transcripts, and voice commands.
- AI avatarAn AI avatar is a synthetic on-screen presenter, either a stock character or a digital likeness of a real person, that speaks any script you type. Tools such as Synthesia and HeyGen use it to turn text into talking-head video without a camera or crew.
- AI watermarkingAI watermarking embeds a hidden or attached signal in AI-generated content, such as an invisible pattern in an image or metadata in a file, so that its origin can be identified later. It helps with provenance but is far from foolproof.