Voice cloning explained

Voice Clone replaces a speaker's voice in a video with a different one — built from a short reference recording of the new voice — while keeping the original timing, pacing, and pauses intact. It's a separate tool from face swapping, works alongside it or on its own, and has a couple of rules that don't apply to video or photo swapping.

What it actually does

Starting from a video, Voice Clone identifies the speech of whoever you choose to replace, learns the vocal character of a reference recording you provide, and resynthesizes that person's lines in the new voice — matched to the original's timing, so the result still lines up with their mouth movements and the rest of the scene. It isn't text-to-speech reading a script; it's a transformation of the original spoken audio, so hesitations, pacing, and emphasis carry over.

Why it needs an account

Unlike video and photo face swapping, which work for anyone without signing up, Voice Clone requires a free account. The reason is simple: a voice profile is something you build once from a reference recording and then reuse across future projects, and reusing it only makes sense if it's actually saved somewhere tied to you. An anonymous, account-less session has nowhere to keep that profile for next time.

Consent applies here too

The same rule that governs face swapping governs voice cloning: you need permission from the person whose voice you're cloning, and from whoever provided the reference recording if that's a different person. A convincing clone of someone's voice used without their agreement is exactly the kind of misuse Recast's Terms of Service and watermarking exist to prevent — every cloned-voice result carries an inaudible audio watermark for that reason, alongside the visible label on the video itself.

What makes a good reference recording

• Clear audio with minimal background noise — a quiet room beats a recording with music, wind, or crowd noise in it.
• A reasonable length of actual speech — a few seconds of mostly silence gives the clone much less to learn from than a clip of continuous talking.
• A single speaker — a reference clip with two people talking over each other makes it hard to isolate the voice you actually want.

What it's not

Voice Clone doesn't generate new dialogue — it transforms what was actually said in the source video into a different voice, timed to match. It won't make someone say words they never said; the words come from the original recording.
← Back to Recast
Recast
AI face and voice swapping — multiple faces, cloned voices, and reusable Characters.
support@recast-face-swap.com