Voice Cloning With Consent: Rules, Risks and Good Practice

05/10/2026

Voice Cloning With Consent: Rules, Risks and Good Practice

Twenty seconds. Shorter than a rambling birthday voice message. The voice cloning feature in s2speech needs only 20 seconds to 2 minutes of recording to create a new voice, on the condition that the voice owner consents. Precisely because it is that easy, the question worth asking is no longer “how do we do it?” but “whose voice is it, who said yes, what is it for, and until when?”

This piece is for the people about to press the button: a marketing team that wants a brand voice, a training center that wants an instructor to “read” a whole series of lessons without spending a week in a studio, a call center that wants to keep its familiar greeting voice. All legitimate needs. And all of them can turn into trouble if handled carelessly.

Why consent is not a box-ticking exercise

A voice is deeply personal. Your family recognizes you from a single “hello”. So, at times, do colleagues and partners. A copy of someone's voice therefore carries the trust other people place in that person. Using it without permission is taking something that does not belong to you.

Everyone has heard about the dark side: calls faking a relative's voice to borrow money urgently, or faking a boss's voice to push through a transfer. Scammers don't need a studio, just a few videos people post on social media themselves. A serious platform has to pick a side. On s2speech, voice cloning comes with the condition that the voice owner consents, and businesses that use it should carry the same principle into their own processes.

AIVISION solution demo
Short video: s2speech

A minimum rulebook before you clone

  • Written consent with a clear scope. Which channels the voice will be used on, for what kind of content, for how long, and whether it may be edited. “I agree the company can use my voice” is far too broad.
  • The right to withdraw. Voice owners can change their minds. The agreement should state where the voice will be removed from, and within what time frame.
  • Fair pay. If someone's voice helps a business make money, compensation should be discussed up front.
  • No cloning of celebrities or people who have passed away without valid permission from whoever holds the rights.
  • Labeling. Listeners should be told they are hearing an AI voice, especially on a phone call.

Legally, a voice is tied to a person's identity and may be treated as personal data under data protection rules in many places. Rules differ by country and change over time, so before going live, have your legal team or a lawyer review the consent form. AI can help draft it; people sign it off.

The real risk lives in the process

Picture this: Mai, the company's familiar call center voice for years, leaves her job. Her AI voice keeps greeting customers every day. Is that okay? If the original agreement says nothing about it, the answer is “nobody knows”, and that is exactly the risk. A lot of trouble with cloned voices starts with questions nobody wrote down:

  • Who in the company may create content with the cloned voice? Does anyone approve it before it goes out?
  • Where are the original recordings stored, who can access them, and when are they deleted?
  • Is there a log of what content has been generated with that voice?
  • When the contract ends or the owner withdraws consent, who is responsible for taking it down?

One rule we think every business should have: never use a cloned voice to approve transactions or order money transfers, not even internally. The reverse holds too: don't treat a voice as enough proof of someone's identity over the phone. Call back on a number you know, and ask something only the real person would know.

Doing it right, starting at the recording session

A lean but proper process doesn't need anything fancy. Before the session, the voice owner signs consent, keeps a copy, and knows exactly what they are agreeing to. During the session, pick a quiet room, read clearly, and keep the recording between 20 seconds and 2 minutes; clean audio in means a clean cloned voice out. Before the voice goes into use, let its owner hear their own AI voice. Once it is live, restrict use to the right people and the right projects, and tell listeners clearly when a voice is AI-generated.

Then put a regular review on the calendar: which voices are still in use, which agreements are about to expire, which voices need to come down. It sounds dull, but it is exactly what answers the question about Mai above.

If you don't truly need a custom voice, the natural Vietnamese and English voices already available in s2speech Text-to-Speech may well be enough, and they raise no consent questions at all. Service details are on the s2speech product page, and you can listen to the available voices at s2speech.com.

Voice cloning will only get easier. What keeps it decent is not technical difficulty but the habit of pausing to ask one question before pressing the button: has this person really said yes?

Related insights

See all insights