← All articles
Voice and MusicAugust 12, 2026 · 12 min read

How Much Audio Do You Need to Clone a Voice?

How Much Audio Do You Need to Clone a Voice?

You open a voice cloning tool, it asks for an audio file, and one question stops you: how much audio to clone a voice do you actually need? Some services say fifteen seconds. Others ask for half an hour of recording, sometimes several hours. Both claims are accurate, but they describe two different operations.

The gap comes from the technology involved, not from anyone overselling. Once that confusion clears, a second surprise waits for most people: past a certain point, adding minutes changes nothing, while a single flaw in the take can ruin a ten minute sample. This guide covers the lengths that matter, how to record a usable sample, and the permissions to gather before you click. If the topic is new to you, our complete guide to AI voice over sets the scene first.

The short answer

Fifteen seconds of clean audio is enough to get a recognisable cloned voice with today's instant engines, and thirty seconds to two minutes is the range where the result genuinely stabilises. Past five minutes, the added value fades. The thirty minutes to three hours some services request belong to a different method, trained cloning, which builds a model dedicated to one voice. In both cases the cleanliness of the recording matters more than its length, and consent from the person concerned comes before everything else.

Instant cloning and trained cloning: two methods, two scales

Instant cloning, what research calls zero-shot cloning, retrains nothing. The model has already learned, from thousands of speakers, what makes a human voice. Your sample only tells it which voice to imitate: it extracts an embedding, a compact signature of the timbre, then applies that signature to any text. That is why it works in seconds and needs only a very short file.

Published work gives the scale of it. Microsoft Research presented VALL-E in January 2023, a model that reconstructed a timbre from a three second excerpt. OpenAI announced a system called Voice Engine in March 2024, able to produce a synthetic voice from a fifteen second recording, and explained in the same post why it was deliberately limiting its release. Three seconds is enough to recognise a voice. It is not enough to work with one.

Trained cloning runs the other way: a model is genuinely fine tuned on the supplied voice, which demands far more material. Public documentation from services offering this option converges on a floor of around thirty minutes of recording, with two to three hours recommended for a brand voice used intensively. The cost is not only studio time: the take must be consistent, captured under identical conditions from start to finish, and the training itself runs for hours.

What each length of sample buys you

Between fifteen seconds and five minutes, progress is anything but linear. Each step fixes one specific flaw, and it helps to know which one before deciding whether your take deserves a second try.

Sample length scale for voice cloning: 15 seconds, 2 minutes, 5 minutes and 30 minutes
Instant cloning learns fast, then stops learning almost entirely.

Below ten seconds, most engines reject the file or return an unstable voice. Between ten and fifteen seconds the timbre comes through: someone who knows you will recognise you. What is still missing is prosody, the music of the sentence that makes a question rise, a list settle and an ending fall. The model invents it from its general habits, so the result sounds right in places and generic in others.

Between thirty seconds and two minutes, you give it something to observe: your breathing, the way you attack a sentence, the length of your pauses. This is the step that turns an entertaining clone into a voice over you can put on a three minute script. Beyond that, up to five minutes, the gains sit in the details: rare words, proper nouns, intonation shifts inside a long paragraph.

Why extra minutes stop helping so quickly

An instant engine compresses your sample into a fixed size embedding. That embedding holds a few hundred values, not a recording. Once it is full, feeding it more audio is like pouring water into a full glass: the surplus is not stored, it is averaged with the rest.

That technical detail has a practical consequence people discover too late. A long, uneven sample often produces a less faithful clone than a short, consistent one, because the average of an energetic voice and a tired voice sounds like neither. If you supply five minutes, supply five minutes recorded in one sitting, in the same room, with the same energy.

Recording quality outweighs length

The engine does not separate your voice from what surrounds it. Room reverb, a fan, background music: all of it enters the embedding and comes back out in every generated sentence, including on scripts you never spoke. This is the most common mistake, and the most expensive, because it cannot be fixed afterwards. Music belongs in the video, never in the sample: it goes on at the edit, and our prompt formulas for composing AI music show how to describe it by genre and by mood.

Anatomy of a good voice cloning sample: continuous reading, quiet room, steady distance, and the noises that ruin a take
Whatever the microphone picks up is learned along with the timbre.

Gear matters less than people assume. A phone microphone held about twenty centimetres from the mouth, in a furnished bedroom with a bed and curtains, gives a better sample than a decent microphone standing in the middle of a tiled living room.

  • Record in one take: a file stitched from fragments creates breaks the model reads as your speaking style.
  • Keep a constant distance from the microphone, about a hand's width, without leaning in on important lines.
  • Switch off anything that blows: air conditioning, fans, extractor hoods, phone notifications.
  • Read at your normal pace, the one your videos will use, not a presentation voice.
  • Vary sentence shapes: one long, one short, a question, an exclamation.
  • Close the window, and prefer a furnished room to an empty office.
  • Listen back on headphones before uploading: whatever bothers you will bother the model.

What should you read during those two minutes?

The content of the sample steers the clone as much as its length. Read something close to what the voice will say later. For narration, take a passage of story. For explainer videos, take an explanation. A voice sampled in an advertising tone keeps a trace of that tone even on calm scripts. If your project calls instead for strongly characterised voices, a child, an animal, a storyteller, cloning is the wrong door: our guide to character voices for animated videos shows how to build them from the available voices and a reading style.

Three traps recur. Reading a list of numbers or isolated words produces choppy prosody. Reciting a memorised text too fast removes the breathing. And a sample pulled from a video call has already lost part of its high frequencies to compression, which the model will faithfully reproduce.

Formats, file size and technical limits

Engines generally accept MP3, WAV and M4A. Uncompressed WAV is the best choice when you have it, because it adds no further loss. An MP3 recorded directly by a phone is perfectly fine. The problem cases are MP3 files re-encoded several times, or audio extracted from an already compressed video. Size and duration caps vary between engines: in our studio the premium engine takes samples from ten seconds to five minutes, while the standard engine tolerates heavier files.

Consent comes before the technique

No sample length makes it lawful to clone a voice you have no right to use. The rule fits in one sentence: clone your own voice, or one belonging to a person who gave explicit, written, dated permission for a named use. The topic deserves more than a paragraph, and our review of the legal framework for voice cloning covers the edge cases.

Four checks before cloning a voice: whose voice it is, written permission, defined use, disclosure to the audience
A single negative answer is enough to stop the process.

The rules have tightened. The European regulation on artificial intelligence, published in 2024, requires under its article 50 that audio and video content imitating real people be flagged as artificial, and those transparency obligations have applied since 2 August 2026. In France, article 226-8 of the Criminal Code punishes publishing an edit made with a person's words without their consent when it is not obviously an edit, and the law of 21 May 2024 on securing the digital space extended that text to content produced by algorithmic processing. In the United States, Tennessee brought voice explicitly into protected attributes with the ELVIS Act, in force since 1 July 2024, and the Federal Communications Commission ruled in February 2024 that AI cloned voices in automated calls count as artificial voices, prohibited without prior consent. Celebrity voices are the sharpest case of all, and what you are allowed to do with a famous voice is worth reading first.

Disclosing a synthetic voice on the platforms

Platforms moved before lawmakers did. The YouTube Help Centre has required creators since March 2024 to disclose, at upload, any realistic altered or synthetically generated content, which includes a voice imitating a real person. YouTube extended its privacy process in July 2024 so that a person can request removal of content reproducing their face or voice. TikTok's community guidelines likewise require realistic AI generated content to be labelled. Ticking the box does not hurt a video's reach. Skipping it invites a takedown. The voice is also not the only part of the soundtrack to check before publishing: our article on copyright and AI generated music sets out what you can release and monetise without a claim.

How cloning works in our studio

In the EasyVids studio, cloning starts from the voice over screen: you name the voice, upload an MP3, WAV or M4A file, and the interface recommends at least fifteen seconds. Two engines are available, a standard model and a premium one. The cloned voice joins your personal list and can be reused across every project, from a single video to multi character dialogue. A certification is requested at every cloning and never remembered from one time to the next: you confirm the voice is yours, or that you hold explicit permission from the person concerned. Cloned voices stay visible to your account alone and can be deleted on request, and our plans are set out on the pricing page.

Frequently asked questions

Is fifteen seconds really enough to clone a voice?

Yes, for a recognisable voice with an instant engine, provided those fifteen seconds are clean and continuous. Timbre is captured very quickly. What fifteen seconds lacks is rhythmic stability on long scripts, and moving to one or two minutes visibly improves narration.

Do you need more audio to clone a voice in another language?

No, the required length is the same. Recent engines carry a timbre across languages from one sample. The real caveat is elsewhere: the original accent often stays audible. When you can, record the sample in the language the voice will speak.

Can you clone a voice from a video found online?

Extraction is technically possible. Legally it is not, without the person's agreement. Quality suffers too, since music, editing and compression are learned along with the voice. A public video is not permission, and platform terms grant no rights over the voices of the people filmed.

How long does the cloning itself take?

With an instant engine it takes tens of seconds: the file is analysed, the embedding computed, the voice appears in your list. Trained cloning needs several hours of processing after upload, on top of the time spent producing the recordings.

Can a cloned voice be deleted afterwards?

Yes. A voice created in your account stays attached to that account and can be removed on request, along with the content made from it. When cloning someone else's voice, write that right of withdrawal into the original agreement: it is the clause that prevents most disputes.

Keep the order of priorities in mind: permission first, a clean take second, length only third. One minute recorded in a quiet room, read at your normal pace, beats ten minutes captured in a noisy one almost every time. To try it on your own voice and hear it read your scripts, creating an account opens the full studio, cloning included.

Go from reading to creating

50 free credits when you sign up, no bank card.

Create my first video