Voice Cloning
Record or upload a short voice sample, then generate new speech in that voice. Runs entirely in your browser, no upload, no signup.
1. Reference voice
A clean 6-10 second clip works best: one speaker, minimal background noise. Silence is trimmed and the first 6 seconds of speech are used.
Cloning a voice without consent may be illegal where you live. This runs entirely in your browser, so we can't verify consent: that's on you.
Stays on your device. The voice model runs in your browser; your clip, text and generated audio never leave it. No upload, no signup, unlimited use.
Get the next tool first
New local-AI tools and updates by email. Unsubscribe anytime.
Free voice cloning that runs in your browser
Record or upload a short clip of a voice, then generate new speech in that voice from any text, without uploading a single sound to a server. This tool runs Chatterbox Turbo, an open-weight zero-shot voice cloning model from Resemble AI, entirely inside your browser using WebGPU. No account, no API key, no per-minute fee, and nothing ever leaves your device.
The same engine drives both privacy and price. Because the model lives in your browser, we have no inference bill to recover, no logs to keep, and nothing to upsell. Once the model is cached it even works offline.
How to use it
- Record a 6 to 10 second clip from your microphone, or upload an existing one.
- Wait for the reference voice to encode. The first run also downloads the voice model once, after which it is cached.
- Type or paste up to 2,000 characters of text.
- Confirm the consent checkbox, then click Generate speech.
- Play it back in the built-in player and download the WAV when you are happy.
What it is good for
Narrating a video or presentation in your own voice without re-recording every revision, creating personalized voice messages, dubbing your own voiceovers into new scripts, and accessibility use cases like preserving a voice for someone who may lose the ability to speak. Because the audio is a plain WAV file, it drops straight into any video editor or audio tool, and there is no usage cap to work around.
Use it responsibly
Voice cloning is powerful, and that power cuts both ways. Only clone your own voice, or a voice you have explicit permission to clone. This tool has no server and no account, which means privacy for you, but it also means there is no one checking consent on the other end. That responsibility is yours. Impersonating someone without their permission can be illegal depending on where you live, independent of which tool was used to do it.
Private and free, by design
Cloud voice-cloning services upload your audio to their servers, meter you per character or minute past a small free tier, and require an account. This tool does the opposite. The model runs on your own hardware, so your voice sample never leaves the browser, there is nothing to log, and there is no bill to pass on to you. Chatterbox Turbo and its ONNX export are MIT licensed. Without WebGPU, and on phones, the tool runs Pocket TTS by Kyutai, exported to ONNX by KevinAHM and used under CC BY 4.0.
Want a text-to-speech tool with a curated set of ready-made voices instead of cloning your own? Try Text to Speech.
Frequently asked questions
- Does my voice sample get uploaded anywhere?
- No. The reference clip you record or upload, the text you type, and the generated audio all stay in your browser the entire time. The voice model runs locally, so there is no server in the loop and nothing for us to log or retain.
- How much reference audio do I need?
- A clean 6 to 10 second clip of one person talking, with minimal background noise or music. Silence at the start and end is trimmed and only the first 6 seconds of speech are used, because longer references make the model drift into mumbling and cost time on every sentence it generates. Record straight from your microphone or upload an existing recording.
- Why does it download a model the first time?
- For the cloning to happen on your device, the neural voice model has to live on your device. The first time you submit a reference clip, your browser downloads it and caches it: about 750 MB for Chatterbox Turbo on a computer with WebGPU, or about 140 MB for Pocket TTS everywhere else. After that it loads from the cached copy. You can review or delete cached models from the Manage models button in the page header.
- Why is the download so large?
- Zero-shot voice cloning needs a larger model than a fixed-voice text-to-speech tool because it has to understand and reproduce an arbitrary, never-seen-before voice from a short sample, not just play back one of a small number of pre-trained voices. It is a one-time download, cached by your browser, and it never touches our servers. Phones and browsers without WebGPU get Pocket TTS instead, which is about 140 MB.
- Can I use the audio commercially?
- Chatterbox Turbo and its ONNX export are released under the MIT license, which permits commercial use of the software. Whether you personally have the right to use a specific voice, including your own, is a separate legal and ethical question that is on you to answer honestly, not something this tool can verify.
- Is it okay to clone someone else’s voice?
- Only with that person’s clear permission, or for your own voice. Cloning someone’s voice without consent can violate impersonation, fraud, publicity-rights, or deepfake laws depending on where you live, regardless of the tool used. Because this runs entirely in your browser with no account and no server logs, there is no way for us to verify consent, which means the responsibility sits entirely with you.
- How long can the text be?
- Up to 2,000 characters at a time. Longer text is automatically split into sentence-sized chunks and generated one after another, then stitched into a single continuous WAV with short natural pauses between them.
- Does it work on iPhone and Android?
- Yes. Phones get Pocket TTS, a 100M-parameter model that downloads about 140 MB and runs on the CPU, because Chatterbox Turbo’s ~750 MB of weights exceed what a phone browser tab can hold. Generation on a phone is slower than on a computer, and we have not measured it on every device, so expect a wait on long text.
- How does this compare to ElevenLabs or PlayHT?
- Cloud voice-cloning services typically require an account and meter you by character or minute once you exceed a small free tier. This tool has no account, no API key, no character limit tied to a subscription tier, and no per-minute fee, because the model runs on your own hardware instead of a paid inference server.
- Can I download the audio?
- Yes. Every generation produces a standard WAV file you can download with one click and use in videos, voice messages, dubbing, accessibility recordings, or anywhere else.
- Does voice cloning work without WebGPU?
- Yes. With WebGPU, which recent desktop Chrome, Edge, Safari and Firefox offer, you get Chatterbox Turbo on your graphics card. Without it you get Kyutai’s Pocket TTS on the CPU: a smaller download and a quicker start, with slightly plainer delivery. The chip in the status panel tells you which one you got: “Runs on your GPU” is Chatterbox Turbo, “Runs on your CPU” is Pocket TTS.