Quick Answer
voicebox is a free, open-source voice studio running on your own computer, bundling seven text-to-speech engines and voice cloning across 23 languages. It is MIT licensed, so output is yours to sell. Unlike paid services, local tools enforce no consent checks, and voice cloning law tightened significantly in 2026.
voicebox is a free, open-source voice studio that runs on your own computer, bundling seven text-to-speech engines and voice cloning across 23 languages. It is MIT licensed, so anything you generate is yours to sell. The genuine catch is not quality or cost. It is that local tools have no consent enforcement whatsoever, and voice cloning law tightened significantly in 2026.
Key takeaways
- voicebox has 52,100 GitHub stars and an MIT licence, the most permissive option in this category, so commercial use is unambiguous.
- It ships prebuilt installers for Mac and Windows. The download is small, but the first run pulls roughly 15-20GB of model weights.
- Several rival "free" voice tools are commercially unusable: XTTS-v2's weights are non-commercial and Coqui shut down in 2024, so there is nobody left to buy a licence from.
- The EU AI Act's transparency rules came into force on 2 August 2026 and apply to individuals acting professionally, not just companies.
- YouTube explicitly does not require disclosure when you clone your own voice for voiceovers, which is the use case most creators actually have.
What voicebox does
voicebox is a desktop application for generating speech from text and cloning voices, running entirely on your own machine. Its distinguishing feature is that it bundles seven different TTS engines rather than committing to one: Qwen3-TTS, Qwen CustomVoice, LuxTTS, Chatterbox Multilingual, Chatterbox Turbo, TADA and Kokoro. It covers 23 languages and ships more than 50 preset voices.
That variety matters because these engines trade off differently. Some are small and fast; others are large and higher quality. Being able to switch without reinstalling anything is genuinely useful.
Voice cloning works zero-shot, meaning you supply a reference recording and it imitates that voice without any training step. The documentation does not specify a required sample length, so plan to experiment. For comparison, GPT-SoVITS documents 5 seconds for zero-shot cloning and about a minute for fine-tuning.
What you need to run it
voicebox ships prebuilt installers, which puts it ahead of most open-source AI tools for accessibility. There is a .dmg for Apple Silicon Macs and an .exe installer for Windows, each around 515MB.
The installer size is misleading. On first run the app downloads model weights, and a practitioner review testing on an M2 Pro with 16GB of RAM reported 15 to 20GB of weights pulled down. Budget the disk space and the time.
The same review measured the engines on that machine: LuxTTS around 1GB running at roughly 150 times realtime on CPU, Kokoro about 500MB at around 50 times realtime, Qwen3-TTS 0.6B around 3GB, and the largest TADA model needing about 10GB. That is a single reviewer's measurements on one machine rather than official benchmarks, so treat them as indicative.
The practical takeaway is that the smaller engines are fast enough on ordinary hardware without a GPU, and only the largest models demand serious memory.
The free voice tools you should avoid
This is where the research gets genuinely useful, because several widely recommended "free" tools cannot legally be used for commercial work.
XTTS-v2 (Coqui) is the most important trap. The code is MPL-2.0, which is permissive, but the model weights are released under a non-commercial licence. Worse, Coqui Inc. shut down in January 2024. There is no company left to sell you a commercial licence, so there is no legal route to commercial use. The original repository is unmaintained; a community fork continues under Idiap.
ChatTTS is doubly restricted: AGPLv3 for the code and CC BY-NC 4.0 for the model, so commercial use is prohibited on both counts. It also deliberately degrades its own output, adding high-frequency noise and MP3 compression as an anti-misuse measure.
Real-Time-Voice-Cloning, despite 60,100 stars, carries a note from its own author saying the repository "has quickly gotten old," pointing users elsewhere.
Play.ht is gone entirely. Meta acqui-hired the team in July 2025 and the platform shut down on 31 December 2025. Its domain no longer resolves. It still appears in most 2026 "best TTS tools" listicles, which tells you how much of that content is written without checking.
That leaves voicebox (MIT) and GPT-SoVITS (MIT) as the two genuinely commercial-safe open-source options.
The consent problem nobody enforces
Here is the difference between free and paid voice tools that matters most, and it has nothing to do with audio quality.
ElevenLabs requires a recorded consent statement spoken in the target voice before it will clone professionally. That is a technical gate. You cannot clone someone without their participation.
voicebox ships a responsible-use policy requiring you to clone only your own voice or one you have explicit permission to use, prohibiting impersonation, fraud and commercial use of a person's voice without legal right.
That policy is a text file, not a control. No local open-source tool can enforce consent, because the software runs on your machine with no gatekeeper. The entire legal responsibility transfers to you, and the law has moved quickly.
What the law actually says in 2026
The legal picture changed substantially, and most articles either ignore it or get it wrong.
EU AI Act Article 50 has applied since 2 August 2026. It is already in force, not "coming soon." It binds individuals acting professionally, not only companies. Deepfake disclosure must be clear, distinguishable and perceivable by humans, so embedded metadata alone is not sufficient. Artistic and satirical works get lighter-touch treatment.
In the United States, the NO FAKES Act has not passed; it was reintroduced in May 2026 and reported out of committee in June. State law is ahead of federal: Tennessee's ELVIS Act took effect on 1 July 2024 and was the first US law to add voice to right-of-publicity protections. California's AB 2602, requiring specific consent for digital voice replicas in contracts, took effect on 1 January 2025, and AB 1836 covering deceased personalities followed on 1 January 2026.
In the United Kingdom, the Data (Use and Access) Act 2025 section 138 came into force on 6 February 2026, criminalising the creation and requesting of non-consensual intimate images. This point is consistently misreported: it covers images only, not voice or audio. The UK has no dedicated voice-cloning statute, and recourse runs through passing-off, data protection or fraud law.
Denmark is introducing a neighbouring right over voice and appearance lasting 50 years after death, expected in force July 2026, with Ireland signalling it may follow.
Platform rules matter too. YouTube requires disclosure for realistic synthetic content, including making a real person appear to say something, and non-disclosure risks removal or monetization suspension. Crucially, YouTube does not require disclosure for cloning your own voice for voiceovers or dubbing. TikTok requires labelling and bans synthetic media of real private individuals outright, even when labelled.
Who should use voicebox
Use it if you want unlimited voiceovers at no cost, you are cloning your own voice, you value keeping audio on your own machine, and you can spare 20GB of disk.
Use a paid service if you need the absolute best quality, want indemnity and support, or are cloning anyone else's voice, where a provider's consent workflow is genuine protection rather than bureaucracy.
For the quality and cost comparison, see voicebox versus ElevenLabs and Murf. For the wider open-source field, read the best free AI voice generators.
The verdict
voicebox is the best free voice tool available right now, and its MIT licence removes the commercial ambiguity that makes most of its rivals unusable for paid work. For cloning your own voice for narration, it is genuinely excellent and costs nothing.
Just be clear about what you are taking on. The software will happily clone anyone whose audio you feed it, and it will not ask whether you have permission. In 2026 that question increasingly has a legal answer, and it is yours to get right.
The legal side deserves its own answer: is AI voice cloning legal? covers the 2026 rules across the US, UK and EU.
Related reading: best AI voice generators that sound human and how to clone your voice with AI.
Frequently Asked Questions
Is voicebox free for commercial use?
Yes. voicebox is MIT licensed, the most permissive common open-source licence, so both the software and anything you generate are unrestricted commercially. That is unusual in this category and is its main advantage over rivals.
Which free AI voice tools cannot be used commercially?
XTTS-v2's model weights are non-commercial, and because Coqui Inc. shut down in January 2024 there is no longer anyone to buy a commercial licence from. ChatTTS is restricted on both axes, AGPLv3 code and CC BY-NC model, and deliberately degrades its own output. Play.ht shut down entirely on 31 December 2025.
Is it legal to clone someone's voice with AI?
Only with their consent, and the law tightened in 2026. The EU AI Act's transparency rules have applied since 2 August 2026 and bind individuals acting professionally. Tennessee's ELVIS Act and California's AB 2602 and AB 1836 create US state protections. The UK's 2026 image law covers images only, not voice.
Does YouTube require disclosure for AI voiceovers?
Not for your own cloned voice. YouTube requires disclosure for realistic synthetic content, including making a real person appear to say something, but explicitly does not require it when you clone your own voice for voiceovers or dubbing. TikTok requires labelling and bans synthetic media of real private individuals outright.
What hardware does voicebox need?
It ships prebuilt installers for Apple Silicon Macs and Windows, each around 515MB. First run downloads roughly 15 to 20GB of model weights. Smaller engines like LuxTTS and Kokoro run fast on CPU without a GPU; the largest models need around 10GB of memory.
