Chatterbox is a free open-source voice cloning model with emotional tone control
Resemble AI has released Chatterbox, a free open-source voice cloning model that runs locally and supports emotional tone control like "dramatic" or "monotone." It clones voices using just a few seconds of audio and responds in under 200 milliseconds. The tool works on Windows, Mac, and Linux with 5–6 GB of video memory. All generated speech includes a faint watermark, "PerTh," to identify it as AI-made. According to Resemble AI, it performed better than ElevenLabs in blind tests. Currently, it only supports English.
Decoder EN demo (heightened emotional expression)
Chatterbox is licensed under MIT and targets developers. Check out the demo here.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Subscribe nowRead on for the full picture.
Subscribe for hype-free coverage.
- Full access to every article on THE DECODER
- No ads
- Join the comments and community discussions
- A weekly AI news recap via mail
- 6x/year: "AI Radar" — deep dives on the AI topics that matter most
- Daily AI news, always up to date
- Our full ten-year archive
- Covered by a team with 10+ years in AI