XTTS v2 (Coqui)
Multilingual voice cloning in 6 seconds.
HARDWARE REQUIREMENTS //
Runs locally · Entry GPU (6–8 GB)
Real-time on 4 GB+. Apple Silicon MPS works.
Why we recommend XTTS v2 (Coqui)
- Open source
Source is public — you can audit it, fork it, and you'll never lose access to your workflows if XTTS v2 (Coqui) the company changes direction.
- Runs on 4 GB
Fits on entry-level cards (GTX 1660, RTX 3050, RTX 4060). Rare for this category.
- Apple Silicon
Native Metal / MPS support — runs on M-series Macs without CUDA gymnastics.
- Beginner-friendly
You don't need to read a paper before getting your first result — sensible defaults and a quick install.
Documentation-led datasheet
This page summarizes upstream documentation, release information, and editorially reviewed catalogue fields. It is not presented as a hands-on benchmark. Verify changing requirements at the official project; report stale data through our corrections channel.
AT-A-GLANCE SIGNALS //
DERIVED FROM THIS PAGE'S DATA- Install difficultyEasy
Runs CPU-only — no CUDA / driver gymnastics required.
- Hardware comfortEntry-level
Fits on 4 GB cards — GTX 1660 / RTX 3050 territory.
- EcosystemStrong devkit
Open-source AND ships an API — easy to integrate, possible to host yourself.
- VerificationStale
289 days since the last refresh — treat hardware numbers as a floor, not a ceiling.
Tutorials & deep-dives for XTTS v2 (Coqui)
Hand-picked from YouTube, Reddit, GitHub, and the wider web. Each link goes straight to the source — we don't intercept or rewrite anything.
Other audio & speech generation tools we rate
Three picks across different tradeoffs — so you don't end up with three near-clones of XTTS v2 (Coqui).
RVC (Retrieval-based Voice Conversion)
The voice-changer that took over Discord.
MMAudio
Generate synchronized audio for any silent video.
ElevenLabs
The benchmark commercial TTS / voice clone API.
What is XTTS v2 (Coqui)?
Coqui's XTTS v2 is the production TTS workhorse: clone a voice from 6 seconds of audio, generate speech in 17 languages, run on a 4 GB GPU. Coqui the company is gone but the model lives on under a permissive licence, and it's the backbone of most current OSS voice apps.
Pros & cons
✓ PROS
- 6-second voice cloning that actually works
- 17 languages including cross-lingual cloning
- Real-time on a single mid-range GPU
- Used as the engine inside many higher-level apps
– CONS
- Coqui (the company) shut down — community-maintained from here
- Licence is permissive but not OSI-approved; check before commercial use
What's actually free?
Coqui Public Model Licence — free for personal & commercial use.