Skip to content
DATASHEET // TORTOISE-TTS

Tortoise TTS

Slow, but the quality is worth the wait.

OPEN SOURCE4–8 GB VRAMRuns locally
Actually FreeNo SignupOpen SourceWatermark-FreeHobbyist-OK
Visit Tortoise TTSUPDATED 2024-11-02 · DIRECT LINK
github.com/neonbjb/tortoise-tts
Tortoise TTS — preview image

HARDWARE REQUIREMENTS //

Runs locally · Entry GPU (6–8 GB)

4–8 GB VRAM
Min VRAM
4 GB
Rec. VRAM
8 GB
Min RAM
8 GB
Rec. RAM
16 GB
Disk
5 GB
GPU class
Entry GPU
11.6+No Apple SiliconCPU-CapableQuant: FP16

Workable at 4 GB VRAM; 8 GB recommended. CPU is impractical.

[ EDITORIAL PICK ]

Why we recommend Tortoise TTS

DERIVED FROM METADATA — NOT SPONSORED
  • Open source

    Source is public — you can audit it, fork it, and you'll never lose access to your workflows if Tortoise TTS the company changes direction.

  • Runs on 4 GB

    Fits on entry-level cards (GTX 1660, RTX 3050, RTX 4060). Rare for this category.

  • Beginner-friendly

    You don't need to read a paper before getting your first result — sensible defaults and a quick install.

[ EVIDENCE NOTE ]

Documentation-led datasheet

This page summarizes upstream documentation, release information, and editorially reviewed catalogue fields. It is not presented as a hands-on benchmark. Verify changing requirements at the official project; report stale data through our corrections channel.

AT-A-GLANCE SIGNALS //

DERIVED FROM THIS PAGE'S DATA
  • Install difficulty
    Easy

    Runs CPU-only — no CUDA / driver gymnastics required.

  • Hardware comfort
    Entry-level

    Fits on 4 GB cards — GTX 1660 / RTX 3050 territory.

  • Ecosystem
    Open source

    Source is public — auditable and forkable, no vendor lock.

  • Verification
    Stale

    621 days since the last refresh — treat hardware numbers as a floor, not a ceiling.

[ MORE IN THIS NICHE ]

Other audio & speech generation tools we rate

Three picks across different tradeoffs — so you don't end up with three near-clones of Tortoise TTS.

What is Tortoise TTS?

Tortoise is the high-quality, low-speed TTS that set the bar before XTTS landed. Multi-step diffusion-style generation produces remarkably natural prosody from just a few seconds of reference audio. Now mostly displaced by faster models, but still notable for clone fidelity on a budget.

Pros & cons

✓ PROS

  • Excellent prosody and naturalness vs. its 2022 contemporaries
  • Voice cloning from a few seconds of reference
  • Mature codebase with many community forks

– CONS

  • Glacially slow — minutes per sentence on consumer GPUs
  • Newer models (XTTS, F5-TTS) match quality with 10–50× speedup

What's actually free?

Apache 2.0; weights and code both free.

✓ Actually FreeNo SignupOpen SourceWatermark-Free

Alternatives

Bark

Suno's expressive transformer-based TTS.

OPEN SOURCE8–12 GB VRAM
VRAM fit8–12 GB

F5-TTS

Zero-shot voice cloning TTS — 15 s of audio is enough.

OPEN SOURCE8–12 GB VRAM
VRAM fit8–12 GB