Stable Audio Open
Open-weight text-to-audio — 47-second sound effects and music.
Generate synchronized audio for any silent video.
Runs locally · Entry GPU (6–8 GB)
Small/medium variants on 8 GB; large on 12 GB.
Source is public — you can audit it, fork it, and you'll never lose access to your workflows if MMAudio the company changes direction.
Comfortable on a mid-range consumer card — no need to remortgage for an A100.
Native Metal / MPS support — runs on M-series Macs without CUDA gymnastics.
You don't need to read a paper before getting your first result — sensible defaults and a quick install.
This page summarizes upstream documentation, release information, and editorially reviewed catalogue fields. It is not presented as a hands-on benchmark. Verify changing requirements at the official project; report stale data through our corrections channel.
A standard local install — download, install dependencies, point at your GPU.
Needs 8 GB minimum — RTX 3060 12GB or 4070 territory.
Source is public — auditable and forkable, no vendor lock.
97 days since the last catalogue refresh — flagged for re-verification.
Three picks across different tradeoffs — so you don't end up with three near-clones of MMAudio.
Multilingual voice cloning in 6 seconds.
The voice-changer that took over Discord.
The benchmark commercial TTS / voice clone API.
MMAudio (Sony AI + research collab) generates synchronized sound effects and ambient audio for silent video clips. Drop in a Wan/Hunyuan/SVD output and get matching footsteps, ambient room tone, splashes — automatically aligned to the frame. Best-in-class for the V2A (video-to-audio) niche.
MIT-licensed; full weights public.
Open-weight text-to-audio — 47-second sound effects and music.
Meta's text-to-music & sound-effect model family.