VoiceStudio: Local Voice Cloning, Dubbing, and TTS Suite
VoiceStudio is an open-source application (AGPL-3.0 license) that enables voice cloning, video dubbing, dictation, and long-form audio generation entirely offline, without requiring an account, API key, or subscription. It integrates 16 text-to-speech (TTS) engines, 11 automatic speech recognition (ASR) engines, and supports 646 languages, all runnable on macOS, Windows, Linux, and via Docker containers. Designed for users who prioritize privacy and full control over their data, it serves as a local alternative to cloud-based voice services.
The suite is suitable for content creators, developers, and professionals needing to generate audio or dub videos without relying on external services. It operates on local hardware, including NVIDIA GPUs (CUDA), Apple Silicon (MPS/MLX), and AMD (ROCm on Linux), with CPU-only support as well. Minimum requirements include 8 GB of RAM and 10 GB of disk space, though 16 GB and 20 GB are recommended for optimal performance.
Notable features include zero-shot voice cloning, integration with models like CosyVoice 3, GPT-SoVITS, and OmniVoice, and a user-friendly interface for switching between engines. Additional tools include vocal isolation, speaker diarization, and an MCP server for custom integrations. While currently in active beta, the latest stable release is recommended for production use.