1100+ pretrained models — Covers over 1100 languages with ready-to-use speech synthesis models.
Multi-speaker TTS — Generate speech in different speaker voices using speaker embeddings.
Voice cloning — Clone voices using short audio samples with models like XTTS.
Model training pipeline — Train custom TTS models or fine-tune existing models on new datasets.
Multiple vocoder implementations — Choose from MelGAN, ParallelWaveGAN, HiFiGAN and other vocoder architectures.
Battle-tested production library with models like XTTS that handle 16 languages and stream with sub-200ms latency. Includes multiple spectrogram and end-to-end architectures (Tacotron, Glow-TTS, VITS, Bark) plus comprehensive training utilities for dataset curation and model development.
Python 3.8+, PyTorch; GPU recommended for inference