F5-TTS
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
Written mainly in Python. Released under the MIT licence.
What each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
Written mainly in Python. Released under the MIT licence.
Instant voice cloning by MIT and MyShell. Audio foundation model.
Written mainly in Python. Released under the MIT licence.
Easily train a good VC model with voice data <= 10 mins!
Written mainly in Python. Released under the MIT licence.
🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
Written mainly in Python. Released under the MPL-2.0 licence.
🔊 Text-Prompted Generative Audio Model
Written mainly in Jupyter Notebook. Released under the MIT licence.
A TTS model capable of generating ultra-realistic dialogue in one pass.
Written mainly in Python. Released under the Apache-2.0 licence.
Faster Whisper transcription with CTranslate2
Written mainly in Python. Released under the MIT licence.
SOTA Open Source TTS
Written mainly in Python.
https://hf.co/hexgrad/Kokoro-82M
Written mainly in JavaScript. Released under the Apache-2.0 licence.
Robust Speech Recognition via Large-Scale Weak Supervision
Written mainly in Python. Released under the MIT licence.
Port of OpenAI's Whisper model in C/C++
Written mainly in C++. Released under the MIT licence.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Written mainly in Python. Released under the BSD-2-Clause licence.