Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
github.com/SWivid/F5-TTSWhat each tool is actually for, described plainly. No ratings, no scores, no invented benchmarks.
Official code for "F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching"
github.com/SWivid/F5-TTSInstant voice cloning by MIT and MyShell. Audio foundation model.
github.com/myshell-ai/OpenVoiceEasily train a good VC model with voice data <= 10 mins!
github.com/RVC-Project/Retrieval-based-Voice-Conversion-WebUI🐸💬 - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
github.com/coqui-ai/TTSA TTS model capable of generating ultra-realistic dialogue in one pass.
github.com/nari-labs/diaFaster Whisper transcription with CTranslate2
github.com/SYSTRAN/faster-whisperRobust Speech Recognition via Large-Scale Weak Supervision
github.com/openai/whisperPort of OpenAI's Whisper model in C/C++
github.com/ggml-org/whisper.cppWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
github.com/m-bain/whisperX