huggingface / transformers
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
๐ต Text-to-speech, speech recognition, music generation and audio processing โ sorted by stars, 24 projects.
Source: GitHub topic pages ยท 24 projects indexed
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
ๅฉ็จ AI ๅคงๆจกๅๅ่ชๅจๅๅทฅไฝๆต๏ผๆ นๆฎไธป้ขๆๅ ณ้ฎ่ฏไธ้ฎ็ๆ้ซๆธ ็ญ่ง้ขใGenerate HD short videos from a topic or keyword with an automated AI workflow.
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.
1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video p
Port of OpenAI's Whisper model in C/C++
๐ธ๐ฌ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production
A generative speech model for daily dialogue.
Instant voice cloning by MIT and MyShell. Audio foundation model.
VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
๐Clone a voice in 5 seconds to generate arbitrary speech in real-time
VoiceStudio is the open-source, fully-local ElevenLabs alternative โ voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.
DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.
Faster Whisper transcription with CTranslate2
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
A TTS model capable of generating ultra-realistic dialogue in one pass.
Translate the video from one language to another and embed dubbing & subtitles.
๐ง Leon is your open-source personal assistant.
kaldi-asr/kaldi is the official location of the Kaldi project.
Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.