Trending open-source projects ยท updated daily

Audio Trending

๐ŸŽต Text-to-speech, speech recognition, music generation and audio processing โ€” sorted by stars, 24 projects.

Source: GitHub topic pages ยท 24 projects indexed

1

huggingface / transformers

๐Ÿค— Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.

Pythonโ˜… 166,096โ‘‚ 0
โ†’
2

harry0703 / MoneyPrinterTurbo

ๅˆฉ็”จ AI ๅคงๆจกๅž‹ๅ’Œ่‡ชๅŠจๅŒ–ๅทฅไฝœๆต๏ผŒๆ นๆฎไธป้ข˜ๆˆ–ๅ…ณ้”ฎ่ฏไธ€้”ฎ็”Ÿๆˆ้ซ˜ๆธ…็Ÿญ่ง†้ข‘ใ€‚Generate HD short videos from a topic or keyword with an automated AI workflow.

Pythonโ˜… 123,769โ‘‚ 0
โ†’
3

unslothai / unsloth

Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, DeepSeek-V4, MiniMax-H3, Gemma 4, FLUX and more.

Pythonโ˜… 76,179โ‘‚ 0
โ†’
4

RVC-Boss / GPT-SoVITS

1 min voice data can also be used to train a good TTS model! (few shot voice cloning)

Pythonโ˜… 61,797โ‘‚ 0
โ†’
5

calesthio / OpenMontage

World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video p

Pythonโ˜… 59,198โ‘‚ 0
โ†’
6

ggml-org / whisper.cpp

Port of OpenAI's Whisper model in C/C++

C++โ˜… 53,675โ‘‚ 0
โ†’
7

coqui-ai / TTS

๐Ÿธ๐Ÿ’ฌ - a deep learning toolkit for Text-to-Speech, battle-tested in research and production

Pythonโ˜… 46,012โ‘‚ 0
โ†’
8

2noise / ChatTTS

A generative speech model for daily dialogue.

Pythonโ˜… 39,843โ‘‚ 0
โ†’
9

myshell-ai / OpenVoice

Instant voice cloning by MIT and MyShell. Audio foundation model.

Pythonโ˜… 37,531โ‘‚ 0
โ†’
10

OpenBMB / VoxCPM

VoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning

Pythonโ˜… 37,478โ‘‚ 0
โ†’
11

babysor / MockingBird

๐Ÿš€Clone a voice in 5 seconds to generate arbitrary speech in real-time

Pythonโ˜… 36,906โ‘‚ 0
โ†’
12

debpalash / VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative โ€” voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Pythonโ˜… 29,769โ‘‚ 0
โ†’
13

mozilla / DeepSpeech

DeepSpeech is an open source embedded (offline, on-device) speech-to-text engine which can run in real time on devices ranging from a Raspberry Pi 4 to high power GPU servers.

C++โ˜… 26,774โ‘‚ 0
โ†’
14

SYSTRAN / faster-whisper

Faster Whisper transcription with CTranslate2

Pythonโ˜… 25,404โ‘‚ 0
โ†’
15

m-bain / whisperX

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

Pythonโ˜… 24,038โ‘‚ 0
โ†’
16

index-tts / index-tts

An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Pythonโ˜… 23,975โ‘‚ 0
โ†’
17

QwenAudio / CosyVoice

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

Pythonโ˜… 23,622โ‘‚ 0
โ†’
18

modelscope / FunASR

Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.

Pythonโ˜… 20,334โ‘‚ 0
โ†’
19

nari-labs / dia

A TTS model capable of generating ultra-realistic dialogue in one pass.

Pythonโ˜… 19,399โ‘‚ 0
โ†’
20

jianchang512 / pyvideotrans

Translate the video from one language to another and embed dubbing & subtitles.

Pythonโ˜… 19,017โ‘‚ 0
โ†’
21

leon-ai / leon

๐Ÿง  Leon is your open-source personal assistant.

TypeScriptโ˜… 17,512โ‘‚ 0
โ†’
22

kaldi-asr / kaldi

kaldi-asr/kaldi is the official location of the Kaldi project.

Shellโ˜… 15,484โ‘‚ 0
โ†’
23

alphacep / vosk-api

Offline speech recognition API for Android, iOS, Raspberry Pi and servers with Python, Java, C# and Node

Jupyter Notebookโ˜… 15,127โ‘‚ 0
โ†’
24

NVIDIA / DeepLearningExamples

State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.

Jupyter Notebookโ˜… 14,843โ‘‚ 0
โ†’