abus-aikorea/voice-pro

★ 12,817⑂ 0

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

12,817Star
0Fork
0Watch
0Issue
PythonLanguage
-License
Created · last push · repository size 0 KB · default branch -

README

Voice-Pro

The best AI speech recognition, translation, and multilingual dubbing solution 🚀

https://github.com/abus-aikorea/voice-pro/blob/HEAD/Ask DeepWiki.com https://github.com/abus-aikorea/voice-pro/blob/HEAD/youtube https://github.com/abus-aikorea/voice-pro/blob/HEAD/Buy Me a Coffee https://github.com/abus-aikorea/voice-pro/blob/HEAD/release https://github.com/abus-aikorea/voice-pro/blob/HEAD/GitHub Repo stars

https://github.com/abus-aikorea/voice-pro/blob/HEAD/Dubbing Studio


🎙️ An AI-powered web application for speech recognition, translation, and dubbing

https://github.com/abus-aikorea/voice-pro/blob/HEAD/South Korea Flag 한국어 https://github.com/abus-aikorea/voice-pro/blob/HEAD/United Kingdom Flag English https://github.com/abus-aikorea/voice-pro/blob/HEAD/China Flag 中文简体 https://github.com/abus-aikorea/voice-pro/blob/HEAD/Taiwan Flag 中文繁體 https://github.com/abus-aikorea/voice-pro/blob/HEAD/Japan Flag 日本語 https://github.com/abus-aikorea/voice-pro/blob/HEAD/Germany Flag Deutsch https://github.com/abus-aikorea/voice-pro/blob/HEAD/Spain Flag Español https://github.com/abus-aikorea/voice-pro/blob/HEAD/Portugal Flag Português

Voice-Pro is a state-of-the-art web app that transforms multimedia content creation. It integrates YouTube video downloading, voice separation, speech recognition, translation, and text-to-speech into a single, powerful tool for creators, researchers, and multilingual professionals.

A robust alternative to ElevenLabs, Voice-Pro empowers podcasters, developers, and creators with advanced voice solutions.

⚠️ Please Note

📰 News & History

version 4.0
  • Migrated the installer from Miniconda/pip to uv — dramatically faster, fully reproducible installs from a committed uv.lock. Everything stays inside installer_files/ (uv, Python, packages).
  • 🐍 Upgraded runtime: Python 3.12, Torch 2.8.0+cu128 (RTX 50-series supported), Gradio 6.20.
  • 🎙️ Latest ASR stack: faster-whisper 1.2.1 (large-v3-turbo, distil-large-v3.5), openai-whisper 20250625, whisper-timestamped 1.15.9. whisperX was removed (its dependency pins blocked the Gradio 6 upgrade; existing configs fall back to faster-whisper).
  • 🗣️ Latest TTS stack: F5-TTS 1.1.21, kokoro 0.9.4, edge-tts 7.x, and re-vendored CosyVoice (upstream main).
  • 🇰🇷 New optional TTS model: Fun-CosyVoice3-0.5B — 9 languages including Korean, selectable in the CosyVoice tab (downloads from the official HF repo on first use).
  • 🧹 CUDA Toolkit and Visual Studio Build Tools are no longer required — all dependencies ship prebuilt wheels, and PyTorch bundles the CUDA runtime.
  • 🛡️ Friendly to restricted / corporate PCs: no administrator rights needed — start.bat auto-downloads a portable ffmpeg if it is not installed, Whisper model downloads self-heal after interrupted/corrupted transfers, and translation automatically retries with backoff when the network rate-limits the free Google endpoint (failed lines are reported, originals kept).
  • 🚨 Errors are now visible in the WebUI: every failure shows a red error toast that stays on screen until you close it (previously a 10-second warning that was easy to miss), with actionable messages for common causes (missing ffmpeg, no media registered, etc.).
  • 🖥️ UI: migrated to Gradio 6 (full-width layout for all tabs, subtitle tracks shown directly in the video players).
  • 🧽 uninstall.bat no longer requires administrator rights and no longer force-reboots; uninstall.bat silent runs unattended.
version 3.2
  • We have been focusing on WeConnect development for the past few months and have not been able to manage Voice-Pro at all.
  • We have decided to open source all Voice-Pro code.
  • Voice-Pro is completely free and supports Windows, Mac, Linux.
  • WeConnect is an application for global cultural exchange.
  • Connect with people from all over the world for meaningful cultural exchanges, language learning, and international friendships.

https://github.com/abus-aikorea/voice-pro/blob/HEAD/ScreenShot 0 https://github.com/abus-aikorea/voice-pro/blob/HEAD/ScreenShot 1 https://github.com/abus-aikorea/voice-pro/blob/HEAD/ScreenShot 2 https://github.com/abus-aikorea/voice-pro/blob/HEAD/ScreenShot 3 https://github.com/abus-aikorea/voice-pro/blob/HEAD/ScreenShot 4

version 3.1
version 3.0
  • 🔥 Removed the AI Cover feature.
  • 🚀 Added support for m-bain/whisperX.
version 2.0
  • 🐍 Built with Python 3.10.15, Torch 2.5.1+cu124, and Gradio 5.14.0.
  • 🆓 Free trial supports media up to 60 seconds in length.
  • 🔥 Added the AI Cover feature.
  • 🎤 Introduced support for CosyVoice and kokoro.
  • ⏳ Initial run downloads CozyVoice2-0.5B (9GB), which may take over an hour depending on network speed.
  • 🎧 Voice samples for cloning will be continuously updated.
  • 📝 Added spaCy for natural sentence-by-sentence translation and TTS.
  • ☁️ Subscription version includes Microsoft Azure Translator and TTS.
  • 🏪 Subscription offers unlimited usage (no 60-second limit) during the subscription period, available via Shopify.

🎥 YouTube Showcase

https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 1
Demo for Voice-Pro (v2.0)
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 2
F5-TTS: Voice Cloning
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 3
Live Transcription & Translation
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 4
Multi-Lingual Voice Cloning: Korean - German
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 5
Multi-Lingual Voice Cloning: English - Korean
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 6
Multi-Lingual Voice Cloning: Korean - Japanese
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 7
NVIDIA RTX Video Super-Resolution
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 8
AI Karaoke
https://github.com/abus-aikorea/voice-pro/blob/HEAD/Demo Video 5
Multi-Lingual Voice Cloning: English - Korean

⭐ Key Features

1. Dubbing Studio

2. Speech Technologies

3. Real-Time Translation

🤖 WebUI

Dubbing Studio Tab

https://github.com/abus-aikorea/voice-pro/blob/HEAD/Multilingual Voice Conversion and Subtitle Generation Web UI Interface

Whisper Caption Tab

Translate Tab

https://github.com/abus-aikorea/voice-pro/blob/HEAD/WebUI for Real-Time Speech Recognition and Translation

Speech Generation Tab

https://github.com/abus-aikorea/voice-pro/blob/HEAD/Podcast Production WebUI Using Voice-Cloning Technology

🎤✨ Reference Voice

English

Andrew Bustamante

Andrew Huberman

Avi Loeb

Ben Shapiro

Brett Johnson

Brian Keating

Coffeezilla

Dan Carlin

David Buss

David Fravor

David Kipping

Dennis Whyte

Donald Hoffman

Donald Trump

Douglas Murray

Duncan Trussell

Elon Musk

Garry Nolan

Jack Barsky

James Sexton

Jeff Bezos

Joe Rogan

John Mearsheimer

Jordan Peterson

Kanye 'Ye' West

Mark Zuckerberg

Michael Levin

Michael Saylor

Michio Kaku

MrBeast

Nick Lane

Paul Rosolie

Ryan Graves

Sam Altman

Sam Harris

Stephen Wolfram

Tucker Carlson

Vitalik Buterin

Yuval Harari
Chinese

迪丽热巴 (Dílì Rèbā)

蔡依林 (Cài Yīlín)

吴亦凡 (Wú Yìfán)

李易峰 (Lǐ Yìfēng)

杨幂 (Yáng Mì)

赵丽颖 (Zhào Lìyǐng)
Korean

BTS 진 (Jin)

BTS RM

IU (아이유)

이병헌

이정재

유재석
Japanese

綾瀬はるか (Ayase Haruka)

💻 System Requirements

More Audio Trending projects

1

huggingface / transformers

Python★ 166,108⑂ 0
2

harry0703 / MoneyPrinterTurbo

Python★ 123,776⑂ 0
3

unslothai / unsloth

Python★ 76,181⑂ 0
4

RVC-Boss / GPT-SoVITS

Python★ 61,798⑂ 0
5

calesthio / OpenMontage

Python★ 59,205⑂ 0
6

ggml-org / whisper.cpp

C++★ 53,674⑂ 0