espnet/espnet

★ 9,980⑂ 2,440

End-to-End Speech Processing Toolkit

About espnet/espnet

espnet/espnet is an open-source project on GitHub, mainly written in Python. End-to-End Speech Processing Toolkit It currently holds 9,980 stars and 2,440 forks with 105 open issues, and was last pushed on 2026-10-08 (repository created 2017-12-13).

Project Overview

Git Homed tracks it on the Audio Trending board and on the AI Audio Trending list.

GitHub Repository Details

Repository espnet/espnet · default branch master · size 1840901 KB · watchers 171 · source: GitHub REST API and repository README

README

https://github.com/espnet/espnet/blob/HEAD/ESPnet

End-to-end speech processing toolkit

PyPI Python Downloads License codecov Hugging Face Discord

Documentation · Installation · Recipes · Model Zoo · Notebooks · Discord

______________________________________________________________________

ESPnet is an end-to-end speech processing toolkit built on PyTorch. It covers speech recognition, text-to-speech, speech translation, speech enhancement, speaker diarization, spoken language understanding, singing voice synthesis, speech language models, and more — with Kaldi-style reproducible recipes from data preparation to evaluation, and hundreds of pretrained models on Hugging Face.

What's new

Earlier releases
  • ESPnet 202610.post1 — the command line grows espnet demo (the OWSM browser demo) and --live (the microphone, transcribed as you speak); one Speech2Text now loads either kind of OWSM checkpoint, with best_path() for CTC decoding without a search; espnet/espnet:inference-cpu-latest and -gpu-latest run a published model with nothing installed; three more demo Spaces (TTS, enhancement, speaker verification).
  • ESPnet 202610 — one-line inference from the command line (pip install espnet && espnet asr audio.wav), two OWSM v4 demos as Hugging Face Spaces, a core install without the training stack (training is espnet[train]), batched beam search, PyTorch 2.11-2.14.
  • ESPnet 202609 — ESPnet3 complete on egs3/librispeech_100 at ESPnet2 parity, CI rebuilt on a prebuilt image (compute per run halved), OpenBEATs pretraining, ten new recipes (ASR, TTS, SER, ST, audio SSL), Python 3.12-3.13.
  • ESPnet 202604 — Docker-based CI, PyTorch 2.9.1 support, FastSpeech2 inference ~1.9x faster at batch 8, new recipes (Kinyarwanda, Emilia, kosp2e).
  • ESPnet 202511 — parallel-processing primitives, refactored inference and evaluation pipeline, expanded SpeechLM support.
  • ESPnet 202509 — Python 3.9-3.13, Debian 12 CI, the LID subsystem completed, multi-optimizer training (HybridOptim / HybridLRS).
  • ESPnet 202506 — ESPnet3 groundwork (data organizer, trainer, model), LID training and task setup, codec1 recipes, USES2 speech enhancement, IPAPack++ S2T recipes.
  • ESPnet 202503 — PyTorch Lightning trainer support, Hugging Face front-end, scaled dot-product attention, ML-SUPERB 2024 recipe.
Full history: Releases.

Install

# Install PyTorch first: https://pytorch.org/get-started/locally/
pip install espnet              # run pretrained models
pip install "espnet[train]"     # also train them (Lightning, TensorBoard, W&B, Dask)
Other installation options
pip install "espnet[all]"                       # training plus every task extra (except sds)
pip install git+https://github.com/espnet/espnet  # latest master
Tested environments (CI status)

|system/pytorch ver.|2.11.0|2.13.0|2.14.0| | :---- | :---: | :---: | :---: | |ubuntu/python3.12/pip|ci on ubuntu|ci on ubuntu|ci on ubuntu| |ubuntu/python3.13/pip|ci on ubuntu|ci on ubuntu|ci on ubuntu| |debian12/python3.12/conda|ci on debian12||| |windows/python3.12/pip|ci on windows||| |macos/python3.12/pip|ci on macos||| |macos/python3.12/conda|ci on macos|||

Each badge is its workflow's aggregate status on master, where the full grid runs. Coverage is not uniform — some suites run on one pytorch only, and a pull request runs less than master does. What each column covers.

pre-commit.ci Code style: black Imports: isort Mergify

Quick start

pip install espnet

Then pick how you want to call it. Same models, same six tasks.

From the terminal

espnet transcribe audio.wav                # detects the language
espnet translate audio.wav --to eng
espnet synthesize "Hello from ESPnet" -o out.wav
espnet enhance noisy.wav -o clean.wav

phonemize and align are there too, and espnet models names the default model of each. Every command but that one takes --model and --device cuda.

From Python

from espnet2.bin.s2t_inference import Speech2Text

OWSM-CTC v4: multilingual ASR, translation and language ID in one model

s2t = Speech2Text.from_pretrained("espnet/owsm_ctc_v4_1B") for start, end, text in s2t.decode_long("audio.wav"): # any length or rate print(text)

Any model from the ESPnet organization, cached after the first download.

From an agent

pip install "espnet[mcp]"
claude mcp add espnet -- espnet-mcp

An MCP server offering the same six tasks as tools, so Claude, Cursor or another agent calls them itself.

Or without installing anything, in a container
docker run --rm -v "$PWD:/data" -v "$HOME/.cache/huggingface:/cache/huggingface" \
    espnet/espnet:inference-cpu-latest asr /data/audio.wav

The second mount keeps the downloaded model between runs. With a GPU, use espnet/espnet:inference-gpu-latest, --gpus all and --device cuda; on a Linux host add --user "$(id -u):$(id -g)". The other images are in docker/.

Train a recipe — every corpus follows the same interface:

cd egs2/librispeech/asr1
./run.sh                              # data → features → training → scoring
./run.sh --stage 11 --stop_stage 13   # or selected stages

New to ESPnet? Start with egs2/mini_an4/asr1 — it runs end to end in minutes.

Supported tasks

| | Task | Template | Highlights | | :-- | :-- | :-- | :-- | | 🗣️ | ASR — speech recognition | asr1, asr2 | Hybrid CTC/attention, Transducer, streaming, Conformer / E-Branchformer, Whisper, SSL front-ends | | 🌏 | S2T — multilingual multitask | s2t1 | OWSM: open Whisper-style models trained on public data | | 🔊 | TTS — text-to-speech | tts1, tts2 | Tacotron 2, FastSpeech 2, VITS, JETS, multi-speaker / multilingual | | 🎤 | SVS — singing voice synthesis | svs1, svs2 | VISinger 1/2, Xiaoice, DiffSinger; merged from Muskits | | 🎧 | SE/SS — enhancement & separation | enh1, enh_asr1 | Unified encoder–separator–decoder, TasNet / DPRNN / beamformers, ASR-integrated | | 🌐 | ST / MT / S2ST — translation | st1, mt1, s2st1 | End-to-end and cascaded speech translation, speech-to-speech translation | | 💬 | SLU — language understanding | slu1 | Intent + transcript multitasking, pretrained ASR/NLP encoders | | 👤 | SPK / LID / DIAR — speaker & language | spk1, lid1, diar1 | Speaker embeddings, verification, language ID, diarization | | 🧠 | SSL — self-supervised learning | ssl1, hubert1 | HuBERT pretraining; S3PRL upstreams as front-ends | | 🤖 | SpeechLM — speech language models | speechlm1 | Unified sequence modeling across speech and text tasks | | 📦 | Codec — neural audio codecs | codec1 | Discrete speech tokens for downstream tasks | | ➕ | More | uasr1, cls1, asvspoof1, lm1, sds1 | Unsupervised ASR (EURO), audio classification, anti-spoofing, LM, spoken dialogue |

Each template ships a corpus-agnostic pipeline; see egs2/README.md for the full list of 200+ corpora recipes.

Why ESPnet

Demos

Four ways to run a published model — a hosted Space, a notebook, the MCP server an agent calls, and the command line. 🟢 is there today, 🚧 is in review, ❌ is not there yet.

| Task | Space | Notebook | MCP | CLI | | :-- | :-- | :-- | :-- | :-- | | ASR — transcription | 🟢 owsm-ctc-v4 | 🟢 asr_demo | 🟢 transcribe | 🟢 espnet transcribe | | ST — speech translation | 🟢 owsm-ctc-v4 | 🟢 st_demo | 🟢 translate | 🟢 espnet translate | | TTS — synthesis | 🟢 ljspeech-vits | 🟢 tts_demo | 🟢 synthesize | 🟢 espnet synthesize | | SE — enhancement | 🟢 universal-se | 🟢 enh_demo | 🟢 enhance | 🟢 espnet enhance | | PR — phone recognition | 🟢 powsm-ctc | 🟢 s2t_pr_demo | 🟢 phonemize | 🟢 espnet phonemize | | ALIGN — forced alignment | 🟢 forced-alignment | 🟢 s2t_align_demo | 🟢 align | 🟢 espnet align |

Every Space is built from a directory in this repository, and doc/front_ends.md is how a seventh task gets all four columns.

pip install "espnet[demo]" adds espnet demo, the same app on localhost; espnet transcribe --live reads the microphone.

Beyond the six: a Space and a notebook for speaker verification, notebooks for neural codecs and spoken dialogue, and egs2/TEMPLATE/sds1 for the full spoken dialogue system, which runs locally.

Every notebook, and whether it still runs

| Demo | | Last run | | :-- | :-- | :-- | | Spoken dialogue — listen, think, speak | Open in Colab | sds_demo | | Speech recognition, in any of 151 languages | Open in Colab | asr_demo | | The words appearing as the audio arrives | Open in Colab | asr_streaming_demo | | Speech translation — the same model, a different task symbol | Open in Colab | st_demo | | Text-to-speech, one voice and then 128 | Open in Colab | tts_demo | | Speech enhancement, and what it did to the signal-to-noise | Open in Colab | enh_demo | | Speaker verification — two recordings, one score | Open in Colab | spk_demo | | Neural codecs — a waveform as a few integers a frame | Open in Colab | codec_demo |

Each runs top to bottom on a CPU and pins the release it was checked against. The second badge is that notebook being executed cell by cell every Sunday, so a red one names the demo that broke rather than leaving you to find out by opening it. More, including the CMU course material: espnet/notebook.

Publish your own. Every ESPnet3 recipe can wrap its trained model in a Gradio app and push it to Hugging Face Spaces — the UI, the Space README.md and requirements.txt are all generated from conf/demo.yaml.

The three stages
cd egs3/librispeech_100/esp2_asr
train=conf/tuning/training_e_branchformer.yaml   # the config the model was trained with
python run.py --stages pack_model  --training_config $train --publication_config conf/publication.yaml  # -> exp/.../model_pack
python run.py --stages pack_demo   --training_config $train --demo_config conf/demo.yaml                # -> demo/
python run.py --stages upload_demo --training_config $train --demo_config conf/demo.yaml                # needs hf auth login

Run the packed app locally with python demo/app.py.

Learn

Contributing

Contributions, questions, and feature requests are all welcome — open an issue or a pull request. First time here? Read the contribution guide.

https://github.com/espnet/espnet/blob/HEAD/Contributors

Details

Full feature list by task

Kaldi-style complete recipe

  • Support numbers of ASR recipes (WSJ, Switchboard, CHiME-4/5, Librispeech, TED, CSJ, AMI, HKUST, Voxforge, REVERB, Gigaspeech, etc.)
  • Support numbers of TTS recipes in a similar manner to the ASR recipe (LJSpeech, LibriTTS, M-AILABS, etc.)
  • Support numbers of ST recipes (Fisher-CallHome Spanish, Libri-trans, IWSLT'18, How2, Must-C, Mboshi-French, etc.)
  • Support numbers of MT recipes (IWSLT'14, IWSLT'16, the above ST recipes etc.)
  • Support numbers of SLU recipes (CATSLU-MAPS, FSC, Grabo, IEMOCAP, JDCINAL, SNIPS, SLURP, SWBD-DA, etc.)
  • Support numbers of SE/SS recipes (DNS-IS2020, LibriMix, SMS-WSJ, VCTK-noisyreverb, WHAM!, WHAMR!, WSJ-2mix, etc.)
  • Support voice conversion recipe (VCC2020 baseline)
  • Support speaker diarization recipe (mini_librispeech, librimix)
  • Support singing voice synthesis recipe (ofuton_p_utagoe_db, opencpop, m4singer, etc.)

ASR: Automatic Speech Recognition

  • State-of-the-art performance in several ASR benchmarks (comparable/superior to hybrid DNN/HMM and CTC)
  • Hybrid CTC/attention based end-to-end ASR
  • Fast/accurate training with CTC/attention multitask training
  • CTC/attention joint decoding to boost monotonic alignment decoding
  • Encoder: VGG-like CNN + BiRNN (LSTM/GRU), sub-sampling BiRNN (LSTM/GRU), Transformer, Conformer, Branchformer, or E-Branchformer
  • Decoder: RNN (LSTM/GRU), Transformer, or S4
  • Attention: Flash Attention, Dot product, location-aware attention, variants of multi-head
  • Incorporate RNNLM/LSTMLM/TransformerLM/N-gram trained only with text data
  • Batch GPU decoding
  • Data augmentation
  • Transducer based end-to-end ASR
  • Architecture:
  • Custom encoder supporting RNNs, Conformer, Branchformer (w/ variants), 1D Conv / TDNN.
  • Decoder w/ parameters shared across blocks supporting RNN, stateless w/ 1D Conv, MEGA, and RWKV.
  • Pre-encoder: VGG2L or Conv2D available.
  • Search algorithms:
  • Greedy search constrained to one emission by timestep.
  • Default beam search algorithm [[Graves, 2012]](https://arxiv.org/abs/1211.3711) without prefix search.
  • Alignment-Length Synchronous decoding [[Saon et al., 2020]](https://ieeexplore.ieee.org/abstract/document/9053040).
  • Time Synchronous Decoding [[Saon et al., 2020]](https://ieeexplore.ieee.org/abstract/document/9053040).
  • N-step Constrained beam search modified from [[Kim et al., 2020]](https://arxiv.org/abs/2002.03577).
  • modified Adaptive Expansion Search based on [[Kim et al., 202

GitHub Stars & Activity

9,980Stars
2,440Forks
105Open issues
PythonLanguage

GitHub Popularity

GitHub stars9,980
Forks2,440
Open issues105
Primary languagePython
LicenseApache-2.0
Stars gained today0
Created2017-12-13
Last pushed2026-10-08

Trending History

Trending statusnot on today's boards

Related GitHub Projects

1

huggingface / transformers

Python★ 167,150⑂ 0
→
2

RVC-Boss / GPT-SoVITS

Python★ 62,631⑂ 0
→
3

debpalash / VoiceStudio

Python★ 57,061⑂ 0
→
4

coqui-ai / TTS

Python★ 46,108⑂ 0
→
5

2noise / ChatTTS

Python★ 39,896⑂ 0
→
6

OpenBMB / VoxCPM

Python★ 38,514⑂ 0
→
7

myshell-ai / OpenVoice

Python★ 37,830⑂ 0
→
8

babysor / MockingBird

Python★ 36,897⑂ 0
→

More Trending Repositories