About espnet/espnet
espnet/espnet is an open-source project on GitHub, mainly written in Python. End-to-End Speech Processing Toolkit It currently holds 9,980 stars and 2,440 forks with 105 open issues, and was last pushed on 2026-10-08 (repository created 2017-12-13).
Project Overview
Git Homed tracks it on the Audio Trending board and on the AI Audio Trending list.
GitHub Repository Details
README
End-to-end speech processing toolkit
Documentation · Installation · Recipes · Model Zoo · Notebooks · Discord
______________________________________________________________________
ESPnet is an end-to-end speech processing toolkit built on PyTorch. It covers speech recognition, text-to-speech, speech translation, speech enhancement, speaker diarization, spoken language understanding, singing voice synthesis, speech language models, and more — with Kaldi-style reproducible recipes from data preparation to evaluation, and hundreds of pretrained models on Hugging Face.
What's new
- ESPnet 202610.post2 — the command line has one name a task and two more of them:
espnet phonemizereads the phones with POWSM,espnet alignlines text up with the audio it was said in, andasrandttsbecametranscribeandsynthesize(the old names still work); models published before June 2025 load again, afterinit: chainerwas removed from the toolkit.
Earlier releases
- ESPnet 202610.post1 — the command line grows
espnet demo(the OWSM browser demo) and--live(the microphone, transcribed as you speak); oneSpeech2Textnow loads either kind of OWSM checkpoint, withbest_path()for CTC decoding without a search;espnet/espnet:inference-cpu-latestand-gpu-latestrun a published model with nothing installed; three more demo Spaces (TTS, enhancement, speaker verification). - ESPnet 202610 — one-line inference from the command line (
pip install espnet && espnet asr audio.wav), two OWSM v4 demos as Hugging Face Spaces, a core install without the training stack (training isespnet[train]), batched beam search, PyTorch 2.11-2.14. - ESPnet 202609 — ESPnet3 complete on
egs3/librispeech_100at ESPnet2 parity, CI rebuilt on a prebuilt image (compute per run halved), OpenBEATs pretraining, ten new recipes (ASR, TTS, SER, ST, audio SSL), Python 3.12-3.13. - ESPnet 202604 — Docker-based CI, PyTorch 2.9.1 support, FastSpeech2 inference ~1.9x faster at batch 8, new recipes (Kinyarwanda, Emilia, kosp2e).
- ESPnet 202511 — parallel-processing primitives, refactored inference and evaluation pipeline, expanded SpeechLM support.
- ESPnet 202509 — Python 3.9-3.13, Debian 12 CI, the LID subsystem completed, multi-optimizer training (
HybridOptim/HybridLRS). - ESPnet 202506 — ESPnet3 groundwork (data organizer, trainer, model), LID training and task setup,
codec1recipes, USES2 speech enhancement, IPAPack++ S2T recipes. - ESPnet 202503 — PyTorch Lightning trainer support, Hugging Face front-end, scaled dot-product attention, ML-SUPERB 2024 recipe.
Install
# Install PyTorch first: https://pytorch.org/get-started/locally/
pip install espnet # run pretrained models
pip install "espnet[train]" # also train them (Lightning, TensorBoard, W&B, Dask)
Other installation options
pip install "espnet[all]" # training plus every task extra (except sds)
pip install git+https://github.com/espnet/espnet # latest master
- Full setup (recipes, DNN training, Kaldi-style tooling): see the installation guide.
- Docker: see
docker/and the Docker docs. - Task-specific tools live in
tools/installers. - ESPnet1 is no longer supported — use ESPnet2 (
egs2/) or ESPnet3 (egs3/). See the ESPnet1 notice.
Tested environments (CI status)
|system/pytorch ver.|2.11.0|2.13.0|2.14.0|
| :---- | :---: | :---: | :---: |
|ubuntu/python3.12/pip||
|
|
|ubuntu/python3.13/pip|
|
|
|
|debian12/python3.12/conda|
|||
|windows/python3.12/pip|
|||
|macos/python3.12/pip|
|||
|macos/python3.12/conda|
|||
Each badge is its workflow's aggregate status on master, where the full grid runs. Coverage is not uniform — some suites run on one pytorch only, and a pull request runs less than master does. What each column covers.
Quick start
pip install espnet
Then pick how you want to call it. Same models, same six tasks.
From the terminal
espnet transcribe audio.wav # detects the language
espnet translate audio.wav --to eng
espnet synthesize "Hello from ESPnet" -o out.wav
espnet enhance noisy.wav -o clean.wav
phonemize and align are there too, and espnet models names the default model of each. Every command but that one takes --model and --device cuda.
From Python
from espnet2.bin.s2t_inference import Speech2Text
OWSM-CTC v4: multilingual ASR, translation and language ID in one model
s2t = Speech2Text.from_pretrained("espnet/owsm_ctc_v4_1B")
for start, end, text in s2t.decode_long("audio.wav"): # any length or rate
print(text)
Any model from the ESPnet organization, cached after the first download.
From an agent
pip install "espnet[mcp]"
claude mcp add espnet -- espnet-mcp
An MCP server offering the same six tasks as tools, so Claude, Cursor or another agent calls them itself.
Or without installing anything, in a container
docker run --rm -v "$PWD:/data" -v "$HOME/.cache/huggingface:/cache/huggingface" \
espnet/espnet:inference-cpu-latest asr /data/audio.wav
The second mount keeps the downloaded model between runs. With a GPU, use espnet/espnet:inference-gpu-latest, --gpus all and --device cuda; on a Linux host add --user "$(id -u):$(id -g)". The other images are in docker/.
Train a recipe — every corpus follows the same interface:
cd egs2/librispeech/asr1
./run.sh # data → features → training → scoring
./run.sh --stage 11 --stop_stage 13 # or selected stages
New to ESPnet? Start with egs2/mini_an4/asr1 — it runs end to end in minutes.
Supported tasks
| | Task | Template | Highlights |
| :-- | :-- | :-- | :-- |
| 🗣️ | ASR — speech recognition | asr1, asr2 | Hybrid CTC/attention, Transducer, streaming, Conformer / E-Branchformer, Whisper, SSL front-ends |
| 🌏 | S2T — multilingual multitask | s2t1 | OWSM: open Whisper-style models trained on public data |
| 🔊 | TTS — text-to-speech | tts1, tts2 | Tacotron 2, FastSpeech 2, VITS, JETS, multi-speaker / multilingual |
| 🎤 | SVS — singing voice synthesis | svs1, svs2 | VISinger 1/2, Xiaoice, DiffSinger; merged from Muskits |
| 🎧 | SE/SS — enhancement & separation | enh1, enh_asr1 | Unified encoder–separator–decoder, TasNet / DPRNN / beamformers, ASR-integrated |
| 🌐 | ST / MT / S2ST — translation | st1, mt1, s2st1 | End-to-end and cascaded speech translation, speech-to-speech translation |
| 💬 | SLU — language understanding | slu1 | Intent + transcript multitasking, pretrained ASR/NLP encoders |
| 👤 | SPK / LID / DIAR — speaker & language | spk1, lid1, diar1 | Speaker embeddings, verification, language ID, diarization |
| 🧠 | SSL — self-supervised learning | ssl1, hubert1 | HuBERT pretraining; S3PRL upstreams as front-ends |
| 🤖 | SpeechLM — speech language models | speechlm1 | Unified sequence modeling across speech and text tasks |
| 📦 | Codec — neural audio codecs | codec1 | Discrete speech tokens for downstream tasks |
| ➕ | More | uasr1, cls1, asvspoof1, lm1, sds1 | Unsupervised ASR (EURO), audio classification, anti-spoofing, LM, spoken dialogue |
Each template ships a corpus-agnostic pipeline; see egs2/README.md for the full list of 200+ corpora recipes.
Why ESPnet
- Reproducible — one
run.shper corpus, from download to scoring, with published results. - Unified — the same recipe structure, config format, and trainer across every task above.
- Scalable — DDP, multi-node training, Slurm/MPI, DeepSpeed, sharded training, on-the-fly feature extraction.
- Open — hundreds of pretrained models and demos on Hugging Face, plus W&B and TensorBoard logging.
Demos
Four ways to run a published model — a hosted Space, a notebook, the MCP server an agent calls, and the command line. 🟢 is there today, 🚧 is in review, ❌ is not there yet.
| Task | Space | Notebook | MCP | CLI |
| :-- | :-- | :-- | :-- | :-- |
| ASR — transcription | 🟢 owsm-ctc-v4 | 🟢 asr_demo | 🟢 transcribe | 🟢 espnet transcribe |
| ST — speech translation | 🟢 owsm-ctc-v4 | 🟢 st_demo | 🟢 translate | 🟢 espnet translate |
| TTS — synthesis | 🟢 ljspeech-vits | 🟢 tts_demo | 🟢 synthesize | 🟢 espnet synthesize |
| SE — enhancement | 🟢 universal-se | 🟢 enh_demo | 🟢 enhance | 🟢 espnet enhance |
| PR — phone recognition | 🟢 powsm-ctc | 🟢 s2t_pr_demo | 🟢 phonemize | 🟢 espnet phonemize |
| ALIGN — forced alignment | 🟢 forced-alignment | 🟢 s2t_align_demo | 🟢 align | 🟢 espnet align |
Every Space is built from a directory in this repository, and doc/front_ends.md is how a seventh task gets all four columns.
pip install "espnet[demo]" adds espnet demo, the same app on localhost; espnet transcribe --live reads the microphone.
Beyond the six: a Space and a notebook for speaker verification, notebooks for neural codecs and spoken dialogue, and egs2/TEMPLATE/sds1 for the full spoken dialogue system, which runs locally.
Every notebook, and whether it still runs
| Demo | | Last run |
| :-- | :-- | :-- |
| Spoken dialogue — listen, think, speak | |
|
| Speech recognition, in any of 151 languages |
|
|
| The words appearing as the audio arrives |
|
|
| Speech translation — the same model, a different task symbol |
|
|
| Text-to-speech, one voice and then 128 |
|
|
| Speech enhancement, and what it did to the signal-to-noise |
|
|
| Speaker verification — two recordings, one score |
|
|
| Neural codecs — a waveform as a few integers a frame |
|
|
Each runs top to bottom on a CPU and pins the release it was checked against. The second badge is that notebook being executed cell by cell every Sunday, so a red one names the demo that broke rather than leaving you to find out by opening it. More, including the CMU course material: espnet/notebook.
Publish your own. Every ESPnet3 recipe can wrap its trained model in a
Gradio app and push it to Hugging Face Spaces — the UI,
the Space README.md and requirements.txt are all generated from
conf/demo.yaml.
The three stages
cd egs3/librispeech_100/esp2_asr
train=conf/tuning/training_e_branchformer.yaml # the config the model was trained with
python run.py --stages pack_model --training_config $train --publication_config conf/publication.yaml # -> exp/.../model_pack
python run.py --stages pack_demo --training_config $train --demo_config conf/demo.yaml # -> demo/
python run.py --stages upload_demo --training_config $train --demo_config conf/demo.yaml # needs hf auth login
Run the packed app locally with python demo/app.py.
Learn
- Documentation · ESPnet2 tutorial
- Course tutorials at CMU: usage · adding new models/tasks (materials)
- Interspeech 2019 tutorial
Contributing
Contributions, questions, and feature requests are all welcome — open an issue or a pull request. First time here? Read the contribution guide.
Details
Full feature list by task
Kaldi-style complete recipe
- Support numbers of
ASRrecipes (WSJ, Switchboard, CHiME-4/5, Librispeech, TED, CSJ, AMI, HKUST, Voxforge, REVERB, Gigaspeech, etc.) - Support numbers of
TTSrecipes in a similar manner to the ASR recipe (LJSpeech, LibriTTS, M-AILABS, etc.) - Support numbers of
STrecipes (Fisher-CallHome Spanish, Libri-trans, IWSLT'18, How2, Must-C, Mboshi-French, etc.) - Support numbers of
MTrecipes (IWSLT'14, IWSLT'16, the above ST recipes etc.) - Support numbers of
SLUrecipes (CATSLU-MAPS, FSC, Grabo, IEMOCAP, JDCINAL, SNIPS, SLURP, SWBD-DA, etc.) - Support numbers of
SE/SSrecipes (DNS-IS2020, LibriMix, SMS-WSJ, VCTK-noisyreverb, WHAM!, WHAMR!, WSJ-2mix, etc.) - Support voice conversion recipe (VCC2020 baseline)
- Support speaker diarization recipe (mini_librispeech, librimix)
- Support singing voice synthesis recipe (ofuton_p_utagoe_db, opencpop, m4singer, etc.)
ASR: Automatic Speech Recognition
- State-of-the-art performance in several ASR benchmarks (comparable/superior to hybrid DNN/HMM and CTC)
- Hybrid CTC/attention based end-to-end ASR
- Fast/accurate training with CTC/attention multitask training
- CTC/attention joint decoding to boost monotonic alignment decoding
- Encoder: VGG-like CNN + BiRNN (LSTM/GRU), sub-sampling BiRNN (LSTM/GRU), Transformer, Conformer, Branchformer, or E-Branchformer
- Decoder: RNN (LSTM/GRU), Transformer, or S4
- Attention: Flash Attention, Dot product, location-aware attention, variants of multi-head
- Incorporate RNNLM/LSTMLM/TransformerLM/N-gram trained only with text data
- Batch GPU decoding
- Data augmentation
- Transducer based end-to-end ASR
- Architecture:
- Custom encoder supporting RNNs, Conformer, Branchformer (w/ variants), 1D Conv / TDNN.
- Decoder w/ parameters shared across blocks supporting RNN, stateless w/ 1D Conv, MEGA, and RWKV.
- Pre-encoder: VGG2L or Conv2D available.
- Search algorithms:
- Greedy search constrained to one emission by timestep.
- Default beam search algorithm [[Graves, 2012]](https://arxiv.org/abs/1211.3711) without prefix search.
- Alignment-Length Synchronous decoding [[Saon et al., 2020]](https://ieeexplore.ieee.org/abstract/document/9053040).
- Time Synchronous Decoding [[Saon et al., 2020]](https://ieeexplore.ieee.org/abstract/document/9053040).
- N-step Constrained beam search modified from [[Kim et al., 2020]](https://arxiv.org/abs/2002.03577).
- modified Adaptive Expansion Search based on [[Kim et al., 202