denizsafak/abogen
Generate audiobooks from EPUBs, PDFs and text with synchronized captions.
README
abogen 
Abogen is a powerful text-to-speech conversion tool that makes it easy to turn ePub, PDF, text, markdown, or subtitle files into high-quality audio with matching subtitles in seconds. Use it for audiobooks, voiceovers for Instagram, YouTube, TikTok, or any project that needs natural-sounding text-to-speech, using Kokoro-82M.
Demo
https://github.com/user-attachments/assets/094ba3df-7d66-494a-bc31-0e4b41d0b865
This demo was generated in just 5 seconds, producing ∼1 minute of audio with perfectly synced subtitles. To create a similar video, see the demo guide.
How to install? 
Windows
Go to espeak-ng latest release download and run the *.msi file.
OPTION 1: Install using script
1. Download the repository 2. Extract the ZIP file 3. RunWINDOWS_INSTALL.bat by double-clicking it
This method handles everything automatically - installing all dependencies including CUDA in a self-contained environment without requiring a separate Python installation. (You still need to install espeak-ng.)
[!NOTE]
You don't need to install Python separately. The script will install Python automatically.
OPTION 2: Install using uv
First, install uv if you haven't already.# For NVIDIA GPUs (CUDA 12.8) - Recommended
uv tool install --python 3.12 abogen[cuda] --extra-index-url https://download.pytorch.org/whl/cu128 --index-strategy unsafe-best-match
For NVIDIA GPUs (CUDA 12.6) - Older drivers
uv tool install --python 3.12 abogen[cuda126] --extra-index-url https://download.pytorch.org/whl/cu126 --index-strategy unsafe-best-match
For NVIDIA GPUs (CUDA 13.0) - Newer drivers
uv tool install --python 3.12 abogen[cuda130] --extra-index-url https://download.pytorch.org/whl/cu130 --index-strategy unsafe-best-match
For AMD GPUs or without GPU - If you have AMD GPU, you need to use Linux for GPU acceleration, because ROCm is not available on Windows.
uv tool install --python 3.12 abogen
Alternative: Install using pip (click to expand)
# Create a virtual environment (optional)
mkdir abogen && cd abogen
python -m venv venv
venv\Scripts\activate
For NVIDIA GPUs:
We need to use an older version of PyTorch (2.8.0) until this issue is fixed: https://github.com/pytorch/pytorch/issues/166628
pip install torch==2.8.0+cu128 torchvision==0.23.0+cu128 torchaudio==2.8.0 --index-url https://download.pytorch.org/whl/cu128
For AMD GPUs:
Not supported yet, because ROCm is not available on Windows. Use Linux if you have AMD GPU.
Install abogen
pip install abogen
Mac
First, install uv if you haven't already.
# Install espeak-ng
brew install espeak-ng
For Silicon Mac (M1, M2 etc.)
uv tool install --python 3.13 abogen --with "kokoro @ git+https://github.com/hexgrad/kokoro.git,numpy<2"
For Intel Mac
uv tool install --python 3.12 abogen --with "kokoro @ git+https://github.com/hexgrad/kokoro.git,numpy<2"
Alternative: Install using pip (click to expand)
# Install espeak-ng
brew install espeak-ng
Create a virtual environment (recommended)
mkdir abogen && cd abogen
python3 -m venv venv
source venv/bin/activate
Install abogen
pip3 install abogen
For Silicon Mac (M1, M2 etc.)
After installing abogen, we need to install Kokoro's development version which includes MPS support.
pip3 install git+https://github.com/hexgrad/kokoro.git
Linux
First, install uv if you haven't already.
# Install espeak-ng
sudo apt install espeak-ng # Ubuntu/Debian
sudo pacman -S espeak-ng # Arch Linux
sudo dnf install espeak-ng # Fedora
For NVIDIA GPUs or without GPU - No need to include [cuda] in here.
uv tool install --python 3.12 abogen
For AMD GPUs (ROCm 6.4)
uv tool install --python 3.12 abogen[rocm] --extra-index-url https://download.pytorch.org/whl/nightly/rocm6.4 --index-strategy unsafe-best-match
Alternative: Install using pip (click to expand)
# Install espeak-ng
sudo apt install espeak-ng # Ubuntu/Debian
sudo pacman -S espeak-ng # Arch Linux
sudo dnf install espeak-ng # Fedora
Create a virtual environment (recommended)
mkdir abogen && cd abogen
python3 -m venv venv
source venv/bin/activate
Install abogen
pip3 install abogen
For NVIDIA GPUs:
Already supported, no need to install CUDA separately.
For AMD GPUs:
After installing abogen, we need to uninstall the existing torch package
pip3 uninstall torch
pip3 install --pre torch torchvision torchaudio --index-url https://download.pytorch.org/whl/nightly/rocm6.4
See How to fix "CUDA GPU is not available. Using CPU" warning?
See How to fix "WARNING: The script abogen-cli is installed in '/home/username/.local/bin' which is not on PATH" error in Linux?
See How to fix "No matching distribution found" error?
See [How to fix "[WinError 1114] A dynamic link library (DLL) initialization routine failed" error?](#WinError-1114)
Special thanks to @hg000125 for his contribution in #23. AMD GPU support is possible thanks to his work.
Interfaces
Abogen offers two interfaces, but currently they have different feature sets. The Web UI contains newer features that are still being integrated into the desktop application.
| Command | Interface | Features |
|---------|-----------|----------|
| abogen | PyQt6 Desktop GUI | Stable core features |
| abogen-web | Flask Web UI | Core features + Supertonic TTS, LLM Normalization, Audiobookshelf Integration and more! |
Note: The Web UI is under active development. We are working to integrate these new features into the PyQt desktop app. until then, the Web UI provides the most feature-rich experience.
Special thanks to @jeremiahsb for making this possible! I was honestly surprised by his massive contribution (>55,000 lines!) that brought the entire Web UI to life.
🖥️ Desktop Application (PyQt)
How to run?
You can simply run this command to start Abogen Desktop GUI:
abogen
[!TIP]
If you installed Abogen using the Windows installer(WINDOWS_INSTALL.bat), It should have created a shortcut in the same folder, or your desktop. You can run it from there. If you lost the shortcut, Abogen is located inpython_embedded/Scripts/abogen.exe. You can run it from there directly.
How to use?
1) Drag and drop any ePub, PDF, text, markdown, or subtitle file (or use the built-in text editor)
2) Configure the settings:
- Set speech speed
- Select a voice (or create a custom voice using voice mixer)
- Select subtitle generation style (by sentence, word, etc.)
- Select output format
- Select where to save the output
In action
Here’s Abogen in action: in this demo, it processes ∼3,000 characters of text in just 11 seconds and turns it into 3 minutes and 28 seconds of audio, and I have a low-end RTX 2060 Mobile laptop GPU. Your results may vary depending on your hardware.
Configuration
| Options | Description |
|---------|-------------|
| Input Box | Drag and drop ePub, PDF, .TXT, .MD, .SRT, .ASS or .VTT files (or use built-in text editor) |
| Queue options | Add multiple files to a queue and process them in batch, with individual settings for each file. See Queue mode for more details. |
| Speed | Adjust speech rate from 0.1x to 2.0x |
| Select Voice | First letter of the language code (e.g., a for American English, b for British English, etc.), second letter is for m for male and f for female. |
| Voice mixer | Create custom voices by mixing different voice models with a profile system. See Voice Mixer for more details. |
| Voice preview | Listen to the selected voice before processing. |
| Generate subtitles | Disabled, Line, Sentence, Sentence + Comma, Sentence + Highlighting, 1 word, 2 words, 3 words, etc. (Represents the number of words in each subtitle entry) |
| Output voice format | .WAV, .FLAC, .MP3, .OPUS (best compression) and M4B (with chapters) |
| Output subtitle format | Configures the subtitle format as SRT (standard), ASS (wide), ASS (narrow), ASS (centered wide), or ASS (centered narrow). |
| Replace single newlines with spaces | Replaces single newlines with spaces in the text. This is useful for texts that have imaginary line breaks. |
| Save location | Save next to input file, Save to desktop, or Choose output folder |
Special thanks to @brianxiadong for adding markdown support in PR #75
Special thanks to @jborza for chapter support in PR #10
Special thanks to @mleg for adding Line option in subtitle generation in PR #94
| Book handler options | Description |
|---------|-------------|
| Chapter Control | Select specific chapters from ePUBs or markdown files or chapters + pages from PDFs. |
| Save each chapter separately | Save each chapter in e-books as a separate audio file. |
| Create a merged version | Create a single audio file that combines all chapters. (If Save each chapter separately is disabled, this option will be the default behavior.) |
| Save in a project folder with metadata | Save the converted items in a project folder with available metadata files. |
| Menu options | Description |
|---------|-------------|
| Theme | Change the application's theme using System, Light, or Dark options. |
| Configure max words per subtitle | Configures the maximum number of words per subtitle entry. |
| Configure silence between chapters | Configures the duration of silence between chapters (in seconds). |
| Configure max lines in log window | Configures the maximum number of lines to display in the log window. |
| Separate chapters audio format | Configures the audio format for separate chapters as wav, flac, mp3, or opus. |
| Create desktop shortcut | Creates a shortcut on your desktop for easy access. |
| Open config directory | Opens the directory where the configuration file is stored. |
| Open cache directory | Opens the cache directory where converted text files are stored. |
| Clear cache files | Deletes cache files created during the conversion or preview. |
| Use silent gaps between subtitles | Prevents unnecessary audio speed-up by letting speech continue into the silent gaps between subtitle etries. In short, it ignores the end times in subtitle entries and uses the silent space until the beginning of the next subtitle entry. When disabled, it speeds up the audio to fit the exact time interval specified in the subtitle. (for subtitle files). |
| Subtitle speed adjustment method | Choose how to speed up audio when needed: TTS Regeneration (better quality) re-generates the audio at a faster speed, while FFmpeg Time-stretch (better speed) quickly speeds up the generated audio. (for subtitle files). |
| Use spaCy for sentence segmentation | When this option is enabled, Abogen uses spaCy to detect sentence boundaries more accurately, instead of using punctuation marks (like periods, question marks, etc.) to split sentences, which could incorrectly cut off phrases like "Mr." or "Dr.". With spaCy, sentences are divided more accurately. For non-English text, spaCy runs before audio generation to create sentence chunks. For English text, spaCy runs during subtitle generation to improve timing and readability. spaCy is only used when subtitle mode is Sentence or Sentence + Comma. If you prefer the old punctuation splitting method, you can turn this option off. |
| Pre-download models and voices for offline use | Opens a window that displays the available models and voices. Click Download all button to download all required models and voices, allowing you to use Abogen completely offline without any internet connection. |
| Disable Kokoro's internet access | Prevents Kokoro from downloading models or voices from HuggingFace Hub, useful for offline use. |
| Check for updates at startup | Automatically checks for updates when the program starts. |
| Reset to default settings | Resets all settings to their default values. |
Special thanks to @robmckinnon for adding Sentence + Highlighting feature in PR #65
Voice Mixer
With voice mixer, you can create custom voices by mixing different voice models. You can adjust the weight of each voice and save your custom voice as a profile for future use. The voice mixer allows you to create unique and personalized voices.
Special thanks to @jborza for making this possible through his contributions in #5
Queue Mode
Abogen supports queue mode, allowing you to add multiple files to a processing queue. This is useful if you want to convert several files in one batch.
- You can add text files (
.txt) and subtitle files (.srt,.ass,.vtt) directly using the Add files button in the Queue Manager or by dragging and dropping them into the queue list. To add PDF, EPUB, or markdown files, use the input box in the main window and click the Add to Queue button. - Each file in the queue keeps the configuration settings that were active when it was added. Changing the main window configuration afterward does not affect files already in the queue.
- You can enable the Override item settings with current selection option to force all items in the queue to use the configuration currently selected in the main window, overriding their saved settings.
- You can view each file's configuration by hovering over them.
Special thanks to @jborza for adding queue mode in PR #35
---
🌐 Web Application (WebUI)
How to run?
Run this command to start the Web UI:
abogen-web
Then open http://localhost:8808 and drag in your documents. Jobs run in the background worker and the browser updates automatically.
Using the web UI
1. Upload a document (drag & drop or use the upload button).
2. Choose voice, language, speed, subtitle style, and output format.
3. Click Create job. The job immediately appears in the queue.
4. Watch progress and logs update live. Download audio/subtitle assets when complete.
5. Cancel or delete jobs any time. Download logs for troubleshooting.
Multiple jobs can run sequentially; the worker processes them in order.
Container image
You can build a lightweight container image directly from the repository root:
docker build -t abogen .
mkdir -p ~/abogen-data/uploads ~/abogen-data/outputs
docker run --rm \
-p 8808:8808 \
-v ~/abogen-data:/data \
--name abogen \
abogen
Browse to http://localhost:8808. Uploaded source files are stored in /data/uploads and rendered audio/subtitles appear in /data/outputs.
Container environment variables
| Variable | Default | Purpose | |----------|---------|---------| |ABOGEN_HOST | 0.0.0.0 | Bind address for the Flask server |
| ABOGEN_PORT | 8808 | HTTP port |
| ABOGEN_DEBUG | false | Enable Flask debug mode |
| ABOGEN_UPLOAD_ROOT | /data/uploads | Directory where uploaded files are stored |
| ABOGEN_OUTPUT_ROOT | /data/outputs | Directory for generated audio and subtitles (legacy alias of ABOGEN_OUTPUT_DIR) |
| ABOGEN_OUTPUT_DIR | /data/outputs | Container path for rendered audio/subtitles |
| ABOGEN_SETTINGS_DIR | /config | Container path for JSON settings/configuration |
| ABOGEN_TEMP_DIR | /data/cache (Docker) or platform cache dir | Container path for temporary audio working files |
| ABOGEN_UID | 1000 | UID that the container should run as (matches host user) |
| ABOGEN_GID | 1000 | GID that the container should run as (matches host group) |
| ABOGEN_LLM_BASE_URL | "" | OpenAI-compatible endpoint used to seed the Settings → LLM panel |
| ABOGEN_LLM_API_KEY | "" | API key passed to the endpoint above |
| ABOGEN_LLM_MODEL | "" | Default model selected when you refresh the model list |
| ABOGEN_LLM_TIMEOUT | 30 | Timeout (seconds) for server-side LLM requests |
| ABOGEN_LLM_CONTEXT_MODE | sentence | Default prompt context window (sentence, paragraph, document) |
| ABOGEN_LLM_PROMPT | "" | Custom normalization prompt template seeded into the UI |
Set any of these with -e VAR=value when starting the container.
To discover your local UID/GID for matching file permissions inside the container, run:
id -u
id -g
Use those values to populate ABOGEN_UID / ABOGEN_GID in your .env file.
When running via Docker Compose, set ABOGEN_SETTINGS_DIR,
ABOGEN_OUTPUT_DIR, and ABOGEN_TEMP_DIR in your .env file to the host
directories you want mounted into the container. Compose maps them to
/config, /data/outputs, and /data/cache respectively while exporting
those in-container paths to the application. Non-audio caches (e.g., Hugging
Face downloads) stick to the container's internal cache under /tmp/abogen-home/.cache
by default, so only conversion scratch data touches the mounted ABOGEN_TEMP_DIR.
Ensure each host directory exists and is writable by the UID/GID you configure
before starting the stack.
Docker Compose (GPU by default)
The repo includesdocker-compose.yaml, which targets GPU hosts out of the box. Install the NVIDIA Container Toolkit and run:
docker compose up -d --build
Key build/runtime knobs:
TORCH_VERSION– pin a specific PyTorch release that matches your driver (leave blank for the latest on the configured index).TORCH_INDEX_URL– swap out the PyTorch download index when targeting a different CUDA build.ABOGEN_DATA– host path that stores uploads/outputs (defaults to./data).
deploy.resources.reservations.devices block (and the optional runtime: nvidia line) inside the compose file. Compose will then run without requesting a GPU. If you prefer the classic CLI:
docker build -f abogen/Dockerfile -t abogen-gpu .
docker run --rm \
--gpus all \
-p 8808:8808 \
-v ~/abogen-data:/data \
abogen-gpu
LLM-assisted text normalization
Abogen can hand tricky apostrophes and contractions to an OpenAI-compatible large language model. Configure it from Settings → LLM:
1. Enter the base URL for your endpoint (Ollama, OpenAI proxy, etc.) and an API key if required. Use the server root (for Ollama: http://localhost:11434)—Abogen appends /v1/... automatically, but it also accepts inputs that already end in /v1.
2. Click Refresh models to load the catalog, pick a default model, and adjust the timeout or prompt template.
3. Use the preview box to test the prompt, then save the settings. The Normalization panel can synthesize a short audio preview with the current configuration.
When you are running inside Docker or a CI pipeline, seed the form automatically with ABOGEN_LLM_* variables in your .env file. The .env.example file includes sample values for a local Ollama server.
Audiobookshelf integration
Abogen can push finished audiobooks directly into Audiobookshelf. Configure this under Settings → Integrations → Audiobookshelf by providing:
- Base URL – the HTTPS origin (and optional path prefix) where your Audiobookshelf server is reachable, for example
https://abs.example.comorhttps://media.example.com/abs. Do not append/api. - Library ID – the identifier of the target Audiobookshelf library (copy it from the library’s settings page in ABS).
- Folder (name or ID) – the destination folder inside that library. Enter the folder name exactly as it appears in Audiobookshelf (Abogen resolves it to the correct ID automatically), paste the raw
folderId, or click Browse folders to fetch the available folders and populate the field. - API token – a personal access token generated in Audiobookshelf under Account → API tokens.
Reverse proxy checklist (Nginx Proxy Manager)
When Audiobookshelf sits behind Nginx Proxy Manager (NPM), make sure the API paths and headers reach the backend untouched:1. Create a Proxy Host that points to your ABS container or host (default forward port 13378).
2. Under the SSL tab, enable you