Blaizzy/mlx-audio
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
README
MLX-Audio
The best audio processing library built on Apple's MLX framework, providing fast and efficient text-to-speech (TTS), speech-to-text (STT), speech-to-speech (STS), music generation, and more on Apple Silicon.
Table of Contents
- Features
- Installation
- Quick Start
- Supported Models
- Model Examples
- Web Interface \& API Server
- Quantization
- Swift
- Requirements
- License
- Citation
- Acknowledgements
Features
- Fast inference optimized for Apple Silicon (M series chips)
- Multiple model architectures for TTS, STT, STS, and music generation
- Multilingual support across models
- Voice customization and cloning capabilities
- Adjustable speech speed control
- Interactive web interface with 3D audio visualization
- OpenAI-compatible REST API
- Quantization support (3-bit, 4-bit, 6-bit, 8-bit, and more) for optimized performance
- Swift package for iOS/macOS integration
Installation
Using pip
pip install mlx-audio
Using uv to install only the command line tools
Latest release from pypi:uv tool install --force mlx-audio --prerelease=allow
Latest code from github:
uv tool install --force git+https://github.com/Blaizzy/mlx-audio.git --prerelease=allow
For development or web interface:
git clone https://github.com/Blaizzy/mlx-audio.git
cd mlx-audio
pip install -e ".[dev, server]"
Quick Start
Command Line
# Basic TTS generation
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello, world!' --voice Vivian
With a different voice and language hint
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Welcome to MLX-Audio!' --voice Ryan --lang_code English
Play audio immediately
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --play
Save to a specific directory
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --output_path ./my_audio
Stream audio during generation
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream
Stream audio during generation and save it to disk
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream --save
Join multiple generated segments into one file
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text $'Hello!\nHow are you?' --voice Vivian --join_audio
By default, when generation yields multiple segments, mlx-audio saves numbered files such as audio_000.wav and audio_001.wav. Use --join_audio to save one combined file instead. When using --stream, add --save to write the streamed audio to disk.
Python API
from mlx_audio.tts.utils import load_model
Load model
model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit")
Generate speech
for result in model.generate(
"Hello from MLX-Audio!",
voice="Vivian",
lang_code="English",
):
print(f"Generated {result.audio.shape[0]} samples")
# result.audio contains the waveform as mx.array
Supported Models
Text-to-Speech (TTS)
| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Kokoro | Fast, high-quality multilingual TTS | EN, JA, ZH, FR, ES, IT, PT, HI | bf16, 8bit, 6bit, 4bit | | KittenTTS | Compact KittenTTS 0.8 models for edge-friendly TTS | EN | nano, micro, mini, collection | | Qwen3-TTS | Alibaba's multilingual TTS with voice design | ZH, EN, JA, KO, + more | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16 | | Higgs Audio v3 | 4B conversational TTS with voice cloning and inline control tokens | 100 languages | bosonai/higgs-audio-v3-tts-4b | | OmniVoice | Zero-shot multilingual TTS with voice cloning, batch generation, and nonverbal tags | 646+ languages | mlx-community/OmniVoice-bf16 | | CSM / MisoTTS | Sesame-style conversational speech models with voice cloning | EN | mlx-community/csm-1b, MisoTTS bf16, MisoTTS 8bit | | Dia | Dialogue-focused TTS | EN | mlx-community/Dia-1.6B-fp16 | | OuteTTS | Efficient TTS model | EN | mlx-community/OuteTTS-1.0-0.6B-fp16 | | Spark | SparkTTS model | EN, ZH | mlx-community/Spark-TTS-0.5B-bf16 | | Chatterbox | Expressive multilingual TTS (v2/v3) | 23 languages | v3, v2 | | Soprano | High-quality TTS | EN | mlx-community/Soprano-1.1-80M-bf16 | | Ming Omni TTS (BailingMM) | Multimodal generation with voice cloning, style control, and speech/music/event generation | EN, ZH | mlx-community/Ming-omni-tts-16.8B-A3B-bf16 | | Ming Omni TTS (Dense) | Lightweight dense Ming Omni variant for voice cloning and style control | EN, ZH | mlx-community/Ming-omni-tts-0.5B-bf16 | | KugelAudio | SOTA 7B AR+Diffusion TTS for European languages | EN, DE, FR, ES, IT, PT, NL, PL, RU, UK, + 14 more | kugelaudio/kugelaudio-0-open | | Voxtral TTS | Mistral's 4B multilingual TTS (20 voices, 9 languages) | EN, FR, ES, DE, IT, PT, NL, AR, HI | mlx-community/Voxtral-4B-TTS-2603-mlx-bf16 | | rumik-oss 1 | 3B expressive multilingual Indic TTS with 22-language support, description-conditioned delivery and inline vocalizations | 22 Indic languages + EN | rumik-ai/rumik-oss-1, 8bit, 4bit | | VoxCPM2 | 2B tokenizer-free TTS with 48kHz output, voice design, voice cloning, and continuation | 30 languages | bf16, 8bit, 4bit | | LongCat-AudioDiT | SOTA diffusion TTS in waveform latent space with voice cloning | ZH, EN | mlx-community/LongCat-AudioDiT-1B-bf16 | | MeloTTS | Lightweight VITS2-based TTS with streaming | EN (more coming) | mlx-community/MeloTTS-English-MLX | | MOSS-TTS | 8B delay-pattern and local-transformer multilingual TTS with voice cloning | 31 languages | OpenMOSS-Team/MOSS-TTS-v1.5, OpenMOSS-Team/MOSS-TTS, OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5, OpenMOSS-Team/MOSS-TTS-Local-Transformer | | MOSS-TTS-Nano | Tiny multilingual voice-cloning TTS | 20 languages | mlx-community/MOSS-TTS-Nano-100M | | Higgs Audio v2 | 3B Llama-backed TTS with real-time voice cloning | EN, ZH, KO, DE, ES | bf16 (upstream), q8, q6 |
Speech-to-Text (STT)
| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Whisper | OpenAI's robust STT model | 99+ languages | mlx-community/whisper-large-v3-turbo-asr-fp16 | | Distil-Whisper | Distilled fast Whisper variants | EN | distil-whisper/distil-large-v3 | | Qwen3-ASR | Alibaba's multilingual ASR | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ASR-1.7B-8bit | | Mega-ASR | Routed Qwen3-ASR with automatic clean/base vs degraded/LoRA switching | EN (fixtures), multilingual Qwen3-ASR backbone | README | | Qwen3-ForcedAligner | Word-level audio alignment | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ForcedAligner-0.6B-8bit | | MOSS-Transcribe-Diarize | Timestamped transcription with speaker labels | Multiple major languages | https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize | | Parakeet | NVIDIA's accurate STT | EN (v2), 25 EU languages (v3) | mlx-community/parakeet-tdt-0.6b-v3 | | Nemotron 3.5 ASR (streaming) | NVIDIA's cache-aware streaming FastConformer-RNNT with language-ID prompting | 40 language-locales | mlx-community/nemotron-3.5-asr-streaming-0.6b · README | | Voxtral | Mistral's speech model | Multiple | mlx-community/Voxtral-Mini-3B-2507-bf16 | | Voxtral Realtime | Mistral's 4B streaming STT | Multiple | 4bit, fp16 | | VibeVoice-ASR | Microsoft's 3B/9B ASR with diarization, timestamps, hotwords, and native chunk streaming | 10 streaming / 50+ long-form | Streaming 1.5B · Streaming 7B · Long-form · README | | Canary | NVIDIA's multilingual ASR with translation | 25 EU + RU, UK | README | | Moonshine | Useful Sensors' lightweight ASR | EN | README | | MMS | Meta's massively multilingual ASR with adapters | 1000+ | README | | Granite Speech | IBM's ASR + speech translation | EN, FR, DE, ES, PT, JA | README | | Granite Speech 5.0 TurboCTC | IBM's fast encoder-only CTC ASR | EN | README | | Qwen2-Audio | Alibaba's multimodal audio understanding (ASR, captioning, emotion, translation) | Multiple | mlx-community/Qwen2-Audio-7B-Instruct-4bit | | MOSS-Music | OpenMOSS music understanding and lyrics ASR | EN, ZH | README |
Voice Activity Detection / Speaker Diarization (VAD)
| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Silero VAD | Lightweight speech/non-speech detection with streaming state | Language-agnostic | mlx-community/silero-vad | | Sortformer v1 | NVIDIA's end-to-end speaker diarization (up to 4 speakers) | Language-agnostic | mlx-community/diar_sortformer_4spk-v1-fp32 | | Sortformer v2.1 | NVIDIA's streaming speaker diarization with AOSC compression | Language-agnostic | mlx-community/diar_streaming_sortformer_4spk-v2.1-fp32 |
See the model READMEs for API details, streaming examples, and conversion steps.
Speech-to-Speech (STS)
| Model | Description | Use Case | Repo | |-------|-------------|----------|------| | SAM-Audio | Text-guided source separation | Extract specific sounds | mlx-community/sam-audio-large | | DialogueSidon | Two-speaker separation and restoration | Separate dialogue into speaker tracks | mlx-community/DialogueSidon (FP32), mlx-community/DialogueSidon-bf16 (BF16) | | Liquid2.5-Audio* | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | mlx-community/LFM2.5-Audio-1.5B-8bit | | MiMo-Audio | English/Chinese TTS, ASR, audio understanding and dialogue; Base few-shot speech tasks | Speech interactions and audio completion | Instruct, Base, audio tokenizer, guide | | MossFormer2 SE | Speech enhancement | Noise removal | starkdmi/MossFormer2_SE_48K_MLX | | DeepFilterNet (1/2/3) | Speech enhancement | Noise suppression | mlx-community/DeepFilterNet-mlx | | NemotronLabs VoiceChat | Full-duplex speech-to-speech with streaming transcription and function calling | Real-time voice conversation | mlx-community/NemotronLabs-VoiceChat-11B-4bit |
Music Generation
| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | MiniMax Music 3 | Hierarchical AR + flow-matching song generation with lyrics and 44.1 kHz stereo output | Multilingual lyrics | BF16, 8-bit, 6-bit, 4-bit, MXFP8, MXFP4, NVFP4, guide |
Model Examples
Qwen3-TTS
Alibaba's state-of-the-art multilingual TTS with voice cloning, emotion control, and voice design capabilities.
from mlx_audio.tts.utils import load_model
model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-bf16")
results = list(model.generate_custom_voice(
text="Hello, welcome to MLX-Audio!",
speaker="Vivian",
language="English",
))
audio = results[0].audio # mx.array
See the Qwen3-TTS README for voice cloning, CustomVoice, VoiceDesign, and all available models.
OmniVoice
OmniVoice is a zero-shot multilingual TTS model for 646+ languages with voice cloning, batch generation, pronunciation controls, and nonverbal tags such as [laughter] and [sigh]. It uses a bidirectional Qwen3 backbone with iterative masked generation and a HiggsAudioV2 acoustic tokenizer.
from mlx_audio.tts.utils import load_model
model = load_model("mlx-community/OmniVoice-bf16")
Basic multilingual TTS
for result in model.generate(
text="Hello from OmniVoice running on Apple Silicon.",
language="english",
duration_s=5.0,
num_steps=32,
):
audio = result.audio
Zero-shot voice cloning
for result in model.generate(
text="This sentence uses the reference speaker.",
language="english",
ref_audio="reference.wav",
ref_text="Transcript of the reference audio.",
duration_s=5.0,
):
audio = result.audio
For stable voice cloning, provide ref_text that matches the reference clip. OmniVoice also supports generate_batch() for batched TTS and inline pronunciation controls.
Ming Omni TTS (BailingMM)
mlx_audio.tts.generate \
--model mlx-community/Ming-omni-tts-16.8B-A3B-bf16 \
--prompt "Please generate speech based on the following description.\n" \
--text "This is a quick Ming Omni test." \
--lang_code en \
--output_path audio_io \
--file_prefix ming_basic \
--verbose
See the Ming Omni TTS README for CLI and Python cookbook examples, and the Ming Omni Dense README for the mlx-community/Ming-omni-tts-0.5B-bf16 workflow.
Kokoro TTS
Kokoro is a fast, multilingual TTS model with 54 voice presets.
from mlx_audio.tts.utils import load_model
model = load_model("mlx-community/Kokoro-82M-bf16")
Or use a quantized variant for lower memory usage:
model = load_model("mlx-community/Kokoro-82M-8bit")
model = load_model("mlx-community/Kokoro-82M-4bit")
Generate with different voices
for result in model.generate(
text="Welcome to MLX-Audio!",
voice="af_heart", # American female
speed=1.0,
lang_code="a" # American English
):
audio = result.audio
Available Voices:
- American English:
af_heart,af_bella,af_nova,af_sky,am_adam,am_echo, etc. - British English:
bf_alice,bf_emma,bm_daniel,bm_george, etc. - Japanese:
jf_alpha,jm_kumo, etc. - Chinese:
zf_xiaobei,zm_yunxi, etc.
pip install misaki for text processing. Japanese and Mandarin may additionally require pip install misaki[ja] or pip install misaki[zh].
Language Codes:
| Code | Language | Note |
|------|----------|------|
| a | American English | Default; requires pip install misaki |
| b | British English | Requires pip install misaki |
| j | Japanese | Requires pip install misaki[ja] |
| z | Mandarin Chinese | Requires pip install misaki[zh] |
| e | Spanish | Requires pip install misaki |
| f | French | Requires pip install misaki |
CSM (Voice Cloning)
Clone any voice using a reference audio sample:
mlx_audio.tts.generate \
--model mlx-community/csm-1b \
--text "Hello from Sesame." \
--ref_audio ./reference_voice.wav \
--play
Whisper STT
from mlx_audio.stt.generate import generate_transcription
result = generate_transcription(
model="mlx-community/whisper-large-v3-turbo-asr-fp16",
audio="audio.wav",
)
print(result.text)
Qwen3-ASR & ForcedAligner
Alibaba's multilingual speech models for transcription and word-level alignment.
from mlx_audio.stt import load
Speech recognition
model = load("mlx-community/Qwen3-ASR-0.6B-8bit")
result = model.generate("audio.wav", language="English")
print(result.text)
Word-level forced alignment
aligner = load("mlx-community/Qwen3-ForcedAligner-0.6B-8bit")
result = aligner.generate("audio.wav", text="I have a dream", language="English")
for item in result:
print(f"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}")
See the Qwen3-ASR README for CLI usage, all models, and more examples.
Phonon-1
Fermion Research's compact English Qwen3-ASR derivatives load directly from their public transport repositories:
from mlx_audio.stt import load
model = load("FermionResearch/Phonon-1")
result = model.generate("audio.wav", language="English")
print(result.text)
Available builds are Phonon-1-Micro (285 MB), Phonon-1 (415 MB), and
Phonon-1-Big (581 MB). See the Phonon-1 README.
VibeVoice-ASR
Microsoft's 9B parameter speech-to-text model with speaker diarization and timestamps. Supports long-form audio (up to 60 minutes) and outputs structured JSON.
from mlx_audio.stt.utils import load
model = load("mlx-community/VibeVoice-ASR-bf16")
Basic transcription
result = model.generate(audio="meeting.wav", max_tokens=8192, temperature=0.0)
print(result.text)
[{"Start":0,"End":5.2,"Speaker":0,"Content":"Hello everyone, let's begin."},
{"Start":5.5,"End":9.8,"Speaker":1,"Content":"Thanks for joining today."}]
Access parsed segments
for seg in result.segments:
print(f"[{seg['start_time']:.1f}-{seg['end_time']:.1f}] Speaker {seg['speaker_id']}: {seg['text']}")
Streaming transcription:
# Stream tokens as they are generated
for text in model.stream_transcribe(audio="speech.wav", max_tokens=4096):
print(text, end="", flush=True)
With context (hotwords/metadata):
result = model.generate(
audio="technical_talk.wav",
context="MLX, Apple Silicon, PyTorch, Transformer",
max_tokens=8192,
temperature=0.0,
)
CLI usage:
# Basic transcription
python -m mlx_audio.stt.generate \
--model mlx-community/VibeVoice-ASR-bf16 \
--audio meeting.wav \
--output-path output \
--format json \
--max-tokens 8192 \
--verbose
With context/hotwords
python -m mlx_audio.stt.generate \
--model mlx-community/VibeVoice-ASR-bf16 \
--audio technical_talk.wav \
--output-path output \
--format json \
--max-tokens 8192 \
--context "MLX, Apple Silicon, PyTorch, Transformer" \
--verbose
Parakeet (Multilingual STT)
NVIDIA's high-accuracy speech-to-text model. Parakeet v3 supports 25 European languages.
from mlx_audio.stt.utils import load
Load the multilingual v3 model
model = load("mlx-community/parakeet-tdt-0.6b-v3")
Transcribe audio
result = model.generate("audio.wav")
print(f"Text: {result.text}")
Access word-level timestamps
for sentence in result.sentences:
print(f"[{sentence.start:.2f}s - {sentence.end:.2f}s] {sentence.text}")
Streaming transcription:
```python for chunk in model.generate("long_audio.wav", stre