Blaizzy/mlx-audio

★ 7,888⑂ 0

A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.

7,888Star
0Fork
0Watch
0Issue
PythonLanguage
-License
Created · last push · repository size 0 KB · default branch -

README

MLX-Audio

https://github.com/Blaizzy/mlx-audio/blob/HEAD/Blaizzy%2Fmlx-audio | Trendshift

PyPI version Python License: MIT GitHub stars

The best audio processing library built on Apple's MLX framework, providing fast and efficient text-to-speech (TTS), speech-to-text (STT), speech-to-speech (STS), music generation, and more on Apple Silicon.

Table of Contents

Features

Installation

Using pip

pip install mlx-audio

Using uv to install only the command line tools

Latest release from pypi:
uv tool install --force mlx-audio --prerelease=allow

Latest code from github:

uv tool install --force git+https://github.com/Blaizzy/mlx-audio.git --prerelease=allow

For development or web interface:

git clone https://github.com/Blaizzy/mlx-audio.git
cd mlx-audio
pip install -e ".[dev, server]"

Quick Start

Command Line

# Basic TTS generation
mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello, world!' --voice Vivian

With a different voice and language hint

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Welcome to MLX-Audio!' --voice Ryan --lang_code English

Play audio immediately

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --play

Save to a specific directory

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --output_path ./my_audio

Stream audio during generation

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream

Stream audio during generation and save it to disk

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text 'Hello!' --voice Vivian --stream --save

Join multiple generated segments into one file

mlx_audio.tts.generate --model mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit --text $'Hello!\nHow are you?' --voice Vivian --join_audio

By default, when generation yields multiple segments, mlx-audio saves numbered files such as audio_000.wav and audio_001.wav. Use --join_audio to save one combined file instead. When using --stream, add --save to write the streamed audio to disk.

Python API

from mlx_audio.tts.utils import load_model

Load model

model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-8bit")

Generate speech

for result in model.generate( "Hello from MLX-Audio!", voice="Vivian", lang_code="English", ): print(f"Generated {result.audio.shape[0]} samples") # result.audio contains the waveform as mx.array

Supported Models

Text-to-Speech (TTS)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Kokoro | Fast, high-quality multilingual TTS | EN, JA, ZH, FR, ES, IT, PT, HI | bf16, 8bit, 6bit, 4bit | | KittenTTS | Compact KittenTTS 0.8 models for edge-friendly TTS | EN | nano, micro, mini, collection | | Qwen3-TTS | Alibaba's multilingual TTS with voice design | ZH, EN, JA, KO, + more | mlx-community/Qwen3-TTS-12Hz-1.7B-VoiceDesign-bf16 | | Higgs Audio v3 | 4B conversational TTS with voice cloning and inline control tokens | 100 languages | bosonai/higgs-audio-v3-tts-4b | | OmniVoice | Zero-shot multilingual TTS with voice cloning, batch generation, and nonverbal tags | 646+ languages | mlx-community/OmniVoice-bf16 | | CSM / MisoTTS | Sesame-style conversational speech models with voice cloning | EN | mlx-community/csm-1b, MisoTTS bf16, MisoTTS 8bit | | Dia | Dialogue-focused TTS | EN | mlx-community/Dia-1.6B-fp16 | | OuteTTS | Efficient TTS model | EN | mlx-community/OuteTTS-1.0-0.6B-fp16 | | Spark | SparkTTS model | EN, ZH | mlx-community/Spark-TTS-0.5B-bf16 | | Chatterbox | Expressive multilingual TTS (v2/v3) | 23 languages | v3, v2 | | Soprano | High-quality TTS | EN | mlx-community/Soprano-1.1-80M-bf16 | | Ming Omni TTS (BailingMM) | Multimodal generation with voice cloning, style control, and speech/music/event generation | EN, ZH | mlx-community/Ming-omni-tts-16.8B-A3B-bf16 | | Ming Omni TTS (Dense) | Lightweight dense Ming Omni variant for voice cloning and style control | EN, ZH | mlx-community/Ming-omni-tts-0.5B-bf16 | | KugelAudio | SOTA 7B AR+Diffusion TTS for European languages | EN, DE, FR, ES, IT, PT, NL, PL, RU, UK, + 14 more | kugelaudio/kugelaudio-0-open | | Voxtral TTS | Mistral's 4B multilingual TTS (20 voices, 9 languages) | EN, FR, ES, DE, IT, PT, NL, AR, HI | mlx-community/Voxtral-4B-TTS-2603-mlx-bf16 | | rumik-oss 1 | 3B expressive multilingual Indic TTS with 22-language support, description-conditioned delivery and inline vocalizations | 22 Indic languages + EN | rumik-ai/rumik-oss-1, 8bit, 4bit | | VoxCPM2 | 2B tokenizer-free TTS with 48kHz output, voice design, voice cloning, and continuation | 30 languages | bf16, 8bit, 4bit | | LongCat-AudioDiT | SOTA diffusion TTS in waveform latent space with voice cloning | ZH, EN | mlx-community/LongCat-AudioDiT-1B-bf16 | | MeloTTS | Lightweight VITS2-based TTS with streaming | EN (more coming) | mlx-community/MeloTTS-English-MLX | | MOSS-TTS | 8B delay-pattern and local-transformer multilingual TTS with voice cloning | 31 languages | OpenMOSS-Team/MOSS-TTS-v1.5, OpenMOSS-Team/MOSS-TTS, OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5, OpenMOSS-Team/MOSS-TTS-Local-Transformer | | MOSS-TTS-Nano | Tiny multilingual voice-cloning TTS | 20 languages | mlx-community/MOSS-TTS-Nano-100M | | Higgs Audio v2 | 3B Llama-backed TTS with real-time voice cloning | EN, ZH, KO, DE, ES | bf16 (upstream), q8, q6 |

Speech-to-Text (STT)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Whisper | OpenAI's robust STT model | 99+ languages | mlx-community/whisper-large-v3-turbo-asr-fp16 | | Distil-Whisper | Distilled fast Whisper variants | EN | distil-whisper/distil-large-v3 | | Qwen3-ASR | Alibaba's multilingual ASR | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ASR-1.7B-8bit | | Mega-ASR | Routed Qwen3-ASR with automatic clean/base vs degraded/LoRA switching | EN (fixtures), multilingual Qwen3-ASR backbone | README | | Qwen3-ForcedAligner | Word-level audio alignment | ZH, EN, JA, KO, + more | mlx-community/Qwen3-ForcedAligner-0.6B-8bit | | MOSS-Transcribe-Diarize | Timestamped transcription with speaker labels | Multiple major languages | https://huggingface.co/OpenMOSS-Team/MOSS-Transcribe-Diarize | | Parakeet | NVIDIA's accurate STT | EN (v2), 25 EU languages (v3) | mlx-community/parakeet-tdt-0.6b-v3 | | Nemotron 3.5 ASR (streaming) | NVIDIA's cache-aware streaming FastConformer-RNNT with language-ID prompting | 40 language-locales | mlx-community/nemotron-3.5-asr-streaming-0.6b · README | | Voxtral | Mistral's speech model | Multiple | mlx-community/Voxtral-Mini-3B-2507-bf16 | | Voxtral Realtime | Mistral's 4B streaming STT | Multiple | 4bit, fp16 | | VibeVoice-ASR | Microsoft's 3B/9B ASR with diarization, timestamps, hotwords, and native chunk streaming | 10 streaming / 50+ long-form | Streaming 1.5B · Streaming 7B · Long-form · README | | Canary | NVIDIA's multilingual ASR with translation | 25 EU + RU, UK | README | | Moonshine | Useful Sensors' lightweight ASR | EN | README | | MMS | Meta's massively multilingual ASR with adapters | 1000+ | README | | Granite Speech | IBM's ASR + speech translation | EN, FR, DE, ES, PT, JA | README | | Granite Speech 5.0 TurboCTC | IBM's fast encoder-only CTC ASR | EN | README | | Qwen2-Audio | Alibaba's multimodal audio understanding (ASR, captioning, emotion, translation) | Multiple | mlx-community/Qwen2-Audio-7B-Instruct-4bit | | MOSS-Music | OpenMOSS music understanding and lyrics ASR | EN, ZH | README |

Voice Activity Detection / Speaker Diarization (VAD)

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | Silero VAD | Lightweight speech/non-speech detection with streaming state | Language-agnostic | mlx-community/silero-vad | | Sortformer v1 | NVIDIA's end-to-end speaker diarization (up to 4 speakers) | Language-agnostic | mlx-community/diar_sortformer_4spk-v1-fp32 | | Sortformer v2.1 | NVIDIA's streaming speaker diarization with AOSC compression | Language-agnostic | mlx-community/diar_streaming_sortformer_4spk-v2.1-fp32 |

See the model READMEs for API details, streaming examples, and conversion steps.

Speech-to-Speech (STS)

| Model | Description | Use Case | Repo | |-------|-------------|----------|------| | SAM-Audio | Text-guided source separation | Extract specific sounds | mlx-community/sam-audio-large | | DialogueSidon | Two-speaker separation and restoration | Separate dialogue into speaker tracks | mlx-community/DialogueSidon (FP32), mlx-community/DialogueSidon-bf16 (BF16) | | Liquid2.5-Audio* | Speech-to-Speech, Text-to-Speech and Speech-to-Text | Speech interactions | mlx-community/LFM2.5-Audio-1.5B-8bit | | MiMo-Audio | English/Chinese TTS, ASR, audio understanding and dialogue; Base few-shot speech tasks | Speech interactions and audio completion | Instruct, Base, audio tokenizer, guide | | MossFormer2 SE | Speech enhancement | Noise removal | starkdmi/MossFormer2_SE_48K_MLX | | DeepFilterNet (1/2/3) | Speech enhancement | Noise suppression | mlx-community/DeepFilterNet-mlx | | NemotronLabs VoiceChat | Full-duplex speech-to-speech with streaming transcription and function calling | Real-time voice conversation | mlx-community/NemotronLabs-VoiceChat-11B-4bit |

Music Generation

| Model | Description | Languages | Repo | |-------|-------------|-----------|------| | MiniMax Music 3 | Hierarchical AR + flow-matching song generation with lyrics and 44.1 kHz stereo output | Multilingual lyrics | BF16, 8-bit, 6-bit, 4-bit, MXFP8, MXFP4, NVFP4, guide |

Model Examples

Qwen3-TTS

Alibaba's state-of-the-art multilingual TTS with voice cloning, emotion control, and voice design capabilities.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/Qwen3-TTS-12Hz-0.6B-CustomVoice-bf16") results = list(model.generate_custom_voice( text="Hello, welcome to MLX-Audio!", speaker="Vivian", language="English", ))

audio = results[0].audio # mx.array

See the Qwen3-TTS README for voice cloning, CustomVoice, VoiceDesign, and all available models.

OmniVoice

OmniVoice is a zero-shot multilingual TTS model for 646+ languages with voice cloning, batch generation, pronunciation controls, and nonverbal tags such as [laughter] and [sigh]. It uses a bidirectional Qwen3 backbone with iterative masked generation and a HiggsAudioV2 acoustic tokenizer.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/OmniVoice-bf16")

Basic multilingual TTS

for result in model.generate( text="Hello from OmniVoice running on Apple Silicon.", language="english", duration_s=5.0, num_steps=32, ): audio = result.audio

Zero-shot voice cloning

for result in model.generate( text="This sentence uses the reference speaker.", language="english", ref_audio="reference.wav", ref_text="Transcript of the reference audio.", duration_s=5.0, ): audio = result.audio

For stable voice cloning, provide ref_text that matches the reference clip. OmniVoice also supports generate_batch() for batched TTS and inline pronunciation controls.

Ming Omni TTS (BailingMM)

mlx_audio.tts.generate \
    --model mlx-community/Ming-omni-tts-16.8B-A3B-bf16 \
    --prompt "Please generate speech based on the following description.\n" \
    --text "This is a quick Ming Omni test." \
    --lang_code en \
    --output_path audio_io \
    --file_prefix ming_basic \
    --verbose

See the Ming Omni TTS README for CLI and Python cookbook examples, and the Ming Omni Dense README for the mlx-community/Ming-omni-tts-0.5B-bf16 workflow.

Kokoro TTS

Kokoro is a fast, multilingual TTS model with 54 voice presets.

from mlx_audio.tts.utils import load_model

model = load_model("mlx-community/Kokoro-82M-bf16")

Or use a quantized variant for lower memory usage:

model = load_model("mlx-community/Kokoro-82M-8bit")

model = load_model("mlx-community/Kokoro-82M-4bit")

Generate with different voices

for result in model.generate( text="Welcome to MLX-Audio!", voice="af_heart", # American female speed=1.0, lang_code="a" # American English ): audio = result.audio

Available Voices:

Kokoro requires pip install misaki for text processing. Japanese and Mandarin may additionally require pip install misaki[ja] or pip install misaki[zh].

Language Codes: | Code | Language | Note | |------|----------|------| | a | American English | Default; requires pip install misaki | | b | British English | Requires pip install misaki | | j | Japanese | Requires pip install misaki[ja] | | z | Mandarin Chinese | Requires pip install misaki[zh] | | e | Spanish | Requires pip install misaki | | f | French | Requires pip install misaki |

CSM (Voice Cloning)

Clone any voice using a reference audio sample:

mlx_audio.tts.generate \
    --model mlx-community/csm-1b \
    --text "Hello from Sesame." \
    --ref_audio ./reference_voice.wav \
    --play

Whisper STT

from mlx_audio.stt.generate import generate_transcription

result = generate_transcription( model="mlx-community/whisper-large-v3-turbo-asr-fp16", audio="audio.wav", ) print(result.text)

Qwen3-ASR & ForcedAligner

Alibaba's multilingual speech models for transcription and word-level alignment.

from mlx_audio.stt import load

Speech recognition

model = load("mlx-community/Qwen3-ASR-0.6B-8bit") result = model.generate("audio.wav", language="English") print(result.text)

Word-level forced alignment

aligner = load("mlx-community/Qwen3-ForcedAligner-0.6B-8bit") result = aligner.generate("audio.wav", text="I have a dream", language="English") for item in result: print(f"[{item.start_time:.2f}s - {item.end_time:.2f}s] {item.text}")

See the Qwen3-ASR README for CLI usage, all models, and more examples.

Phonon-1

Fermion Research's compact English Qwen3-ASR derivatives load directly from their public transport repositories:

from mlx_audio.stt import load

model = load("FermionResearch/Phonon-1") result = model.generate("audio.wav", language="English") print(result.text)

Available builds are Phonon-1-Micro (285 MB), Phonon-1 (415 MB), and Phonon-1-Big (581 MB). See the Phonon-1 README.

VibeVoice-ASR

Microsoft's 9B parameter speech-to-text model with speaker diarization and timestamps. Supports long-form audio (up to 60 minutes) and outputs structured JSON.

from mlx_audio.stt.utils import load

model = load("mlx-community/VibeVoice-ASR-bf16")

Basic transcription

result = model.generate(audio="meeting.wav", max_tokens=8192, temperature=0.0) print(result.text)

[{"Start":0,"End":5.2,"Speaker":0,"Content":"Hello everyone, let's begin."},

{"Start":5.5,"End":9.8,"Speaker":1,"Content":"Thanks for joining today."}]

Access parsed segments

for seg in result.segments: print(f"[{seg['start_time']:.1f}-{seg['end_time']:.1f}] Speaker {seg['speaker_id']}: {seg['text']}")

Streaming transcription:

# Stream tokens as they are generated
for text in model.stream_transcribe(audio="speech.wav", max_tokens=4096):
    print(text, end="", flush=True)

With context (hotwords/metadata):

result = model.generate(
    audio="technical_talk.wav",
    context="MLX, Apple Silicon, PyTorch, Transformer",
    max_tokens=8192,
    temperature=0.0,
)

CLI usage:

# Basic transcription
python -m mlx_audio.stt.generate \
    --model mlx-community/VibeVoice-ASR-bf16 \
    --audio meeting.wav \
    --output-path output \
    --format json \
    --max-tokens 8192 \
    --verbose

With context/hotwords

python -m mlx_audio.stt.generate \ --model mlx-community/VibeVoice-ASR-bf16 \ --audio technical_talk.wav \ --output-path output \ --format json \ --max-tokens 8192 \ --context "MLX, Apple Silicon, PyTorch, Transformer" \ --verbose

Parakeet (Multilingual STT)

NVIDIA's high-accuracy speech-to-text model. Parakeet v3 supports 25 European languages.

from mlx_audio.stt.utils import load

Load the multilingual v3 model

model = load("mlx-community/parakeet-tdt-0.6b-v3")

Transcribe audio

result = model.generate("audio.wav") print(f"Text: {result.text}")

Access word-level timestamps

for sentence in result.sentences: print(f"[{sentence.start:.2f}s - {sentence.end:.2f}s] {sentence.text}")

Streaming transcription:

```python for chunk in model.generate("long_audio.wav", stre

More Audio Trending projects

1

huggingface / transformers

Python★ 166,108⑂ 0
2

harry0703 / MoneyPrinterTurbo

Python★ 123,776⑂ 0
3

unslothai / unsloth

Python★ 76,181⑂ 0
4

RVC-Boss / GPT-SoVITS

Python★ 61,798⑂ 0
5

calesthio / OpenMontage

Python★ 59,205⑂ 0
6

ggml-org / whisper.cpp

C++★ 53,674⑂ 0