OpenMOSS/MOSS-TTS
An open-source model family for long-form speech, dialogue synthesis, voice design, sound effects, and real-time streaming TTS
README
MOSS-TTS Family
MOSS‑TTS Family is an open‑source speech and sound generation model family from MOSI.AI and the OpenMOSS team. It is designed for high‑fidelity, high‑expressiveness, and complex real‑world scenarios, covering stable long‑form speech, multi‑speaker dialogue, voice/character design, environmental sound effects, and real‑time streaming TTS.
Start here: Quickstart · Model weights · Listen to samples · Fine-tuning · Serving backends
Choose a model for your task
| What you want to build | Start with | | --- | --- | | Speech and voice cloning on a CPU or in a browser | MOSS-TTS-Nano | | Multilingual long-form narration and voice cloning | MOSS-TTS-v1.5 · Local Transformer v1.5 | | Multi-speaker dialogue, podcasts, and dubbing | MOSS-TTSD | | Real-time streaming speech | MOSS-TTS-Realtime | | Voice design or environmental sound effects | Released models · MOSS-SoundEffect v2 |
News
- 2026.6.18: 🚀 MOSS-TTS-Local-Transformer-v1.5 receives Day-0 support in SGLang-Omni — the first inference backend to support the
MossTTSLocalarchitecture, with an OpenAI-compatible/v1/audio/speechendpoint, streaming, and voice cloning. See the cookbooks:moss_tts_local,moss_tts. - 2026.6.18: 🚀 Released MOSS-TTS-Local-Transformer-v1.5, a 4B
MossTTSLocalcheckpoint that inherits all v1.5 features (language tags, stable cloning, explicit pause control, etc.), scales the backbone from Qwen3-1.7B to Qwen3-4B, and uses MOSS-Audio-Tokenizer-v2 for native 48 kHz stereo output. - 2026.6.7: 🚀 Released MOSS-Audio-Tokenizer-v2, natively supporting 48 kHz stereo input and output. Check out the MOSS-Audio-Tokenizer repository for more details!
Earlier updates
- 2026.6.2: 🚀 vLLM-Omni now supports the full MOSS-TTS series (
MossTTSDelay,MossTTSRealtime, andMossTTSNanoarchitectures), including MOSS-TTS-v1.5, MOSS-TTS, MOSS-TTSD, MOSS-SoundEffect, MOSS-VoiceGenerator, MOSS-TTS-Realtime, and MOSS-TTS-Nano. See the recipe and examples. - 2026.5.26: 🚀 Released MOSS-SoundEffect-v2.0, a new text-to-audio model using a DiT backbone with the Flow Matching objective, generating 48 kHz bilingual sound effects up to 30 seconds — see
moss_soundeffect_v2/. - 2026.5.26: 🚀 Released MOSS-TTS-v1.5, with stronger multilingual synthesis when language tags are provided, more stable voice cloning, better long-reference short-text cloning, punctuation-following prosody, and explicit pause control via
[pause X.Ys]. - 2026.5.6: 🚀 MOSS-TTS and MOSS-Audio-Tokenizer now support
mlx-audio. Visit the mlx-audio GitHub repository for details. - 2026.4.29: 📝 MOSS-TTS 2.0 is coming soon! We are collecting TTS feedback, suggestions, and feature requests via the requirements collection form.
- 2026.4.13: 🚀 MOSS-TTS-Nano, our ~100M-parameter model, is now available! It supports multilingual voice cloning, 48 kHz stereo input/output, and streaming output on just 4 CPU cores. Check the GitHub repository and our blog for more details.
- 2026.3.31: 📄 Our technical reports for MOSS-TTSD and MOSS-VoiceGenerator are now available on arXiv!
- 2026.3.26: 📘 Added a tutorial on fine-tuning the MOSS-TTS-Realtime!
- 2026.3.20: 📄 Our technical report is now available on arXiv!
- 2026.3.18: 🚀 Added a first-class MOSS-TTS
llama.cppimplementation in the companion repositoryOpenMOSS/llama.cpp, including end-to-end docs and a runnable pipeline for GGUF backbone inference plus ONNX audio codec decoding. See the first-class e2e guide. - 2026.3.16: 📘 Added a tutorial on fine-tuning the MossTTSLocal architecture, suitable for MOSS-TTS-Local-Transformer!
- 2026.3.12: 🚀 Added SGLang backend support for the
MossTTSDelayarchitecture, enabling efficient inference for MOSS-TTS (Delay) and MOSS-SoundEffect, with around 3× faster generation throughput! - 2026.3.11: 📘 Added a tutorial on fine-tuning the MossTTSDelay architecture, suitable for MOSS-TTS(Delay), MOSS-TTSD, MOSS-VoiceGenerator, and MOSS-SoundEffect!
- 2026.3.10: ⚡️ Significantly optimized the VRAM usage of llama.cpp inference pipeline. Now 8B model fits onto 8GB GPUs!
- 2026.3.4: 🚀 Added PyTorch-free inference support — enabling lightweight on-device deployment via llama.cpp + ONNX Runtime. Quantized GGUF weights are released at OpenMOSS-Team/MOSS-TTS-GGUF, and the ONNX audio tokenizer is available at OpenMOSS-Team/MOSS-Audio-Tokenizer-ONNX. See the llama.cpp backend for details.
- 2026.3.4: 🎉 We add MOSS-TTS skills in ClawHub of 🦞 OpenClaw: feishu-voice-tts and moss-tts-voice.
- 2026.2.10: 🎉🎉🎉 We have released MOSS-TTS Family. Check our Blog for more details! Our Huggingface Space is here: MOSS-TTS, MOSS-TTSD-v1.0, MOSS-VoiceGenerator.
Demo
Contents
- MOSS-TTS Family
- News
- Demo
- Contents
- Introduction
- Model Architecture
- Released Models
- Supported Languages
- MOSS-TTS-v1.5
- MOSS-TTS-Local-Transformer-v1.5
- Quickstart
- OpenClaw API Skills
- Environment Setup
- Using Conda
- Using
uv - (Optional) Install FlashAttention 2
- MOSS‑TTS Basic Usage
- Fine-Tuning
- llama.cpp Backend (Torch-Free Inference)
- Quick Start
- Installation Profiles
- Model Weights
- Configuration
- Accelerated Inference Backends
- SGLang-Omni
- vLLM-Omni
- Evaluation
- MOSS‑TTS
- MOSS‑TTSD
- Objective Evaluation
- Subjective Evaluation
- MOSS‑VoiceGenerator
- MOSS‑TTS-Realtime
- MOSS-TTS-Nano
- Introduction
- Model Weights
- MOSS-Audio-Tokenizer
- Introduction
- Model Weights
- Objective Reconstruction Evaluation
- 📚 More Information
- 🌟 Community Projects
- LICENSE
- Citation
- Star History
Introduction
When a single piece of audio needs to sound like a real person, pronounce every word accurately, switch speaking styles across content, remain stable over tens of minutes, and support dialogue, role‑play, and real‑time interaction, a single TTS model is often not enough. The MOSS‑TTS Family breaks the workflow into five production‑ready models that can be used independently or composed into a complete pipeline.
- MOSS‑TTS: The flagship production model featuring high fidelity and optimal zero-shot voice cloning. It supports long-speech generation, fine-grained control over Pinyin, phonemes, and duration, as well as multilingual/code-switched synthesis.
- MOSS‑TTSD: A spoken dialogue generation model for expressive, multi-speaker, and ultra-long dialogues. The new v1.0 version achieves industry-leading performance on objective metrics and outperformed top closed-source models like Doubao and Gemini 2.5-pro in subjective evaluations. You can visit the MOSS-TTSD repository for details.
- MOSS‑VoiceGenerator: An open-source voice design model capable of generating diverse voices and styles directly from text prompts, without any reference speech. It unifies voice design, style control, and synthesis, functioning independently or as a design layer for downstream TTS. Its performance surpasses other top-tier voice design models in arena ratings.
- MOSS‑TTS‑Realtime: A multi-turn context-aware model for real-time voice agents. It uses incremental synthesis to ensure natural and coherent replies, making it ideal for building low-latency voice agents when paired with text models. The TTFB (Time To First Byte) of MOSS-TTS-Realtime reaches 180 ms, and the $T_{\text{LLM-first-sentence}} + T_{\text{MOSS-TTS-Realtime-TTFB}}$ is 377 ms.
- MOSS‑SoundEffect: A content creation model specialized in sound effect generation with wide category coverage and controllable duration. It generates audio for natural environments, urban scenes, biological sounds, human actions, and musical fragments, suitable for film, games, and interactive experiences.
Model Architecture
We train MossTTSDelay and MossTTSLocal as complementary baselines under one training/evaluation setup: Delay emphasizes long-context stability, inference speed, and production readiness, while Local emphasizes lightweight flexibility and strong objective performance for streaming-oriented systems. Together they provide reproducible references for deployment and research.
MossTTSRealtime is not a third comparison baseline; it is a capability-driven design for voice agents. By modeling multi-turn context from both prior text and user acoustics, it delivers low-latency streaming speech that stays coherent and voice-consistent across turns.
| Architecture | Core Mechanism | Arch Details |
|---|---|---|
| MossTTSDelay | Multi‑head parallel RVQ prediction with delay‑pattern scheduling | |
|
MossTTSLocal | Time‑synchronous RVQ blocks with a depth transformer | |
|
MossTTSRealtime | Hierarchical text–audio inputs for realtime synthesis | |
Released Models
| Model | Architecture | Size | Model Card | Hugging Face | ModelScope |
|---|---|---:|---|---|---|
| MOSS-TTS-v1.5 | MossTTSDelay | 8B | |
|
|
| MOSS-TTS 1.0 |
MossTTSDelay | 8B | |
|
|
| MOSS-TTS-Local-Transformer-v1.5 |
MossTTSLocal | 4B | |
|
|
| MOSS-TTS-Local-Transformer |
MossTTSLocal | 1.7B | |
|
|
| MOSS‑TTSD‑V1.0 |
MossTTSDelay | 8B | |
|
|
| MOSS‑VoiceGenerator |
MossTTSDelay | 1.7B | |
|
|
| MOSS‑SoundEffect |
MossTTSDelay | 8B | |
|
|
| MOSS‑SoundEffect‑v2.0 |
MossSoundEffectPipeline | 1.3B DiT | |
|
|
| MOSS‑TTS‑Realtime |
MossTTSRealtime | 1.7B | |
|
|
Supported Languages
MOSS-TTS-v1.5 and MOSS-TTS-Local-Transformer-v1.5 currently support 31 languages. They keep the 20 languages supported by MOSS-TTS 1.0 and extend multilingual continued training to Cantonese, Dutch, Finnish, Hindi, Macedonian, Malay, Romanian, Swahili, Tagalog, Thai, and Vietnamese.
MOSS-TTSD and MOSS-TTS-Realtime follow their own model cards for supported language coverage.
| Language | Code | Flag | Language | Code | Flag | Language | Code | Flag | |---|---|---|---|---|---|---|---|---| | Chinese | zh | 🇨🇳 | Cantonese | yue | 🇭🇰 | English | en | 🇺🇸 | | Arabic | ar | 🇸🇦 | Czech | cs | 🇨🇿 | Danish | da | 🇩🇰 | | Dutch | nl | 🇳🇱 | Finnish | fi | 🇫🇮 | French | fr | 🇫🇷 | | German | de | 🇩🇪 | Greek | el | 🇬🇷 | Hebrew | he | 🇮🇱 | | Hindi | hi | 🇮🇳 | Hungarian | hu | 🇭🇺 | Italian | it | 🇮🇹 | | Japanese | ja | 🇯🇵 | Korean | ko | 🇰🇷 | Macedonian | mk | 🇲🇰 | | Malay | ms | 🇲🇾 | Persian (Farsi) | fa | 🇮🇷 | Polish | pl | 🇵🇱 | | Portuguese | pt | 🇵🇹 | Romanian | ro | 🇷🇴 | Russian | ru | 🇷🇺 | | Spanish | es | 🇪🇸 | Swahili | sw | 🇹🇿 | Swedish | sv | 🇸🇪 | | Tagalog | tl | 🇵🇭 | Thai | th | 🇹🇭 | Turkish | tr | 🇹🇷 | | Vietnamese | vi | 🇻🇳 | | | | | | |
MOSS-TTS-v1.5
MOSS-TTS-v1.5 is continued from MOSS-TTS 1.0. It preserves the main 1.0 capabilities, including zero-shot voice cloning, long-form speech generation, token-level duration control, Pinyin/IPA pronunciation control, multilingual synthesis, and code-switching.
Compared with MOSS-TTS 1.0, v1.5 focuses on these improvements:
- Stronger multilingual synthesis with language tags: when the language is known, set it in
processor.build_user_message(text=text, language="French")or the equivalent API field. - More stable voice cloning: v1.5 improves speaker similarity and reduces cloning variance across repeated generations.
- Better long-reference, short-text cloning: v1.5 handles reference audio that is much longer than the target text more reliably.
- More stable punctuation-following prosody: v1.5 follows punctuation-driven pauses more closely, especially in long sentences.
- Explicit pause control: use inline pause markers such as
[pause 3.2s], for example我今天学习了一首中国的古诗,它的名字是[pause 3.2s]静夜思!.
MOSS-TTS-Local-Transformer-v1.5
MOSS-TTS-Local-Transformer-v1.5 is the 48 kHz stereo local-transformer release. It uses MOSS-Audio-Tokenizer-v2 as the audio tokenizer and adds realtime streaming decode examples in moss_tts_local_v1.5/.
Compared with MOSS-TTS-Local-Transformer-v1.0, v1.5 focuses on these improvements:
- Higher-fidelity stereo audio modeling: v1.5 uses MOSS-Audio-Tokenizer-v2 as the audio tokenizer, supporting native 48 kHz stereo input and output for richer spatial detail and more natural perceived audio quality.
- Stronger multilingual synthesis with language tags: when the language is known, set it in
processor.build_user_message(text=text, language="French")or the equivalent API field. - More stable voice cloning: v1.5 improves speaker similarity and reduces cloning variance across repeated generations.
- Better long-reference, short-text cloning: v1.5 handles reference audio that is much longer than the target text more reliably.
- More stable punctuation-following prosody: v1.5 follows punctuation-driven pauses more closely, especially in long sentences.
- Explicit pause control: use inline pause markers such as
[pause 3.2s], for example我今天学习了一首中国的古诗,它的名字是[pause 3.2s]静夜思!.
Quickstart
OpenClaw API Skills
We add MOSS-TTS skills in ClawHub of 🦞 OpenClaw. You can get your API key from MOSI AI Studio.
| Skill | Description | Install |
|---|---|---|
| feishu-voice-tts | Send voice messages in Feishu | clawhub install feishu-voice-tts |
| moss-tts-voice | Call MOSS-TTS API to generate speech | clawhub install moss-tts-voice |
Environment Setup
We recommend a clean, isolated Python envi