FluidInference/FluidAudio
Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.
About FluidInference/FluidAudio
FluidInference/FluidAudio is an open-source project on GitHub, mainly written in Swift. Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. It currently holds 2,833 stars and 429 forks with 22 open issues, and was last pushed on 2026-09-21 (repository created 2025-06-21).
Project Overview
Git Homed tracks it on the Today's Trending board, currently at rank #70 with 44 new stars today.
GitHub Repository Details
README
FluidAudio - Transcription, Text-to-speech, VAD, Speaker diarization with CoreML Models
FluidAudio is a Swift SDK for fully local, low-latency audio AI on Apple devices, with inference offloaded to the Apple Neural Engine (ANE), resulting in less memory and generally faster inference.
The SDK includes state-of-the-art speaker diarization, transcription, and voice activity detection via open-source models (MIT/Apache 2.0) that can be integrated with just a few lines of code. Models are optimized for background processing, ambient computing and always on workloads by running inference on the ANE, minimizing CPU usage and avoiding GPU/MPS entirely.
For custom use cases, feedback, additional model support, or platform requests, join our Discord. We're also bringing visual, language, and TTS models to device and will share updates there.
Below are some featured local AI apps using Fluid Audio models on macOS and iOS:
Want to convert your own model? Check möbius
Highlights
- Automatic Speech Recognition (ASR): Parakeet TDT v3 (0.6b) and other TDT/CTC models for batch transcription supporting 25 European languages and Japanese, plus SenseVoice and Paraformer for Mandarin Chinese; Parakeet EOU (120m) for streaming ASR with end-of-utterance detection (English only). See all ASR models.
- Inverse Text Normalization (ITN): Post-process ASR output to convert spoken-form to written-form ("two hundred" → "200"). See text-processing-rs. Optional: ASR-only apps can drop the engine (~8 MB per slice) with
traits: [](Swift 6.2+), see PostProcessing.md - Text-to-Speech (TTS): Kokoro (82m) for parallel synthesis with SSML and pronunciation control across 9 languages (EN, ES, FR, HI, IT, JA, PT, ZH); PocketTTS for streaming TTS with voice cloning support (EN, DE, ES, FR, IT, PT — 6L and 24L variants); Chatterbox Multilingual (520M, 18 languages) and Chatterbox Nano (110M, English with
[laugh]/[chuckle]paralinguistic tags) in beta — see Documentation/TTS/Chatterbox.md - Speaker Diarization (Online + Offline): Speaker separation and identification across audio streams. Streaming pipeline for real-time processing and offline batch pipeline with advanced clustering.
- Speaker Embedding Extraction: Generate speaker embeddings for voice comparison and clustering, you can use this for speaker identification
- Voice Activity Detection (VAD): Voice activity detection with Silero models
- Apple Neural Engine: Models run efficiently on Apple's ANE for maximum performance with minimal power consumption
- Open-Source Models: All models are publicly available on HuggingFace — converted and optimized by our team; permissive licenses. See full model catalog.
Video Demos
| Link | Description | | --- | --- | | Spokenly Real-time ASR | Video demonstration of FluidAudio's transcription accuracy and speed | | Senko Integration | Python Speaker diarization on Mac using FluidAudio's segmentation model | | Kokoro TTS | Text-to-speech demo using FluidAudio's Kokoro and Silero models on iOS | | Parakeet Realtime EOU | Parakeet streaming ASR with end-of-utterance detection on iOS | | Sortformer Diarization | Sortformer for speaker diarization with overlapping speech on iOS | | PocketTTS | Streaming text-to-speech using PocketTTS on iOS | | Parakeet EOU Ultra-Low Latency | Real-time Parakeet EOU transcription on iOS demonstrating ultra-low latency speech-to-text | | Action Phrase Live Production Control | Voice-controlled live production workflow using FluidAudio's ASR and speaker diarization to trigger cameras, graphics, and layouts with natural voice commands | | talat - VAD, ASR, Speaker ID | A video demo showcasing FluidAudio's VAD, two different ASR models, and speaker diarization during a talat.app meeting recording | | Kyutai PocketTTS on ANE | Kyutai labs PocketTTS in iOS running fast on the ANE & background-capable | | Supertonic-3 on iPhone 17 Pro ANE | Supertonic-3 running on iPhone 17 Pro via ANE/CoreML with 2 minutes of audio generated in 3 seconds with low RAM & background support |
Showcase
Make a PR if you want to add your app, please keep it in chronological order.
| App | GitHub | Description | | --- | :---: | --- | | Voice Ink | — | Local AI for instant, private transcription with near-perfect accuracy. Uses Parakeet ASR. | | Spokenly | — | Mac dictation app for fast, accurate voice-to-text; supports real-time dictation and file transcription. Uses Parakeet ASR and speaker diarization. | | Senko | ✓ | A very fast and accurate speaker diarization pipeline. A good example for how to integrate FluidAudio into a Python app | | Slipbox | — | Privacy-first meeting assistant for real-time conversation intelligence. Uses Parakeet ASR (iOS) and speaker diarization across platforms. | | Whisper Mate | — | Transcribes movies and audio locally; records and transcribes in real time from speakers or system apps. Uses speaker diarization. | | Altic/Fluid Voice | ✓ | Lightweight Fully free and Open Source Voice to Text dictation for macOS built using FluidAudio. Never pay for dictation apps | | Paraspeech | — | AI powered voice to text. Fully offline. No subscriptions. | | BoltAI | — | Write content 10x faster using parakeet models | | Dictate Anywhere | ✓ | Native macOS dictation app with global Fn key activation. Dictate into any app with 25 language support. Uses Parakeet ASR. | | hongbomiao.com | ✓ | A personal R&D lab that facilitates knowledge sharing. Uses Parakeet ASR. | | Hex | ✓ | macOS app that lets you press-and-hold a hotkey to record your voice, transcribe it, and paste into any application. Uses Parakeet ASR. | | Super Voice Assistant | ✓ | Open-source macOS voice assistant with local transcription. Uses Parakeet ASR. | | VoiceTypr | ✓ | Open-source voice-to-text dictation for macOS and Windows. Uses Parakeet ASR. | | Summit AI Notes | — | Local meeting transcription and summarization with speaker identification. Supports 100+ languages. | | Snaply | — | Free, Fast, 100% local AI dictation for Mac. | | OpenOats | ✓ | Open-source meeting note-taker that transcribes conversations in real time and surfaces relevant notes from your knowledge base. Uses FluidAudio for local transcription. | | Enconvo | — | AI Agent Launcher for macOS with voice input, live captions, and text-to-speech. Uses Parakeet ASR for local speech recognition. | | Meeting Transcriber | ✓ | macOS menu bar app that auto-detects, records, and transcribes meetings (Teams, Zoom, Webex) with dual-track speaker diarization. Uses Parakeet ASR and speaker diarization. | | Hitoku Draft | ✓ | A local, private, voice writing assistant on your macOS menu bar. Uses Parakeet ASR. | | Muesli | ✓ | Native macOS dictation and meeting transcription with ~0.13s latency. Captures microphone and system audio with automatic speaker diarization. Uses Parakeet TDT ASR. | | NanoVoice | — | Free iOS voice keyboard for fast, private dictation in any app. Uses Parakeet ASR. | | MiniWhisper | ✓ | Open-source macOS menu bar for quick local voice-to-text with minimal setup. Pick a shortcut, start talking. Uses Parakeet ASR. | | Talat | — | Privacy-focused AI meeting notes app. Records and transcribes meetings locally on your Mac with speaker identification and LLM-powered summaries. Featured in TechCrunch. Uses Parakeet ASR. | | VivaDicta | ✓ | Open-source iOS voice-to-text app with system-wide AI voice keyboard — dictate and AI-process text in any app. 15+ AI providers, 40+ AI presets. Uses Parakeet ASR. | | MimicScribe | — | macOS menu bar app combining Parakeet TDT streaming ASR, PyanNote Community 1 speaker diarization, and cloud LLMs to provide AI-generated talking points during meetings, derived from the live transcript and user-provided instructions. Features meeting summarization, natural language search, an MCP server for agent integration, and a keyboard- and voice-forward UI. | | Action Phrase | — | Voice-controlled live production app for iOS, iPadOS, and macOS. Control cameras, graphics, layouts, and production workflows with natural voice commands. Integrates with popular tools including OBS, vMix, ProPresenter, Bitfocus Companion, and more. Uses Parakeet TDT ASR and Sortformer diarization. | | Sayboard | ✓ | Privacy-first AI voice keyboard for iOS. Local models, no servers, no tracking, no subscriptions, no ads, no in-app purchases. Fully offline and open-source. | | Kesha Voice Kit | ✓ | Local-first voice toolkit for CLI and LLM-agent workflows. Speech-to-text in 25 languages, text-to-speech in 9, VAD, language detection, and MCP, OpenClaw and Hermes integrations. On Apple Silicon, FluidAudio powers ASR, Kokoro TTS and speaker diarization. | | Dictato | — | Turn your voice into text anywhere on your Mac. Fully local, private, and offline — boost your own vocabulary and dictate in multiple languages. Uses Parakeet TDT ASR. | | Utter | ✓ | An ultra-minimal speech-to-text status bar utility for Mac. Register a hotkey and go. | | Resonant | ✓ | macOS voice workspace for dictation, meetings, and ambient work context. Uses FluidAudio for local transcription and speaker diarization. | | Thoth | — | Privacy-first meeting recorder for Mac. Records both sides of any call with dual-channel audio, transcribes locally with speaker diarization, and summarizes with on-device AI or BYOK cloud. Available on the Mac App Store. Featured in MacGeneration. Uses Parakeet EOU and Parakeet TDT ASR. | | Dettivo | — | Local-first Mac app for private dictation, transcripts, and meeting workflows in one place, with developer tooling across CLI, MCP, REST, and app automation. Uses FluidAudio Parakeet TDT ASR and offline speaker diarization. | | Local Narrator | — | Privacy-first iOS audiobook reader that reads EPUB and PDF books aloud with on-device text-to-speech. Uses FluidAudio Kokoro TTS for local English and Spanish narration. | | Hedy | — | Privacy-first AI meeting coach for iOS, macOS, Android, and Windows. Real-time, fully on-device transcription with speaker diarization and AI-powered conversation insights. Uses Parakeet and Nemotron streaming ASR and speaker diarization on the Apple Neural Engine. | | Parakey | ✓ | Open-source (MIT) menu-bar push-to-talk dictation for macOS — hold a key, speak, release; the transcript pastes at the cursor in about 100 ms. Uses Parakeet TDT v3 ASR on the Apple Neural Engine via FluidAudio. | | TypeWhisper | — | Speech-to-text and AI text processing for macOS. Uses FluidAudio's Parakeet ASR for local transcription. | | evoglyph | — | Lightweight, privacy-first macOS menu-bar dictation — press a hotkey, speak, and cleaned-up text is injected at your cursor. Fully local: Parakeet TDT ASR with CTC vocabulary boosting and Silero VAD via FluidAudio on the Apple Neural Engine, plus on-device LLM cleanup. Audio never leaves the Mac. | | echo99 | — | Private call recorder for macOS. Menu-bar app that records the mic and system audio as separate tracks and transcribes them entirely on-device. | | Presspeech | ✓ | Open-source (MIT) menu-bar push-to-talk dictation for macOS — hold a key, speak, release; the transcript pastes at the cursor in about 100 ms. Uses Parakeet TDT v3 ASR on the Apple Neural Engine via FluidAudio. | | Better Voice | — | macOS menu-bar app for on-device dictation and meeting notes that save to Apple Notes. Everything runs locally. Uses speaker diarization. | | Logue | ✓ | Privacy-first AI meeting notes and writing assistant for macOS. Records mic + system audio and transcribes locally on Apple Silicon, with speaker diarization, Smart Minutes, and an on-device AI writing editor — nothing leaves the Mac. Uses FluidAudio streaming Sortformer speaker diarization. | | Goodmeet | — | The AI note-taker that puts your privacy first. Uses FluidAudio models for VAD and transcription. | | Subtitles | ✓ | Live captions for anything your Mac plays: meetings and calls, videos, podcasts and lectures, drawn as an always-on-top overlay that stays put while you switch apps. Captures system audio with a Core Audio process tap and transcribes entirely on-device, with selectable latency and optional speaker breaks. Uses Parakeet EOU and Nemotron streaming ASR, Silero VAD, and Sortformer speaker diarization on the Apple Neural Engine. | | Transkript | — | Offline AI transcription assistant for iPhone, iPad, and Mac. Transcribes audio/video files and live recordings in 25+ European languages, with color-coded speaker labels, AI summaries, translations, and subtitle export. Uses FluidAudio for ASR and speaker diarization. | | Notiva | — | Menu-bar notes app for Mac meetings. Live on-device transcription on the Neural Engine with speaker diarization, merged with your own notes into a single document. Uses FluidAudio for transcription and speaker diarization. | | oats | ✓ | Open-source (MIT) meeting notes app for macOS and Windows built with Tauri and Vue. One-click recording, real-time transcription with speaker labels, and on-device LLM note generation in Markdown. Uses FluidAudio on the Apple Neural Engine for macOS transcription. |
More apps built with FluidAudio are listed in Documentation/Showcase.md.
Installation
Add FluidAudio to your project using Swift Package Manager:
dependencies: [
.package(url: "https://github.com/FluidInference/FluidAudio.git", from: "0.12.4"),
],
In Xcode:
1. Add the FluidAudio package to your project
2. In the "Add Package" dialog, select FluidAudio
3. Add it to your app target
In Package.swift:
.product(name: "FluidAudio", package: "FluidAudio")
CocoaPods: We recommend using cocoapods-spm for better SPM integration, but if needed, you can also use our podspec: pod 'FluidAudio', '~> 0.12.4'
Other Frameworks
Building with a different framework? Use one of our official wrappers:
| Platform | Package | Install |
|----------|---------|---------|
| React Native / Expo | @fluidinference/react-native-fluidaudio | npm install @fluidinference/react-native-fluidaudio |
| Rust / Tauri | fluidaudio-rs | cargo add fluidaudio-rs |
Post-Processing Tools
Enhance ASR output with post-processing:
| Tool | Description | Language | |------|-------------|----------| | text-processing-rs | ✓ | Inverse Text Normalization (ITN) and Text Normalization (TN) across 7 languages (EN, DE, ES, FR, HI, JA, ZH). 100% NeMo test compatibility (3,011 tests). Converts spoken-form ASR output to written form ("two hundred" → "200", "five dollars" → "$5"). Rust port of NVIDIA NeMo Text Processing with Swift wrapper. | Rust, Swift |
Configuration
Quick Reference
Both solve the same problem: "I can't reach HuggingFace directly." They're alternative approaches - pick whichever matches your setup:
| Scenario | Solution | Configuration |
|----------|----------|---|
| You have a local mirror or internal model server | Use Registry URL override | REGISTRY_URL=https://your-mirror.com |
| You're behind a corporate firewall with a proxy that can reach HuggingFace | Use Proxy configuration | https_proxy=http://proxy.company.com:8080 |
How they work:
- Registry URL - App requests from
your-mirror.cominstead ofhuggingface.co - Proxy - App still requests
huggingface.co, but traffic routes through proxy to reach it
Model Registry URL - Change download destination
By default, FluidAudio downloads models from HuggingFace. You can override this to use a mirror, local server, or air-gapped environment.
Programmatic override (recommended for apps):
import FluidAudio
// Set custom registry before using any managers
ModelRegistry.baseURL = "https://your-mirror.example.com"
// Only needed when the mirror does not preserve an upstream pinned commit.
ModelRegistry.revisionOverrides = [
"FluidInference/speaker-diarization-coreml": "your-mirror-revision"
]
// Models will now download from the custom registry
let diarizer = DiarizerManager()
Mirrors should preserve upstream Git revisions when possible. For repositories
that FluidAudio pins to an immutable commit, set revisionOverrides explicitly
if the mirror exposes the same files under a different branch, tag, or commit.
Environment Variables (recommended for CLI/testing):
# Use custom registry
export REGISTRY_URL=https://your-mirror.example.com
swift run fluidaudiocli transcribe audio.wav
Or use the MODEL_REGISTRY_URL alias
export MODEL_REGISTRY_URL=https://models.internal.corp
swift run fluidaudiocli diarization-benchmark --auto-download
Xcode Scheme Configuration:
1. Edit Scheme → Run → Arguments
2. Go to Environment Variables tab
3. Click + and add: REGISTRY_URL = https://your-mirror.example.com
4. The custom registry will apply to all debug runs
Proxy Configuration - Route downloads through a proxy server
If you're behind a corporate firewall and cannot reach HuggingFace (or your registry) directly, configure a proxy to forward requests:
Set the https_proxy environment variable:
export https_proxy=http://proxy.company.com:8080
or for authenticated proxies:
export https_proxy=http://user:[email protected]:8080
swift run fluidaudiocli transcribe audio.wav
Xcode Scheme Configuration for Proxy:
1. Edit Scheme → Run → Arguments
2. Go to Environment Variables tab
3. Click + and add: https_proxy = http://proxy.company.com:8080
4. FluidAudio will automatically route downloads through the proxy
Offline-only mode - Refuse every network fetch, bundle your own models
If your application ships pre-downloaded model assets and never wants FluidAudio to reach HuggingFace at runtime (privacy-sensitive desktop apps, air-gapped deployments, kiosk builds), set the static ModelHub.offlineMode flag at startup:
```swift import FluidAudio
// Set once before any FluidAudio loader runs. ModelHub.offlineMode = true
// Load via manual APIs that read from your bundled directory: let asr = try await AsrMode









