FluidInference/FluidAudio

▲ 44 stars today★ 2,833⑂ 429

Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. In Swift, powered by SOTA open source.

About FluidInference/FluidAudio

FluidInference/FluidAudio is an open-source project on GitHub, mainly written in Swift. Frontier CoreML audio models in your apps — text-to-speech, speech-to-text, voice activity detection, and speaker diarization. It currently holds 2,833 stars and 429 forks with 22 open issues, and was last pushed on 2026-09-21 (repository created 2025-06-21).

Project Overview

Git Homed tracks it on the Today's Trending board, currently at rank #70 with 44 new stars today.

GitHub Repository Details

Repository FluidInference/FluidAudio · default branch main · size 295143 KB · watchers 43 · source: GitHub REST API and repository README

README

banner.png

FluidAudio - Transcription, Text-to-speech, VAD, Speaker diarization with CoreML Models

Swift Platform Documentation Discord Hugging Face ModelsAsk DeepWiki

FluidAudio is a Swift SDK for fully local, low-latency audio AI on Apple devices, with inference offloaded to the Apple Neural Engine (ANE), resulting in less memory and generally faster inference.

The SDK includes state-of-the-art speaker diarization, transcription, and voice activity detection via open-source models (MIT/Apache 2.0) that can be integrated with just a few lines of code. Models are optimized for background processing, ambient computing and always on workloads by running inference on the ANE, minimizing CPU usage and avoiding GPU/MPS entirely.

For custom use cases, feedback, additional model support, or platform requests, join our Discord. We're also bringing visual, language, and TTS models to device and will share updates there.

Below are some featured local AI apps using Fluid Audio models on macOS and iOS:

https://github.com/FluidInference/FluidAudio/blob/HEAD/Voice Ink https://github.com/FluidInference/FluidAudio/blob/HEAD/Spokenly https://github.com/FluidInference/FluidAudio/blob/HEAD/Slipbox https://github.com/FluidInference/FluidAudio/blob/HEAD/Hex https://github.com/FluidInference/FluidAudio/blob/HEAD/BoltAI https://github.com/FluidInference/FluidAudio/blob/HEAD/Paraspeech https://github.com/FluidInference/FluidAudio/blob/HEAD/Fluid Voice https://github.com/FluidInference/FluidAudio/blob/HEAD/Snaply https://github.com/FluidInference/FluidAudio/blob/HEAD/OpenOats https://github.com/FluidInference/FluidAudio/blob/HEAD/Talat

Want to convert your own model? Check möbius

Highlights

Video Demos

| Link | Description | | --- | --- | | Spokenly Real-time ASR | Video demonstration of FluidAudio's transcription accuracy and speed | | Senko Integration | Python Speaker diarization on Mac using FluidAudio's segmentation model | | Kokoro TTS | Text-to-speech demo using FluidAudio's Kokoro and Silero models on iOS | | Parakeet Realtime EOU | Parakeet streaming ASR with end-of-utterance detection on iOS | | Sortformer Diarization | Sortformer for speaker diarization with overlapping speech on iOS | | PocketTTS | Streaming text-to-speech using PocketTTS on iOS | | Parakeet EOU Ultra-Low Latency | Real-time Parakeet EOU transcription on iOS demonstrating ultra-low latency speech-to-text | | Action Phrase Live Production Control | Voice-controlled live production workflow using FluidAudio's ASR and speaker diarization to trigger cameras, graphics, and layouts with natural voice commands | | talat - VAD, ASR, Speaker ID | A video demo showcasing FluidAudio's VAD, two different ASR models, and speaker diarization during a talat.app meeting recording | | Kyutai PocketTTS on ANE | Kyutai labs PocketTTS in iOS running fast on the ANE & background-capable | | Supertonic-3 on iPhone 17 Pro ANE | Supertonic-3 running on iPhone 17 Pro via ANE/CoreML with 2 minutes of audio generated in 3 seconds with low RAM & background support |

Showcase

Make a PR if you want to add your app, please keep it in chronological order.

| App | GitHub | Description | | --- | :---: | --- | | Voice Ink | — | Local AI for instant, private transcription with near-perfect accuracy. Uses Parakeet ASR. | | Spokenly | — | Mac dictation app for fast, accurate voice-to-text; supports real-time dictation and file transcription. Uses Parakeet ASR and speaker diarization. | | Senko | ✓ | A very fast and accurate speaker diarization pipeline. A good example for how to integrate FluidAudio into a Python app | | Slipbox | — | Privacy-first meeting assistant for real-time conversation intelligence. Uses Parakeet ASR (iOS) and speaker diarization across platforms. | | Whisper Mate | — | Transcribes movies and audio locally; records and transcribes in real time from speakers or system apps. Uses speaker diarization. | | Altic/Fluid Voice | ✓ | Lightweight Fully free and Open Source Voice to Text dictation for macOS built using FluidAudio. Never pay for dictation apps | | Paraspeech | — | AI powered voice to text. Fully offline. No subscriptions. | | BoltAI | — | Write content 10x faster using parakeet models | | Dictate Anywhere | ✓ | Native macOS dictation app with global Fn key activation. Dictate into any app with 25 language support. Uses Parakeet ASR. | | hongbomiao.com | ✓ | A personal R&D lab that facilitates knowledge sharing. Uses Parakeet ASR. | | Hex | ✓ | macOS app that lets you press-and-hold a hotkey to record your voice, transcribe it, and paste into any application. Uses Parakeet ASR. | | Super Voice Assistant | ✓ | Open-source macOS voice assistant with local transcription. Uses Parakeet ASR. | | VoiceTypr | ✓ | Open-source voice-to-text dictation for macOS and Windows. Uses Parakeet ASR. | | Summit AI Notes | — | Local meeting transcription and summarization with speaker identification. Supports 100+ languages. | | Snaply | — | Free, Fast, 100% local AI dictation for Mac. | | OpenOats | ✓ | Open-source meeting note-taker that transcribes conversations in real time and surfaces relevant notes from your knowledge base. Uses FluidAudio for local transcription. | | Enconvo | — | AI Agent Launcher for macOS with voice input, live captions, and text-to-speech. Uses Parakeet ASR for local speech recognition. | | Meeting Transcriber | ✓ | macOS menu bar app that auto-detects, records, and transcribes meetings (Teams, Zoom, Webex) with dual-track speaker diarization. Uses Parakeet ASR and speaker diarization. | | Hitoku Draft | | A local, private, voice writing assistant on your macOS menu bar. Uses Parakeet ASR. | | Muesli | ✓ | Native macOS dictation and meeting transcription with ~0.13s latency. Captures microphone and system audio with automatic speaker diarization. Uses Parakeet TDT ASR. | | NanoVoice | — | Free iOS voice keyboard for fast, private dictation in any app. Uses Parakeet ASR. | | MiniWhisper | ✓ | Open-source macOS menu bar for quick local voice-to-text with minimal setup. Pick a shortcut, start talking. Uses Parakeet ASR. | | Talat | — | Privacy-focused AI meeting notes app. Records and transcribes meetings locally on your Mac with speaker identification and LLM-powered summaries. Featured in TechCrunch. Uses Parakeet ASR. | | VivaDicta | ✓ | Open-source iOS voice-to-text app with system-wide AI voice keyboard — dictate and AI-process text in any app. 15+ AI providers, 40+ AI presets. Uses Parakeet ASR. | | MimicScribe | — | macOS menu bar app combining Parakeet TDT streaming ASR, PyanNote Community 1 speaker diarization, and cloud LLMs to provide AI-generated talking points during meetings, derived from the live transcript and user-provided instructions. Features meeting summarization, natural language search, an MCP server for agent integration, and a keyboard- and voice-forward UI. | | Action Phrase | — | Voice-controlled live production app for iOS, iPadOS, and macOS. Control cameras, graphics, layouts, and production workflows with natural voice commands. Integrates with popular tools including OBS, vMix, ProPresenter, Bitfocus Companion, and more. Uses Parakeet TDT ASR and Sortformer diarization. | | Sayboard | ✓ | Privacy-first AI voice keyboard for iOS. Local models, no servers, no tracking, no subscriptions, no ads, no in-app purchases. Fully offline and open-source. | | Kesha Voice Kit | ✓ | Local-first voice toolkit for CLI and LLM-agent workflows. Speech-to-text in 25 languages, text-to-speech in 9, VAD, language detection, and MCP, OpenClaw and Hermes integrations. On Apple Silicon, FluidAudio powers ASR, Kokoro TTS and speaker diarization. | | Dictato | — | Turn your voice into text anywhere on your Mac. Fully local, private, and offline — boost your own vocabulary and dictate in multiple languages. Uses Parakeet TDT ASR. | | Utter | ✓ | An ultra-minimal speech-to-text status bar utility for Mac. Register a hotkey and go. | | Resonant | | macOS voice workspace for dictation, meetings, and ambient work context. Uses FluidAudio for local transcription and speaker diarization. | | Thoth | — | Privacy-first meeting recorder for Mac. Records both sides of any call with dual-channel audio, transcribes locally with speaker diarization, and summarizes with on-device AI or BYOK cloud. Available on the Mac App Store. Featured in MacGeneration. Uses Parakeet EOU and Parakeet TDT ASR. | | Dettivo | — | Local-first Mac app for private dictation, transcripts, and meeting workflows in one place, with developer tooling across CLI, MCP, REST, and app automation. Uses FluidAudio Parakeet TDT ASR and offline speaker diarization. | | Local Narrator | — | Privacy-first iOS audiobook reader that reads EPUB and PDF books aloud with on-device text-to-speech. Uses FluidAudio Kokoro TTS for local English and Spanish narration. | | Hedy | — | Privacy-first AI meeting coach for iOS, macOS, Android, and Windows. Real-time, fully on-device transcription with speaker diarization and AI-powered conversation insights. Uses Parakeet and Nemotron streaming ASR and speaker diarization on the Apple Neural Engine. | | Parakey | ✓ | Open-source (MIT) menu-bar push-to-talk dictation for macOS — hold a key, speak, release; the transcript pastes at the cursor in about 100 ms. Uses Parakeet TDT v3 ASR on the Apple Neural Engine via FluidAudio. | | TypeWhisper | — | Speech-to-text and AI text processing for macOS. Uses FluidAudio's Parakeet ASR for local transcription. | | evoglyph | — | Lightweight, privacy-first macOS menu-bar dictation — press a hotkey, speak, and cleaned-up text is injected at your cursor. Fully local: Parakeet TDT ASR with CTC vocabulary boosting and Silero VAD via FluidAudio on the Apple Neural Engine, plus on-device LLM cleanup. Audio never leaves the Mac. | | echo99 | — | Private call recorder for macOS. Menu-bar app that records the mic and system audio as separate tracks and transcribes them entirely on-device. | | Presspeech | ✓ | Open-source (MIT) menu-bar push-to-talk dictation for macOS — hold a key, speak, release; the transcript pastes at the cursor in about 100 ms. Uses Parakeet TDT v3 ASR on the Apple Neural Engine via FluidAudio. | | Better Voice | — | macOS menu-bar app for on-device dictation and meeting notes that save to Apple Notes. Everything runs locally. Uses speaker diarization. | | Logue | ✓ | Privacy-first AI meeting notes and writing assistant for macOS. Records mic + system audio and transcribes locally on Apple Silicon, with speaker diarization, Smart Minutes, and an on-device AI writing editor — nothing leaves the Mac. Uses FluidAudio streaming Sortformer speaker diarization. | | Goodmeet | — | The AI note-taker that puts your privacy first. Uses FluidAudio models for VAD and transcription. | | Subtitles | | Live captions for anything your Mac plays: meetings and calls, videos, podcasts and lectures, drawn as an always-on-top overlay that stays put while you switch apps. Captures system audio with a Core Audio process tap and transcribes entirely on-device, with selectable latency and optional speaker breaks. Uses Parakeet EOU and Nemotron streaming ASR, Silero VAD, and Sortformer speaker diarization on the Apple Neural Engine. | | Transkript | — | Offline AI transcription assistant for iPhone, iPad, and Mac. Transcribes audio/video files and live recordings in 25+ European languages, with color-coded speaker labels, AI summaries, translations, and subtitle export. Uses FluidAudio for ASR and speaker diarization. | | Notiva | — | Menu-bar notes app for Mac meetings. Live on-device transcription on the Neural Engine with speaker diarization, merged with your own notes into a single document. Uses FluidAudio for transcription and speaker diarization. | | oats | ✓ | Open-source (MIT) meeting notes app for macOS and Windows built with Tauri and Vue. One-click recording, real-time transcription with speaker labels, and on-device LLM note generation in Markdown. Uses FluidAudio on the Apple Neural Engine for macOS transcription. |

More apps built with FluidAudio are listed in Documentation/Showcase.md.

Installation

Add FluidAudio to your project using Swift Package Manager:

dependencies: [
    .package(url: "https://github.com/FluidInference/FluidAudio.git", from: "0.12.4"),
],

In Xcode: 1. Add the FluidAudio package to your project 2. In the "Add Package" dialog, select FluidAudio 3. Add it to your app target

In Package.swift:

.product(name: "FluidAudio", package: "FluidAudio")

CocoaPods: We recommend using cocoapods-spm for better SPM integration, but if needed, you can also use our podspec: pod 'FluidAudio', '~> 0.12.4'

Other Frameworks

Building with a different framework? Use one of our official wrappers:

| Platform | Package | Install | |----------|---------|---------| | React Native / Expo | @fluidinference/react-native-fluidaudio | npm install @fluidinference/react-native-fluidaudio | | Rust / Tauri | fluidaudio-rs | cargo add fluidaudio-rs |

Post-Processing Tools

Enhance ASR output with post-processing:

| Tool | Description | Language | |------|-------------|----------| | text-processing-rs | ✓ | Inverse Text Normalization (ITN) and Text Normalization (TN) across 7 languages (EN, DE, ES, FR, HI, JA, ZH). 100% NeMo test compatibility (3,011 tests). Converts spoken-form ASR output to written form ("two hundred" → "200", "five dollars" → "$5"). Rust port of NVIDIA NeMo Text Processing with Swift wrapper. | Rust, Swift |

Configuration

Quick Reference

Both solve the same problem: "I can't reach HuggingFace directly." They're alternative approaches - pick whichever matches your setup:

| Scenario | Solution | Configuration | |----------|----------|---| | You have a local mirror or internal model server | Use Registry URL override | REGISTRY_URL=https://your-mirror.com | | You're behind a corporate firewall with a proxy that can reach HuggingFace | Use Proxy configuration | https_proxy=http://proxy.company.com:8080 |

How they work:

In most cases, you only need one. (You'd use both only if your mirror is behind the proxy and unreachable without it.)

Model Registry URL - Change download destination

By default, FluidAudio downloads models from HuggingFace. You can override this to use a mirror, local server, or air-gapped environment.

Programmatic override (recommended for apps):

import FluidAudio

// Set custom registry before using any managers ModelRegistry.baseURL = "https://your-mirror.example.com"

// Only needed when the mirror does not preserve an upstream pinned commit. ModelRegistry.revisionOverrides = [ "FluidInference/speaker-diarization-coreml": "your-mirror-revision" ]

// Models will now download from the custom registry let diarizer = DiarizerManager()

Mirrors should preserve upstream Git revisions when possible. For repositories that FluidAudio pins to an immutable commit, set revisionOverrides explicitly if the mirror exposes the same files under a different branch, tag, or commit.

Environment Variables (recommended for CLI/testing):

# Use custom registry
export REGISTRY_URL=https://your-mirror.example.com
swift run fluidaudiocli transcribe audio.wav

Or use the MODEL_REGISTRY_URL alias

export MODEL_REGISTRY_URL=https://models.internal.corp swift run fluidaudiocli diarization-benchmark --auto-download

Xcode Scheme Configuration: 1. Edit Scheme → Run → Arguments 2. Go to Environment Variables tab 3. Click + and add: REGISTRY_URL = https://your-mirror.example.com 4. The custom registry will apply to all debug runs

Proxy Configuration - Route downloads through a proxy server

If you're behind a corporate firewall and cannot reach HuggingFace (or your registry) directly, configure a proxy to forward requests:

Set the https_proxy environment variable:

export https_proxy=http://proxy.company.com:8080

or for authenticated proxies:

export https_proxy=http://user:[email protected]:8080

swift run fluidaudiocli transcribe audio.wav

Xcode Scheme Configuration for Proxy: 1. Edit Scheme → Run → Arguments 2. Go to Environment Variables tab 3. Click + and add: https_proxy = http://proxy.company.com:8080 4. FluidAudio will automatically route downloads through the proxy

Offline-only mode - Refuse every network fetch, bundle your own models

If your application ships pre-downloaded model assets and never wants FluidAudio to reach HuggingFace at runtime (privacy-sensitive desktop apps, air-gapped deployments, kiosk builds), set the static ModelHub.offlineMode flag at startup:

```swift import FluidAudio

// Set once before any FluidAudio loader runs. ModelHub.offlineMode = true

// Load via manual APIs that read from your bundled directory: let asr = try await AsrMode

GitHub Stars & Activity

2,833Stars
429Forks
22Open issues
SwiftLanguage

GitHub Popularity

GitHub stars2,833
Forks429
Open issues22
Primary languageSwift
LicenseApache-2.0
Stars gained today44
Created2025-06-21
Last pushed2026-09-21

Trending History

Daily boardrank #70 · ▲ 44 stars

Related GitHub Projects

1

MonitorControl / MonitorControl

Swift★ 34,276⑂ 1,011▲ 30 stars
2

LiveContainer / LiveContainer

Swift★ 12,312⑂ 1,027▲ 29 stars
3

mrkai77 / Loop

Swift★ 11,647⑂ 273▲ 27 stars
4

jithin-sabu / purge-app

Swift★ 825⑂ 46▲ 29 stars
5

obra / superpowers

Shell★ 289,686⑂ 25,927▲ 544 stars
6

mattpocock / skills

Shell★ 267,091⑂ 22,555▲ 695 stars
7

ossu / computer-science

HTML★ 209,332⑂ 25,890▲ 108 stars
8

ohmyzsh / ohmyzsh

Shell★ 189,870⑂ 27,279▲ 29 stars

More Trending Repositories