Product Introduction
- Definition: FluidVoice is a local-first, open-source speech-to-text (STT) dictation application for macOS. It is a technical tool that performs all audio capture, speech recognition, and AI-powered text post-processing directly on the user's device.
- Core Value Proposition: It exists to provide Mac users with a private, fast, and intelligent dictation solution that transforms natural, unpolished speech into ready-to-use text without ever sending voice data to the cloud. Its primary value is on-device privacy, real-time transcription speed, and context-aware AI polishing via its custom Fluid-1 model.
Main Features
- 100% On-Device Processing: All speech recognition occurs locally using selected models. Audio from the microphone is processed directly on the Mac's hardware (CPU/GPU), ensuring zero bytes of voice data or raw transcripts are transmitted to external servers. This is enabled by leveraging CoreML and Metal for Apple Silicon optimization and supporting local model frameworks.
- Fluid Intelligence with Fluid-1 Model: This is a custom-trained, local AI model (~3.5 GB download) that acts as a post-processor. It takes the raw transcription and performs intelligent formatting: removing filler words (ums, ahs), correcting capitalization, applying proper punctuation, and adapting the tone to match the target application (e.g., casual for Slack, formal for Mail). It runs entirely on-device, requiring no API keys.
- System-Wide Input with Adaptive Tone: Activated by a single global hotkey, FluidVoice can dictate into any text field across the macOS system—including terminals, code editors, email clients, and note-taking apps. It automatically detects the focused application and applies a user-customizable tone profile, ensuring the output text is contextually appropriate.
- Multi-Model, Multi-Language Support: The app supports a range of local speech recognition models, each with different characteristics: Nemotron Speech 3.5 (Ultra Fast Low Latency & Multilingual, ~40 languages), Parakeet Flash/TDT models (English-focused, low latency), Cohere Transcribe, Apple Speech, and Whisper variants (up to 99 languages). Users can select the model that best balances their needed language support, speed, and accuracy.
- Optional Cloud AI Provider Integration: While designed as a local-first tool, users have the optional flexibility to route the post-processing step to external AI providers like OpenAI or Groq for rewriting, instead of using the local Fluid-1 model. This maintains user control over the privacy-performance trade-off.
Problems Solved
- Pain Point: Privacy concerns with cloud-based dictation. Traditional services like Google Dictate or Apple's enhanced dictation send audio to remote servers, creating data security and privacy risks. FluidVoice solves this by guaranteeing zero data exfiltration.
- Pain Point: The inefficiency of raw transcription. Natural speech is messy, filled with hesitations and incomplete thoughts, resulting in transcripts that require significant manual editing. FluidVoice's Fluid-1 model directly solves this by delivering polished, ready-to-send text.
- Target Audience: Privacy-conscious professionals (developers, writers, lawyers, healthcare workers), multilingual users and teams, developers and technical users who work in terminals/IDEs, and anyone seeking a faster, hands-free text input method on Mac without subscription fees.
- Use Cases: Dictating long-form emails or documents with proper formatting; writing code comments or commands in a terminal via voice; communicating in team chat apps (Slack, Discord) without typing; taking notes during meetings; creating content in multiple languages; working in environments with no or limited internet connectivity.
Unique Advantages
- Differentiation vs. Cloud Competitors: Unlike Wispr, Rewind AI, or cloud-dependent services, FluidVoice's core transcription pipeline is permanently offline-capable and collects no voice data. Versus built-in macOS dictation, it offers superior, AI-powered text polishing and app-specific tone adaptation.
- Differentiation vs. Other Local Tools: Compared to other local STT options, its integration of the custom Fluid-1 model for context-aware rewriting is a unique value-add. The combination of system-wide input, multi-model support, and adaptive post-processing in one free, open-source package is currently distinctive.
- Key Innovation: The development and integration of the Fluid-1 model, a custom-trained, locally-runnable Large Language Model (LLM) specifically fine-tuned for dictation post-processing. This allows for sophisticated text normalization and tone matching without any cloud dependency, which is a significant technical achievement in the local AI application space.
Frequently Asked Questions (FAQ)
- Is FluidVoice really free and what is the license? Yes, FluidVoice is completely free forever. It is open-source software, and as of February 23, 2026, it is licensed under the GNU General Public License v3.0 (GPLv3).
- Does FluidVoice work without an internet connection? Yes, the core speech-to-text dictation functionality works entirely offline using the local speech models (e.g., Whisper, Nemotron). The optional Fluid-1 AI post-processing model also runs 100% on-device, requiring no internet. Internet is only needed if you choose to use optional cloud AI providers like OpenAI for post-processing.
- What are the system requirements for FluidVoice? FluidVoice requires macOS 15.0 (Sequoia) or later. It supports both Apple Silicon (M-series) and Intel Macs. It requires microphone access and Accessibility permissions to type into other applications.
- How does FluidVoice's accuracy and speed compare to cloud services? FluidVoice achieves a perceived latency of under 100ms for real-time transcription, making it feel instantaneous. Accuracy is model-dependent; the supported Nemotron and Parakeet models are state-of-the-art for local inference. While cloud services may have marginal accuracy benefits from larger models, FluidVoice provides excellent accuracy with the decisive advantage of complete privacy and no network latency.
- What is the Fluid-1 model and do I have to use it? Fluid-1 is FluidVoice's optional, custom-trained local AI model that rewrites and polishes raw transcriptions. It is not required. You can use FluidVoice for raw, fast transcription without it, or you can enable it to get intelligently formatted and tone-adapted text. It is a separate ~3.5 GB download.