Product Introduction
- Definition: Speek is a native macOS voice assistant application and dictation tool that integrates directly into the Mac's notch. It is a technical productivity application that leverages on-device speech recognition (via models like NVIDIA Parakeet/Canary and OpenAI Whisper) and large language models (via OpenAI, OpenRouter, or local Codex) to execute commands and automate workflows.
- Core Value Proposition: Speek exists to eliminate context-switching and manual input friction on macOS. Its primary value is enabling hands-free Mac operation, voice-driven task automation, and context-aware dictation directly into any active application, allowing users to interact with their computer and connected services through natural speech without interrupting their workflow.
Main Features
- Notch-Integrated Voice Interface: The assistant resides visually within the MacBook Pro's notch, providing a persistent, low-profile UI. It uses a system-wide keyboard shortcut (default: Option-Space) for push-to-talk and supports configurable voice activation ("Hey Speek") with on-device wake-word detection. This allows for instant, system-level invocation from any app or screen state.
- Context-Aware Dictation & Text Manipulation: Beyond basic speech-to-text, Speek's dictation engine analyzes the context around the text cursor and the active application to format output appropriately. It offers raw, cleaned, or polished writing styles. A key technical feature is the ability to edit selected text by voice command (e.g., "make this shorter," "translate to French") by passing the selected content and the vocal instruction to the configured LLM for processing.
- Screen Context & Pointer Integration: For commands requiring visual reference, Speek can capture screen content. Users can circle UI elements with the pointer while holding the activation key to provide precise spatial context for commands like "click this" or "what is this?". This combines screen capture technology with pointer event tracking to understand user intent relative to the graphical interface.
- Modular Integration via MCP (Model Context Protocol): Speek connects to external services and tools not through a closed ecosystem, but via the open Model Context Protocol. This allows it to integrate with hundreds of tools (Gmail, Google Calendar, GitHub, Notion, Linear) and local CLI tools through MCP servers. Users manage granular permissions per integration, choosing "ask," "auto," or "never" for each action.
- Background Task Orchestration & Computer Use: For complex, multi-step tasks, Speek can orchestrate background automation sequences. Using macOS Accessibility APIs (AXSwift) and synthetic input (KeySender), it can perform "computer use" – programmatically clicking, typing, and navigating within native Mac applications (Calendar, Mail, Notes, etc.) to complete tasks like creating an event or filing an email, then notifying the user upon completion.
Problems Solved
- Pain Point: Workflow Interruption and Context Switching. Traditional assistants require switching to a separate app or window. Speek solves this by being omnipresent in the notch, allowing users to issue commands or dictate text without leaving their current application, preserving focus and flow state.
- Target Audience: Power Users and Knowledge Workers on macOS. This includes developers, writers, project managers, and executives who perform repetitive computer tasks, manage communications across multiple apps (Slack, Mail, Calendar), and seek efficiency gains through automation and voice control without sacrificing privacy or control.
- Use Cases: Essential for multi-modal productivity scenarios: Dictating long-form content directly into a code editor or document; quickly checking calendar availability or creating a reminder while in a focused deep work session; using voice to control music/media during presentations; performing complex app-specific actions (e.g., "file this email to the Project X folder") via verbal command while reading.
Unique Advantages
- Differentiation: Unlike cloud-first assistants (Siri, Google Assistant), Speek prioritizes on-device processing for wake-word and speech recognition where possible, enhancing privacy. Unlike pure dictation software, it is an action-oriented automation engine. Compared to other macOS automation tools (Keyboard Maestro, Alfred), it uses natural language as the primary input method, reducing the need to memorize keyboard shortcuts or build complex macros.
- Key Innovation: The "circle to point" gesture combined with screen context is a unique HCI innovation for voice assistants. It solves the inherent ambiguity of referring to screen elements (e.g., "that button") by allowing precise spatial disambiguation. Furthermore, its use of the open Model Context Protocol for integrations future-proofs it against API changes and grants users unprecedented control and extensibility compared to walled-garden assistants.
Frequently Asked Questions (FAQ)
- Does Speek work offline or require an internet connection? Speek uses on-device models for wake-word detection and speech-to-text (STT), allowing for offline dictation and activation. However, for natural language understanding, command execution, and complex task reasoning, it requires an internet connection to communicate with your configured LLM provider (OpenAI, OpenRouter, etc.).
- How does Speek handle privacy and data security for voice commands? Speech recognition for activation and dictation can run locally on your Mac. When using cloud LLMs, your transcriptions and requests are sent to your chosen provider (OpenAI/OpenRouter). Integration credentials (OAuth tokens) and personal memory facts are stored locally in the macOS Keychain and app sandbox. You maintain granular control over which integrations can act automatically.
- Can I use Speek with applications beyond the built-in integrations? Yes, through two primary methods. First, via MCP servers, you can connect to virtually any tool with an API. Second, using the "computer use" feature, Speek can learn to operate any macOS application on-screen via Accessibility APIs, though this may require user guidance for complex layouts.
- What are the system requirements for running Speek? Speek requires macOS 26 (or later) and a Mac with a notch (or the menu bar for non-notch Macs). It requires granting Accessibility, Screen Recording, and Microphone permissions. For optimal performance of local speech models, a Mac with Apple Silicon (M-series) is recommended.
- How does Speek's dictation compare to macOS's built-in dictation? Speek offers superior context-awareness, adapting tone and format based on the target app. It provides real-time transcription preview in the notch, supports voice-driven text editing of selections, and features a learning vocabulary that improves from user corrections. It is designed as a power-user tool for writing and command, not just basic text entry.
