🚀 Maximize your product's SEO. Submit to 240+ directories in 1-click with DirSubmit. Launch Now
Speech To Markdown logo

Speech To Markdown

Harness local AI for notes

2026-07-25

Product Introduction

  1. Definition: Speech-to-Markdown is a native macOS menu-bar application and a separate iOS app that functions as a fully local, AI-powered dictation and document formatting tool. Technically, it is a voice-to-text transcription and natural language processing (NLP) application that leverages on-device models to convert spoken audio into structured text formats like Markdown, HTML, or plain text.
  2. Core Value Proposition: It exists to provide a privacy-first, offline-capable solution for professionals who need to transform unstructured speech into clean, formatted documentation instantly. Its primary value is enabling global dictation and live markdown editing powered by a local LLM, ensuring 100% local processing with no cloud dependency and no API keys.

Main Features

  1. Global Dictation Hotkey (⌘⌥]): This feature allows users to trigger dictation from any application on macOS. When activated, a Spotlight-style indicator appears, and speech is transcribed locally using whisper.cpp. The resulting text is typed directly at the cursor's location in the active window, whether it's a terminal, browser, Slack, or code editor. This enables seamless voice input across the entire operating system.
  2. Agent Mode with Real-Time LLM Streaming: This is the core editing environment. Users speak freely into a dedicated window, and their audio is transcribed in 4-second chunks by whisper.cpp. The transcribed text is buffered and, upon reaching a threshold (30 words), a pause (5 seconds), or a manual "Send" command, is sent to a user-configured local LLM server (e.g., omlx, Ollama). The LLM then streams formatted text (Markdown, TXT, HTML) back into a live editor in real-time, continuously improving the document structure.
  3. Multi-Mode Processing (Format/Edit/Append): This feature provides granular control over how the local LLM interacts with the text. Format Mode sends the entire document plus new words for a complete rewrite. Edit Mode allows voice commands to instruct the LLM to modify a selected portion of text (e.g., "turn that list into a table"). Append Mode optimizes for long documents by sending only the last few sentences plus new words, keeping latency constant and appending the newly formatted output.
  4. Cross-Platform Offline Operation: The product consists of two distinct applications. The macOS app relies on user-managed local dependencies (whisper.cpp, a local LLM server). The iOS app (for iPhone 15 Pro/iOS 26+) requires zero setup, using Apple's on-device SpeechAnalyzer for transcription and Apple Intelligence Foundation Models for formatting, making it 100% offline and airplane-mode friendly.

Problems Solved

  1. Pain Point: The inefficiency and lack of structure in traditional dictation. Standard voice-to-text outputs a wall of plain text, requiring manual editing to add formatting, lists, headers, and punctuation—a time-consuming post-processing task.
  2. Target Audience: Technical writers, journalists, researchers, developers documenting code, students taking lecture notes, and any knowledge worker who needs to produce structured notes, drafts, or documentation quickly without breaking their workflow.
  3. Use Cases: Documentation Drafting: Speaking meeting notes that are instantly formatted into Markdown with headers and bullet points. Content Creation: Dictating blog post drafts or email responses that are immediately structured. Hands-Free Coding Notes: Using the global hotkey to voice comments or documentation directly into a code editor. Mobile Note-Taking: Using the iOS app to capture and format ideas on the go with complete privacy.

Unique Advantages

  1. Differentiation: Unlike cloud-based dictation services (like Otter.ai or Google Docs Voice Typing) or assistants (Siri), Speech-to-Markdown processes all audio and language model inference locally. It differs from other local transcription tools by deeply integrating a customizable local LLM for real-time formatting and structuring, not just transcription.
  2. Key Innovation: The integration of a low-latency, chunk-based whisper.cpp transcription pipeline with a streaming local LLM interface (using the OpenAI-compatible API format). This creates a live feedback loop where speech is continuously turned into structured text. The Append Mode is a specific technical innovation to maintain performance with arbitrarily long documents, a common issue for context-window-limited LLMs.

Frequently Asked Questions (FAQ)

  1. Is Speech-to-Markdown really free and fully offline? Yes, the application is completely free. On macOS, it runs 100% locally if you provide the open-source models (whisper.cpp for speech, a local LLM server like Ollama for text). On supported iOS devices, it uses entirely on-device Apple Intelligence models, requiring no internet connection.
  2. What are the system requirements for the macOS version? You need a Mac with Homebrew to run the installer script. The app itself requires microphone and accessibility permissions. For optimal Agent Mode performance, a local LLM server (e.g., omlx, Ollama) must be running. More powerful Macs (especially Apple Silicon) can run larger, more capable LLMs (like Qwen3.5 27B 4-bit) for better formatting quality.
  3. How does the iOS app work without installing whisper.cpp or an LLM? The iOS app is a separate build that leverages Apple's private, on-device APIs: SpeechAnalyzer for speech-to-text transcription and Apple Intelligence Foundation Models for text structuring and formatting. This requires a compatible device (iPhone 15 Pro or newer) running iOS 26+.
  4. Can I use OpenAI's GPT-4 or another cloud API with this tool? No. The application is explicitly designed for local LLM servers that use the OpenAI API request format (a common local server standard). It does not send data to OpenAI's servers or any other cloud API. The architecture is built for privacy and offline use.
  5. What is the difference between Format, Edit, and Append modes? Format is for free-form dictation, where the LLM continually refines the entire document. Edit is for giving voice instructions to change specific selected text. Append is a performance-optimized mode for long documents where only the newest text is processed, preventing slowdowns and keeping latency low.

Submit to 240+ Directories with 1-Click

Maximize your product's SEO and drive massive traffic by automatically submitting it to over 240 curated startup directories using DirSubmit.

Related Products

Subscribe to Our Newsletter

Get weekly curated tool recommendations and stay updated with the latest product news