audio Tools
216 best audio tools and apps, curated and ranked by community upvotes on ProductCool. Updated daily as new audio products launch.
Nearfield is a native, open-source Mac app that combines the speakers in two Apple Studio Displays into one volume-controllable stereo output. Swap left and right, adjust balance, and use app and window based audio routing. Requires Apple silicon, two Studio Displays, and macOS 14+.
Closed voice platforms make you rent your own agents. Dograh is completely open source- nothing is gated. Visual flow builder, add your model key across 30+ integrations or use local models, telephony, human transfer, and advanced QA & monitoring - all free to self-host in one command. Also connect your claude code with MCP to build voice agents for a use case or call recordings.
Transformers is a comprehensive library providing thousands of pre-trained models for Natural Language Processing (NLP), computer vision, audio, and multimodal tasks. It solves the problem of implementing and deploying cutting-edge machine learning models by offering a unified, easy-to-use API for both inference and training. It is designed for researchers, developers, and practitioners who want to leverage state-of-the-art models without building them from scratch, enabling rapid prototyping and production deployment.
Instant dictation for desktop. Press a shortcut, speak, and instantly get accurate text on your clipboard—perfect for emails, coding, AI prompts, or brain dumps.
The App Store for your voice, where building takes one sentence. Describe what you want "Create an app to help me to journal" and VoiceOS creates the app for you. Share it as a link and anyone can install it in a click.
SpeakoFlow puts your voice over your whole desktop. Speak, and your words land in any app — email, editor, chat, terminal. Say "Hey Flow" and it writes the whole reply from what's on your screen. Ask the assistant about what you're looking at and hear the answer back. It also cleans up your dictation, translates as you speak, and learns how you work. Everything can run on your machine — speech-to-text always does. Free, open source, MIT. Windows, macOS, Linux.
Wondering is the most delightful way to break down complex topics into knowledge you can remember and apply. Tell it what you want to learn, and it creates a personalized path of short lessons with visuals, podcasts, and interactive exercises. It surfaces the most important ideas and lets you explore them in your own way, like a thoughtful tutor beside you.
Every meeting recorder ships your audio to the cloud and bills you monthly. yapyap does neither. It records, transcribes, names your speakers, and turns talk into summaries and action items entirely on your own machine. “Lenses” reshape each recording into whatever you need: a summary, a to-do list, or a decision log. You can install more or build your own. Prefer the cloud? Optionally connect any major provider: OpenAI, Anthropic, Groq. Either way, yapyap is yours forever. No subscription.
Cekura is the testing, observability, and self-improvement platform for production voice and chat AI agents. It simulates thousands of scenarios, catches failures, diagnoses the root cause, rewrites prompts and config, then re-validates with a full regression sweep. Unlike tools that hand failures back to your team, Cekura closes the loop by fixing the agent itself and proving the fix holds without overfitting.
Browser FX transforms your web audio experience with real-time professional audio effects. Capture any tab's audio and shape it with studio-style knobs in a sleek dark interface, complete with an audio-reactive cymatic visualizer, and even hands-on control from your MIDI controller.
Highlight any article, newsletter, blog post, or page on the internet. Liso instantly turns it into beautiful audio. Your personal audiobook, built from the things you actually want to read.
Wispro turns your voice into writing, instantly. Just talk, messy or unfiltered, and Wispro pastes clean, ready-to-use text directly into whatever you're working on. It adapts to how you think, not just how you speak. Basic Mode captures every word exactly as said. Smart Mode cuts filler words and rambling, formatting it into polished prose. Command Mode turns a spoken instruction into finished writing, emails, replies, whole drafts, on the spot. One voice. Three ways to write.
Introducing Universal Dictation on Stream. Push-to-talk across iOS and Mac: Instantly. No app switching, no reconnection. We built Notes so nothing gets lost, and Chat so you can think out loud. With Dictation, your voice goes anywhere. Notes, Chat, Dictation—in one voice ring
Routine AI lets you control your tasks, calendar, notes and projects using your voice. Just talk naturally to schedule meetings, add reminders, capture ideas, write notes, search your knowledge, update projects and automate repetitive work.
Drop in a full mix or an isolated drum stem. Backbeat Forge puts the performance on a five-line drum score you can check against playback and correct by hand. When the chart feels right, export it as a printable PDF or General MIDI file. It all runs locally.
River enables B2B companies to sell with VoiceAI. When a lead enquires, our AI account executive joins a live call instantly, runs the product demo, handles objections, and closes - so no lead ever waits for a rep's calendar. Backed by founders of Ramp, Kalshi, and Lean.
SoundPipe creates virtual audio devices on your Mac so you can send audio from any app, or your microphone, to any other app.
GPT-Live is OpenAI’s new full-duplex voice model for ChatGPT Voice. It can listen and speak at the same time, handle pauses and interruptions more naturally, and delegate harder search or reasoning work to frontier models in the background.
Ellis is an AI notetaker for in-person meetings. Record your meeting, get a clean transcript with each speaker identified, then ask anything — what was decided, what you missed, how it went. No laptop. No extra hardware. Just your iPhone (or Apple Watch).
Cadence is the only screen recorder that effortlessly perfects your video the moment you hit stop. The AI instantly transforms accents, applies clear voice dubbing, removes background noise, and automatically builds transcripts and screenshots for you.
Nada is the easiest way to turn your voice into music. Hum, sing, or whistle an idea, and instantly convert it into MIDI. Arrange your melodies on mobile, choose from a variety of instruments, and build songs wherever inspiration strikes. No AI generation involved, just your own creativity, captured faster.