MLX Whisper
Private, offline, on-device
Transcribe on the Apple Silicon GPU with local MLX models. Audio stays on your Mac and the engine works without internet.
- Local
- MLX
- Offline
Dictate anywhere on macOS. Run MLX Whisper on your Apple Silicon Mac, offload batch audio to your own Local AI Server, or use Deepgram and ElevenLabs in batch or realtime. Then optionally polish the transcript with your AI provider and paste it into whatever you're writing.
A simulation of the real SapoWhisper flow: recording overlay, transcription, AI improvement, and auto-paste. Pick a mode and run it.
In the real app this starts with your hotkey, a combination or a double press of a modifier. Try it here: press Run, or double-tap ⌥ on your keyboard.
Stay fully on-device, use your own speech server, or bring a cloud API key for batch accuracy and realtime streaming. Switch without leaving the app.
Private, offline, on-device
Transcribe on the Apple Silicon GPU with local MLX models. Audio stays on your Mac and the engine works without internet.
Your OpenAI-compatible endpoint
Offload batch transcription to a self-hosted speech-to-text server. Configure its URL, model, and optional bearer token in SapoWhisper.
Nova-3 batch or Flux Live realtime
Choose Nova-3 for high-accuracy batch transcription or Flux Live for low-latency streaming. Bring your own API key.
ElevenLabs Scribe v2 batch or ElevenLabs Scribe Realtime v2
Use Scribe v2 for accurate batch transcription or Scribe Realtime v2 for low-latency committed-text streaming. Bring your own API key.
Each transcript runs through a small, fast model from OpenRouter, OpenAI, Groq, or your own endpoint. Your API key stays in the macOS Keychain.
Model: openai/gpt-5.4-nano
uh send to the team like the latest mockups before lunch i think and also push the release branch to github
Send to the team the latest mockups before lunch, and also push the release branch to GitHub.
Fixes likely errors, drops filler words and improves readability while staying literal: your words, your order. A fidelity guard restores the original text if the AI drifts.
Speak in any language and get the final text in the one you choose. Spanish in, English out, no extra step.
Editable per-project modes for Codex, Claude Code, Slack or issues, plus custom vocabulary and personal context applied to every mode.
Engines and AI handle the transcript. These are the everyday touches that make SapoWhisper feel like part of your Mac, not a separate tool.
Trigger SapoWhisper from any app with a global hotkey, a key combination or a double-tap of a modifier, and the transcript lands at your cursor via auto-paste.
Mic test, input gain controls, sound feedback and auto-ducking that lowers system volume while you talk.
Swap the interface and the transcription language between Spanish, English or Auto, and recognise both without restarting.
Export and import your full configuration as a portable JSON file. Free and open source, no account required.
SapoWhisper keeps working after the paste. Search past transcripts, replay the original audio, pin important entries, and run the same clip through another engine when you want a better result.
That beat there crushes it for me like, I don't know, like a thousand hours in the editing. Really, because the editing, in one day I took like five hours, six hours. And you've done it right now literally in five minutes, ten minutes, and you've done it super well, almost like surpassing the level of what I was doing. So that's perfect.
Find older transcripts quickly instead of dictating the same thing twice.
Listen back to the saved audio or run it through another engine.
Re-polish any entry with your AI provider when you need it.
Pick the transcription engine that fits the moment, let your AI provider polish the result, and start writing with your voice from anywhere on macOS.