Spokenly and Dictato both turn speech into text on your Mac. Spokenly is free for local dictation, with a Pro plan for cloud models and hosted AI cleanup. Dictato is 19.99€ once, everything included. The useful question is not whether Spokenly is free (it is), but what happens to your words after the transcription, once the “um”, the false starts and the mid-sentence corrections are in the text.
What Spokenly gives you for free
Spokenly’s default engines are cloud models. Pick a local one instead (Whisper, Parakeet or Apple’s recognizer) and dictation runs on your Mac with no account, no usage cap, and a Local Only Mode that blocks every network connection. With some engines the words show up live while you talk.
Text cleanup lives in “modes”: grammar, filler removal, translation, your own instructions. Modes are free to create, but each needs a language model behind it, and you write the instruction yourself. Four routes: Spokenly’s hosted AI (Pro), Apple Intelligence on macOS 26, your own API key from OpenAI, Anthropic or any compatible provider, or a model server you run on the Mac. Spokenly’s own comparison pages list “stammer correction” as not included.
Pricing
- Free: unlimited local dictation, modes with a model you provide.
- Pro: $9.99 a month or $99.99 a year. Managed cloud transcription, hosted AI cleanup that works out of the box, sync across devices. No lifetime option, 14-day refund.
What Dictato gives you for 19.99€
Hold (or tap) a shortcut, speak, release. Dictato transcribes on your Mac with one of its on-device engines and drops the text at your cursor; the transcription step takes about 80 ms*. Text arrives when you release the shortcut, not while you speak.
Auto-correct is where Dictato puts its effort, and it never leaves the Mac. On macOS 26 it can use Apple Intelligence. On any Apple Silicon Mac, switch on Profiles and Dictato downloads a small open model once and runs the cleanup with it; that is the path tuned for hesitations, false starts and numbers. No key, no server, no separate bill. Both are off when you install the app, and each dictation takes about a second longer with them on.
Translation into 31 languages (macOS 15 or later), unlimited history, per-app profiles and voice commands are all on-device and included.
Pricing
19.99€ once, yours for life, updates included. Seven-day trial, then a license key by email, no account. 14-day refund. macOS 14 or later, Apple Silicon only.
Spokenly vs Dictato: side by side
| Spokenly | Dictato | |
|---|---|---|
| Price | Free; Pro $9.99/mo or $99.99/yr | 19.99€ once, lifetime |
| Free tier | Yes, unlimited local dictation | 7-day trial |
| Default processing | Cloud models; local is an option | Local only |
| Local engines | Whisper, Parakeet, Apple’s recognizer, and more | Whisper, Parakeet and two more |
| Cloud engines | Yes (Pro or your own key) | No |
| AI cleanup | Your instruction; model from Pro, your key, Apple Intelligence or one you host | Built in, on your Mac, included (off by default) |
| Hesitations, false starts | Not built in; write an instruction | Handled by the bundled model |
| Numbers, dates, amounts as figures | Depends on your instruction | Built in |
| Translation | Through a mode (needs a model) | 31 languages, on-device, macOS 15+ |
| Live text while speaking | With some engines | No, text arrives on release |
| History | Yes, searchable | Yes, unlimited, searchable |
| File transcription | Yes | No |
| Fully offline, AI included | Only with Apple Intelligence or a model you host | Yes |
| Platforms | Mac, iPhone, Windows, Linux | Mac only |
| Requirements | macOS 13.3+, partial Intel support | macOS 14+, Apple Silicon |
| Refund | 14 days | 14 days |
The real difference: what happens after you stop talking
Both apps run the same families of local engines, so the raw transcript is close. The gap is the cleanup.
Dictate a sentence the way people speak: “send it Tuesday, no wait, Wednesday morning, um, around ten.” A raw transcript keeps every word. A good cleanup pass returns “Send it Wednesday morning, around 10”: punctuation, capitals, the hesitation gone, the correction applied, the number written as a figure.
Spokenly can run this pass, and with a strong cloud model it runs it well. But the free tier hands you the tools, not the result: you pick a model, you write the instruction, you test it. Dictato ships the model and the instruction, and runs both on the Mac. So the test worth doing during the trial: does the text still hold up when you hesitate, repeat yourself, or change your mind mid-sentence?
Where Spokenly wins
Free with no clock. Local dictation costs nothing. If a raw transcript is all you need, install it and stop reading.
Your choice of model. With your own key, Spokenly can transcribe and clean up with cloud models from OpenAI, Deepgram, Groq and others. Dictato has no cloud path at all.
Live text, more platforms. Some engines show words while you are still speaking, and the app also runs on iPhone, Windows and Linux, with settings sync in Pro.
Intel Macs and audio files. Spokenly’s Whisper engine still runs on Intel, and Spokenly transcribes audio files. Dictato does neither.
Where Dictato wins
Cleanup included, and nothing leaves the Mac. Auto-correct, translation and history run on-device with no key, no server and no extra bill.
One price, and it is the last one. 19.99€ once, against $99.99 every year for the Pro tier that works out of the box. See the real cost of speech-to-text over time.
Translation without a setup. Speak French, get English, in 31 languages, on-device. In Spokenly, translation is a mode, and a mode needs a model.
Who should use what
Spokenly if you want free local dictation and do not need cleanup, already pay for a cloud AI key, dictate on Windows or an iPhone too, or are on an Intel Mac.
Dictato if you want dictated text that reads like you typed it, with no subscription, key or model server, translate as you dictate, and would rather pay once. Our best dictation app for Mac in 2026 roundup covers the wider field.
Bottom line
Spokenly’s free tier does what it says. Dictato charges 19.99€ once and puts the cleanup model in the box. If you talk the way people talk, with pauses and second thoughts, the cleanup is the feature, and what it takes to get it is the comparison that matters. Related: Dictato vs Voibe and Dictato vs EmberType.