I use AI for a few hours on most days now. Somewhere in the last year, the bottleneck stopped being the model and started being me. Specifically, my hands.
I was typing out long, carefully structured prompts, and by the third or fourth one I could feel myself cutting corners just to be done with the typing. That is a strange problem to have. The intelligence is instant and effectively free, and the slow part is a keyboard.
So I went looking for a way around it.
Why I skipped Wispr Flow and gave up on Mac dictation
Wispr Flow is the one everybody points to, and by all accounts it works well. It also runs about $12 a month on an annual plan and sends your audio to a server.
I am not against paying for tools. I pay for plenty. But I had a hunch that this particular feature was not going to stay premium for very long. Speech is the most natural input we have. Things like that get absorbed into the operating system eventually. They do not stay subscriptions. Paying a yearly fee for something I expect to be free in two years felt like the wrong bet.
So I tried the dictation that already ships with macOS. I would call it 50% useful. Individual words were mostly fine, but it kept guessing wrong whenever it was unsure, and it fell apart at the sentence level. Punctuation landed in odd places, clauses ran into each other, and I spent more time cleaning up than I saved. Fine for a two-line reminder. Useless for dictating an actual paragraph of thinking.
A free, fully local alternative to Wispr Flow
I came across an app called Onit, which has since rebranded to Cloudless Voice. It is a free voice dictation app that runs entirely on your own device, works in any text box across the system, and is positioned directly as an offline alternative to Wispr Flow.
The thing that got my attention was the architecture. It runs a two-stage pipeline entirely on your machine. NVIDIA’s Parakeet, an open speech recognition model, handles the transcription. Then a second small local model cleans it up: filler words, punctuation, lists, formatting. Nothing gets uploaded. You can turn off your wifi and it still works.
It was not smooth at the start. Some of the early updates were buggy, a few were slow, and one or two broke things that had been working. I stuck with it because the idea was right even when the execution was not. Over the last two months, it has improved drastically. It is now part of how I work every day.
They are giving free lifetime access to the first 5,000 sign-ups. As of August 9, 2026, 2,208 seats had been claimed — the counter is live on their site.
Local-first, not open source
When I first explained this choice to someone, I said I preferred open source. That is not quite accurate, and it is worth correcting, because Cloudless’s dictation is not fully open source either.
What I actually want is local-first. Those are two different things. Open source is about whether I can read the code. Local-first is about whether my voice leaves my machine. For something I use a few hundred times a day, across drafts, client notes and half-formed ideas I have not decided to share yet, the second question is the one I care about.
The part I did not expect
Every major LLM has voice built in now. Claude, Gemini, ChatGPT, all of them.
I would still rather own the input layer. Dictation that works in every text box works in my email, my CMS, my terminal, my notes. One habit, everywhere. Voice inside a single app is a feature of that app. Voice at the system level changes how you work.
But the bigger change was not about speed at all. It was in what I started getting back.
When typing is the bottleneck, you write efficient prompts. You compress. You strip out context because context costs keystrokes. Looking back, I think a lot of prompt engineering was really a workaround for a slow input device. We kept things short because typing was expensive, and then we built a whole discipline around it and gave it a name.
Take the cost away and you stop compressing.
Here is what convinced me. I have had a backlog of business ideas for years, and I had eventually accepted that the hard part is not having them, it is validating them honestly. I even built a small tool with ChatGPT to pressure-test them. It helped, but it had a ceiling, because I was still feeding it tidy summaries. It was only ever reacting to my own framing of the problem.
When Claude’s Fable model launched, I tried something different. I talked at it for five straight minutes. No structure at all. What I was thinking about, where the idea came from, what I had already tried, what was actually bothering me about it. Rambling, in the way you would ramble to a friend who happens to be sharp.
It came back with a breakdown across viability, feasibility and financial risk, along with a few angles I had not thought to ask about. I never requested that framework. It built the structure itself, because it finally had enough context to see the shape of what I was circling around.
That is the shift. You stop instructing the model and start briefing it. It is closer to how you would approach a guru you have gone to with a problem. You do not hand them a formatted brief. You tell them the whole messy situation and trust them to find the thread.
Worth a try
A few honest caveats. I have only used it on a Mac. They have since added Windows and iOS, and I cannot tell you how well those hold up. It has been buggy in the past. And it is a small team, so support is a Discord channel rather than a helpdesk.
If you want to try it, here is my referral link (I get credit if you sign up through it): cloudless.so
My actual bet is that none of us will be paying for dictation in two years, because it will simply be part of the operating system, the way spell-check is. The interesting question was never which app to use. It is what happens to your thinking when the friction of getting it out of your head disappears.