How does this compare with Parakeet? I've been using that locally in my projects on an M-series macbook and it's been working great. It's fast and accurate enough for my use cases (meeting transcription, audio transcription for demo videos, etc.)

This definitely seems lighter and faster. How does accuracy compare?

People keep praising Parakeet, but I've found it to be worse than Whisper Large. Yes, it is much, much faster and smaller, but accuracy matters a lot if you are to use dictation regularly and seriously.

I ended up having AI optimize Whisper Large and create a plugin for TypeWhisper, and that's what I use (feeding the results through local Qwen 3.8 running under MTPLX).

I confirm this as my experience as well. Parakeet did not perform well for me (English with heavy accent). Whisper large is much better. But once we get to model of this size look at qwen3-asr (1.7b parameters). I'm quite happy with how it performs (both transcription quality and speed on rtx5080). By the way the model is also available on openrouter and very/very cheap. Performance is reasonable (<1s response), but some requests are delayed (>10s)

When you are running the model locally and trying to do other tasks at the same time, Whisper Large feels too clunky on normal hardware.

I like Parakeet because it's good enough, relatively light, and fairly fast. I'm using this for things like meeting transcription and dictation. Since I'm sending most of the text to an LLM to clean up afterwards, it works well enough.

I'm hoping for something the size of Parakeet (or smaller) but better quality. It feels like with all of the advances in making smaller models better in the LLM space, someone should be able to come up with a lightweight and better quality speech to text model.

For English I've found Parakeet v2 and Parakeet unified to perform significantly better than v3. For deposition audio in a quiet room and neutral american accents I've hit 96-97% on a number of sessions with v2 and I've seen as low as 83% for heavy accents with overlapping speakers.

It’s really domain dependent I think. If you are doing anything conversational interfacing with less AI-familiar users, latency matters a lot.

If you’re feeding the results into a very smart LLM, it will figure out what you meant (but crucially ONLY if you warn it or tell it to do so, in some cases!). If you’re writing code directly or creating something for public consumption, you can’t tolerate mistakes. If you’re taking notes for yourself you just want it to work cheaply.

If you are ok with the complexity you can run both, and a Meta/Google open model with native audio, and let a smart LLM doctor it up. If performance really matters you can pay for a proprietary model or train one yourself. Until you get to that point, I think it probably doesn't matter much either way. Voice is just too easy to fiddle with

Parakeet is the gold standard. With models like moonshine and koroko (TTS model), it’s more about embedding the model in the application itself. If you’re using Parakeet, embedding it in the application is not feasible.

I use parakeet with superwhisper, and I’m making another app that has SST and TTS built in, and I want to use my downloaded parakeet model, but it seems there’s so many different implementations from ONNX to whisper, it’s not easy to use your downloaded models. So models like moonshine and this one allow you to just embed it into your application simply. It might not be as good as parakeet, but it gets you 80% of the way there.