The reason to click: the test set isn't the usual read speech. It's 10 hours of real dictation from over 2,300 people using various mics in their actual environments. Canto posted the lowest word error rate, ahead of Google, OpenAI, AssemblyAI, and Deepgram. On a 3-hour challenge set built to stress-test models, it led among real-time models but trailed Gemini 3.1 Pro overall—a large multimodal model that isn't built for low-latency use.
I'd discount this a bit. The eval is Wispr's own, built from its user base. They enforced speaker separation between train and test, but the app distribution and labeling standards aren't reproducible from the outside. On public benchmarks, Canto tied for first only on LibriSpeech; it didn't lead on FLEURS or Common Voice, which the team acknowledges are read-speech datasets that differ from spontaneous dictation.
The training recipe is worth noting: pretraining on millions of hours of speech and text, then supervised fine-tuning plus GRPO reinforcement learning to optimize full-transcript quality. But the post doesn't give model size, inference latency, or pricing—the details that determine whether this actually works in production. For now, this reads more like a technical credential for Wispr Flow than a general-purpose speech model ranking.