Wispr introduces Canto: a real-time speech model built for real-world dictation
Canto: A speech model built for the real world
Wispr released Canto, a real-time speech model that achieved the lowest word error rate on 10 hours of real-world dictations from over 2,300 speakers, beating models from Google, OpenAI, AssemblyAI, and Deepgram. On a 3-hour challenge set with noise, low volume, and short utterances, Canto led among real-time models but trailed Gemini 3.1 Pro, a large multimodal model unfit for low-latency use. Canto was pretrained on millions of hours of speech and text, then fine-tuned with supervised learning and GRPO reinforcement learning to optimize full-transcript quality. On public benchmarks, Canto tied for first on LibriSpeech and was competitive but not leading on FLEURS and Common Voice; the post notes those datasets consist mostly of read speech, which differs from spontaneous dictation.
Why it matters: Canto brings concrete real-world WER comparisons that satisfy H and K, but Wispr isn't a tier-1 speech vendor so R is weak, landing it right at the featured threshold. Score isn't higher because this reads as a product-level model update, not an industry-shaking event.