Mistral releases Voxtral speech-understanding models in 24B and 3B
Voxtral
Mistral released Voxtral, a speech-understanding model in 24B and 3B versions, both open-sourced under Apache 2.0 and available via API. It supports a 32k token context, handling up to 30 minutes of transcription or 40 minutes of understanding, with built-in Q&A and summarization, multilingual recognition and voice function calling. It keeps the text abilities of Mistral Small 3.1.
Why it matters: Mistral open-sourced two speech-understanding models with 32k context and function calling, priced at less than half comparable APIs, which helps when picking a speech stack.