Skip to content

Model releases

New models, open releases and updates: flagship launches, open weights, and price and performance changes as they happen.

Latest picks

101–107 of 107

Jun 10, 2025Tuesday

Mistral AI

Mistral AI releases its first reasoning model, Magistral, in open and enterprise versions

Mistral AI released Magistral, its first reasoning model, in two versions: the 24B open-source Magistral Small and the enterprise Magistral Medium.

Why it matters: Magistral ships as a 24B open model and an enterprise version, with AIME2024 scores for both.

May 21, 2025Wednesday

Mistral AI

Mistral AI releases agentic coding model Devstral under Apache 2.0

Mistral AI and All Hands AI released Devstral, an agentic LLM for software engineering tasks, under the Apache 2.0 license. It scores 46.8% on SWE-Bench Verified, more than 6 points above the previous open-source state of the art.

Why it matters: A joint Mistral and All Hands AI agentic coding model, with its SWE-Bench Verified score and the bar for local deployment.

May 7, 2025Wednesday

Mistral AI

Mistral AI releases Mistral Medium 3, targeting low cost and enterprise deployment

Mistral AI released Mistral Medium 3, which it says reaches or exceeds 90% of Claude Sonnet 3.7 across benchmarks. Pricing is $0.4 per million input tokens and $2 per million output tokens.

Why it matters: Mistral Medium 3 benchmarks against Claude Sonnet 3.7 at lower cost, and the post gives a path to private enterprise deployment and customization.

Mar 17, 2025Monday

Mistral AI

Mistral AI releases Mistral Small 3.1 with multimodal support and 128k context

Mistral AI released Mistral Small 3.1, which improves text performance and multimodal understanding over Mistral Small 3 and extends the context window to 128k tokens. Inference runs at 150 tokens per second, and the model is open-sourced under Apache 2.0.

Why it matters: Mistral AI open-sources Mistral Small 3.1 under Apache 2.0, extending context to 128k tokens and adding multimodal understanding, runnable on one RTX 4090 or a 32GB Mac.

Mar 12, 2025Wednesday

Hugging Face Blog

Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM

Google announced an open LLM called Gemma 3 and named three traits in the title: multimodal, multilingual, and long context. The RSS snippet has no body, so parameter size, context length, license, and benchmark results are not disclosed. Watch the full post or model card; “open” does not equal open source from the title alone.

Why it matters: Google released the open-weight Gemma 3 models, with 4B and up handling images and text across more than 140 languages, and claims the 4B instruction version outscores the previous 27B.

Oct 1, 2024Tuesday

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Why it matters: Turns distillation into a built-in API step instead of a pipeline teams build themselves.

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: OpenAI's o1 models trade speed for deliberation, lifting math olympiad accuracy from GPT-4o's 13% to 83% and reaching Codeforces top 11%.