Skip to content

Model releases

New models, open releases and updates: flagship launches, open weights, and price and performance changes as they happen.

Latest picks

81–88 of 88

Jul 11, 2025Friday

Mistral AI

Mistral releases Devstral Medium and upgrades Devstral Small 1.1

Mistral AI worked with All Hands AI to launch Devstral Medium and upgrade Devstral Small 1.1.

Why it matters: Mistral and All Hands AI jointly released two coding agent models with SWE-Bench Verified scores and API pricing, making comparison with existing options easier.

Jun 10, 2025Tuesday

Mistral AI

Mistral AI releases its first reasoning model, Magistral, in open and enterprise versions

Mistral AI released Magistral, its first reasoning model, in two versions: the 24B open-source Magistral Small and the enterprise Magistral Medium.

Why it matters: Mistral's first reasoning model comes in two versions with parameter counts and AIME2024 results, so you can judge its open-source and commercial positioning.

May 21, 2025Wednesday

Mistral AI

Mistral AI releases agentic coding model Devstral under Apache 2.0

Mistral AI and All Hands AI released Devstral, an agentic LLM for software engineering tasks, under the Apache 2.0 license. It scores 46.8% on SWE-Bench Verified, more than 6 points above the previous open-source state of the art.

Why it matters: A joint Mistral and All Hands AI agentic coding model, with its SWE-Bench Verified score and the bar for local deployment.

May 7, 2025Wednesday

Mistral AI

Mistral AI releases Mistral Medium 3, targeting low cost and enterprise deployment

Mistral AI released Mistral Medium 3, which it says reaches or exceeds 90% of Claude Sonnet 3.7 across benchmarks. Pricing is $0.4 per million input tokens and $2 per million output tokens.

Why it matters: Mistral Medium 3 benchmarks against Claude Sonnet 3.7 at lower cost, and the post gives a path to private enterprise deployment and customization.

Mar 17, 2025Monday

Mistral AI

Mistral AI releases Mistral Small 3.1 with multimodal support and 128k context

Mistral AI released Mistral Small 3.1, which improves text performance and multimodal understanding over Mistral Small 3 and extends the context window to 128k tokens. Inference runs at 150 tokens per second, and the model is open-sourced under Apache 2.0.

Why it matters: Mistral gives the multimodal, 128k-context and 150 tokens/s figures for Mistral Small 3.1, letting readers compare it with small models of the same class.

Mar 12, 2025Wednesday

Hugging Face Blog

Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM

Google announced an open LLM called Gemma 3 and named three traits in the title: multimodal, multilingual, and long context. The RSS snippet has no body, so parameter size, context length, license, and benchmark results are not disclosed. Watch the full post or model card; “open” does not equal open source from the title alone.

Why it matters: A new Gemma release from Google is inherently newsy, and the multimodal/long-context/open framing hits HKR-H and HKR-R. HKR-K misses because the feed gives no specs, context window, license, or benchmark data, so this lands at the low end of featured.

Oct 1, 2024Tuesday

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: A major OpenAI reasoning-model launch with all three HKR signals: HKR-H from the new “think before answering” hook, HKR-K from concrete benchmark, safety, and pricing numbers, and HKR-R from the tradeoff practitioners must manage between stronger reasoning and missing API basics.