Skip to content

#模型发布

0 today

Yesterday · Sep 29Tuesday

OpenAI News

OpenAI releases GPT-6.1 Sol model

OpenAI released GPT-6.1 Sol, positioned as near-Astra-level intelligence for coding, computer use and professional work. Standard API input and output tokens cost one-fifth of Astra's price.

Why it matters: OpenAI's GPT-6.1 Sol launch shows the capability target for coding and computer use, plus the pricing shift.

Sep 25Friday

Google DeepMind

Google DeepMind releases Gemini 3.8 Live with Live Avatar

Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.

Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.

Aug 12Wednesday

Google DeepMind

Google DeepMind releases SL2T sign language-to-text model, first in Pixel 11 Gboard and Live Transcribe

Google DeepMind released SL2T, a multilingual sign language-to-text model, bringing sign language AI into consumer products for the first time. On Pixel 11, Gboard and Live Transcribe support American Sign Language (ASL) to English dictation, with more devices and languages to follow.

Why it matters: It gives SL2T's training scale, benchmark results and privacy design, so readers can judge the real limits of sign language translation in consumer products.

Jul 21Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber

Google DeepMind released three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and the security-focused 3.5 Flash Cyber.

Why it matters: It gives pricing, token efficiency and benchmark comparisons for all three models, so readers can judge cost and model choice for agent workflows.

Jul 17Friday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash Cyber security model

Google DeepMind released Gemini 3.5 Flash Cyber, fine-tuned from 3.5 Flash to find, verify and patch vulnerabilities quickly. With multiple calls, it approaches larger models on benchmarks such as CyberGym.

Why it matters: It reports how a lightweight security model performs on several benchmarks and inside Google's own codebase, so readers can judge the cost-benefit for vulnerability discovery.

Jun 9Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.5 Live Translate speech model

Google DeepMind released Gemini 3.5 Live Translate, an audio model for near-real-time speech-to-speech translation across more than 70 languages. It detects the language automatically and preserves the speaker's intonation, rhythm and pitch.

Why it matters: The original gives the model's language coverage, how the live translation works and the rollout pace across products, enough to judge where speech translation is usable.

May 18Monday

Google DeepMind

Google DeepMind releases Gemini Omni Flash video model

Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family. It combines image, audio, video and text inputs to generate high-quality video, and supports multi-turn editing in natural language.

Why it matters: Gemini Omni Flash folds video generation and conversational editing into one model, a shift in how multimodal creation gets accessed.

May 16Saturday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

Apr 16Thursday

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 13Monday

Google DeepMind

Google DeepMind releases Gemini Robotics-ER 1.6

Google DeepMind released Gemini Robotics-ER 1.6, an upgrade to its reasoning-first robotics model. It strengthens spatial reasoning and multi-view understanding, and adds gauge-reading ability.

Why it matters: The post details the new model's changes in spatial reasoning, multi-view understanding and gauge reading, plus where it is available, so you can judge progress in high-level robot reasoning.

Apr 3Friday

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Mar 17Tuesday

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Feb 5Thursday

Mistral AI

Mistral releases Voxtral Transcribe 2 speech-to-text model family

Mistral released Voxtral Transcribe 2, a family of two speech-to-text models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live use.

Why it matters: The post gives latency, pricing and open-source licensing for both transcription models, enough to judge the options for real-time voice applications.

Jan 6Tuesday

NVIDIA Blog

NVIDIA unveils new open models, data and tools across agents, robotics, AVs and biomedicine

NVIDIA released open models, datasets and training tools spanning Nemotron, Cosmos, Alpamayo, Isaac GR00T and Clara, plus 10T language tokens, 500K robotics trajectories, 455K protein structures and 100TB of vehicle sensor data. Newly disclosed items include Nemotron Speech/RAG/Safety, Cosmos Reason 2, Transfer 2.5, Predict 2.5, GR00T N1.6 and Alpamayo 1; the key signal is that NVIDIA is opening the data stack across agents, physical AI, AVs and biomedicine.

Dec 17, 2025Wednesday

Mistral AI

Mistral releases OCR 3 with better forms and handwriting, plus Document AI Playground

Mistral released Mistral OCR 3, which wins 74% of head-to-head comparisons against Mistral OCR 2 on forms, scanned documents, complex tables and handwriting. Mistral says its accuracy beats enterprise document-processing tools and AI-native OCR options.

Why it matters: Mistral OCR 3's upgrades on forms, handwriting and complex tables, plus its $2 per 1,000 pages pricing, are useful when evaluating document parsing options.

Dec 9, 2025Tuesday

Mistral AI

Mistral releases Devstral 2 coding models and the Mistral Vibe CLI

Mistral AI released the Devstral 2 coding model family: the 123B Devstral 2 and the 24B Devstral Small 2, under a modified MIT license and Apache 2.0 respectively. Both are open source.

Why it matters: The post gives Devstral 2's SWE-bench scores, open-source licenses and deployment requirements, enough to judge the cost of running open coding models.

Dec 3, 2025Wednesday

Mistral AI

Mistral releases the Mistral 3 family, including 675B-parameter Mistral Large 3

Mistral AI released the Mistral 3 family: three dense models at 14B, 8B and 3B, plus Mistral Large 3, which uses a sparse MoE architecture with 41B active and 675B total parameters. All are open-sourced under Apache 2.0.

Why it matters: Mistral 3 ships an Apache 2.0 family from 3B to 675B in one release, a useful read on where open weights now stand for on-device and frontier capability.

Aug 7, 2025Thursday

OpenAI News

Introducing GPT-5

OpenAI launched GPT-5 on August 7, 2025 and made it available to all ChatGPT users. The system combines a base model, GPT-5 thinking, and a real-time router; Plus gets higher limits, while Pro gets GPT-5 pro. The key change is unified routing with built-in reasoning; the post does not disclose pricing, context window, or API specifics.

Why it matters: An OpenAI frontier-model launch is a top-band event on its own. The excerpt confirms a unified system (base model + GPT-5 thinking + router) and rollout to all ChatGPT users; HKR-H/K/R all pass, and missing price/context/API details do not block p1.

Aug 5, 2025Tuesday

OpenAI News

Open Weights and AI for All

OpenAI said on August 5, 2025 it released its “most capable open-weight reasoning models” and will route them through OpenAI for Countries and its nonprofit grantee programs. The post confirms on-prem deployment and support for data-residency and security-constrained use cases, but does not disclose model names, parameter sizes, licenses, or benchmark results. The key missing piece is distribution detail, not the open-weight claim itself.

Why it matters: OpenAI shipping open-weights reasoning models clears HKR-H/K/R on novelty, a concrete deployment fact, and strategic resonance. Held at 86, not higher, because the post withholds the model name, size, license, and benchmark scores.

OpenAI News

Introducing gpt-oss

OpenAI released gpt-oss-120b and gpt-oss-20b under Apache 2.0, with the 120B model running on one 80GB GPU and the 20B model on devices with 16GB memory. Both are MoE Transformers with 117B and 21B total parameters, 5.1B and 3.6B active params per token, 128k context, and support for the Responses API and Structured Outputs. The part that matters is the lower deployment bar plus open weights; the post excerpt claims strong reasoning, but full benchmark scores are not disclosed here.

Why it matters: Same-day write. OpenAI moving into Apache 2.0 open weights is a strategy story, not a routine update; HKR-H lands on the unexpected move, HKR-K on concrete deployment specs, and HKR-R on cost and open-vs-closed debates. Not 95+ because the excerpt does not disclose full benchmark

Jul 15, 2025Tuesday

Mistral AI

Mistral releases Voxtral speech-understanding models in 24B and 3B

Mistral released Voxtral, a speech-understanding model in 24B and 3B versions, both open-sourced under Apache 2.0 and available via API. It supports a 32k token context, handling up to 30 minutes of transcription or 40 minutes of understanding, with built-in Q&A and summarization, multilingual recognition and voice function calling. It keeps the text abilities of Mistral Small 3.1.

Why it matters: Mistral open-sourced two speech-understanding models with 32k context and function calling, priced at less than half comparable APIs, which helps when picking a speech stack.

Jul 11, 2025Friday

Mistral AI

Mistral releases Devstral Medium and upgrades Devstral Small 1.1

Mistral AI worked with All Hands AI to launch Devstral Medium and upgrade Devstral Small 1.1.

Why it matters: Mistral and All Hands AI jointly released two coding agent models with SWE-Bench Verified scores and API pricing, making comparison with existing options easier.

Jun 10, 2025Tuesday

Mistral AI

Mistral AI releases its first reasoning model, Magistral, in open and enterprise versions

Mistral AI released Magistral, its first reasoning model, in two versions: the 24B open-source Magistral Small and the enterprise Magistral Medium.

Why it matters: Mistral's first reasoning model comes in two versions with parameter counts and AIME2024 results, so you can judge its open-source and commercial positioning.

May 28, 2025Wednesday

Mistral AI

Codestral Embed

Mistral AI 发布首个代码专用嵌入模型 Codestral Embed,官方称其在真实代码数据检索上显著优于 Voyage Code 3、Cohere Embed v4.0 和 OpenAI 的大型嵌入模型。

May 21, 2025Wednesday

Mistral AI

Mistral AI releases agentic coding model Devstral under Apache 2.0

Mistral AI and All Hands AI released Devstral, an agentic LLM for software engineering tasks, under the Apache 2.0 license. It scores 46.8% on SWE-Bench Verified, more than 6 points above the previous open-source state of the art.

Why it matters: A joint Mistral and All Hands AI agentic coding model, with its SWE-Bench Verified score and the bar for local deployment.

May 7, 2025Wednesday

Mistral AI

Mistral AI releases Mistral Medium 3, targeting low cost and enterprise deployment

Mistral AI released Mistral Medium 3, which it says reaches or exceeds 90% of Claude Sonnet 3.7 across benchmarks. Pricing is $0.4 per million input tokens and $2 per million output tokens.

Why it matters: Mistral Medium 3 benchmarks against Claude Sonnet 3.7 at lower cost, and the post gives a path to private enterprise deployment and customization.

Mar 17, 2025Monday

Mistral AI

Mistral AI releases Mistral Small 3.1 with multimodal support and 128k context

Mistral AI released Mistral Small 3.1, which improves text performance and multimodal understanding over Mistral Small 3 and extends the context window to 128k tokens. Inference runs at 150 tokens per second, and the model is open-sourced under Apache 2.0.

Why it matters: Mistral gives the multimodal, 128k-context and 150 tokens/s figures for Mistral Small 3.1, letting readers compare it with small models of the same class.

Mar 12, 2025Wednesday

Hugging Face Blog

Welcome Gemma 3: Google's all new multimodal, multilingual, long context open LLM

Google announced an open LLM called Gemma 3 and named three traits in the title: multimodal, multilingual, and long context. The RSS snippet has no body, so parameter size, context length, license, and benchmark results are not disclosed. Watch the full post or model card; “open” does not equal open source from the title alone.

Why it matters: A new Gemma release from Google is inherently newsy, and the multimodal/long-context/open framing hits HKR-H and HKR-R. HKR-K misses because the feed gives no specs, context window, license, or benchmark data, so this lands at the low end of featured.

Oct 1, 2024Tuesday

OpenAI News

Model Distillation in the API

OpenAI launched an API distillation workflow on October 1, 2024, letting developers use outputs from GPT-4o and o1-preview to fine-tune cheaper models such as GPT-4o mini. The suite includes Stored Completions, Evals in beta, and fine-tuning; setting store:true auto-saves input-output pairs with no added latency, per the post. Pricing includes 2M free GPT-4o mini training tokens per day and 1M for GPT-4o through October 31; Evals are free up to 7 runs per week through year-end if shared with OpenAI.

Sep 12, 2024Thursday

OpenAI News

Introducing OpenAI o1

OpenAI released o1-preview and o1-mini on Sept. 12, 2024, with access for ChatGPT Plus, Team, and tier-5 API developers. The post cites 83% vs 13% on an IMO qualifier, 84 vs 22 on a jailbreak test, and says o1-mini is 80% cheaper than o1-preview. The tradeoff is clear: the API lacks function calling, streaming, and system messages, and the models do not yet support browsing or file and image uploads.

Why it matters: A major OpenAI reasoning-model launch with all three HKR signals: HKR-H from the new “think before answering” hook, HKR-K from concrete benchmark, safety, and pricing numbers, and HKR-R from the tradeoff practitioners must manage between stronger reasoning and missing API basics.