Skip to content

#模型发布

5 today

May 20Wednesday

AI HOT (Curated Pool)

Stability AI Launches Stability Audio 3.0 for Songs Up to 6 Minutes

Stability AI launched the Stability Audio 3.0 audio generation model family with four sizes ranging from 459 million to 2.7 billion parameters; the small model targets on-device use and generates audio under 2 minutes locally, while medium and large models support full music creation beyond 6 minutes and 20 seconds.

Why it matters: HKR-H/K pass because Stability AI gives concrete duration and model-size details. HKR-R is weak: no benchmarks, licensing, pricing, or access terms are disclosed, so this sits at the featured threshold.

AI HOT (Curated Pool)

Kling AI Launches the First Native 4K Video Generation Model

Kling AI launched a native 4K video generation model on April 23, supporting one-click true 4K video generation; the post says Hollywood teams and Wonder Studios have adopted it, but does not disclose pricing, inference cost, or access limits.

Why it matters: HKR-H and HKR-K pass: Kling AI’s native 4K video model has a concrete capability and named adoption. Source is product-side, with no benchmark, pricing, or clip-duration data, so it sits at the featured threshold.

AI HOT (Curated Pool)

Smarter Google AI Edge Gallery: MCP Integration, Notifications, and Session Continuity

Google AI Edge Gallery adds experimental MCP support on Android, letting Gemma 4 coordinate external data sources including Google Workspace and Google Maps; the update also adds scheduled notifications and persistent chat history for faster restoration of long-session context.

Why it matters: HKR-H/K/R all pass: Google’s developer update adds experimental MCP, notifications, and session continuity to AI Edge Gallery. It is a mid-weight product update, not a model release or major capability launch.

AI HOT (Curated Pool)

Gemini 3.5 Released: A New Model Family Combining Intelligence and Action

Google AI Developers announced the Gemini 3.5 model family, saying it combines intelligence with action capabilities; the post does not disclose parameters, benchmarks, pricing, availability, or context window details.

Why it matters: HKR-H and HKR-R pass: an official Gemini 3.5 family launch has flagship-model pull and competitive resonance. HKR-K fails because the post gives no params, benchmarks, pricing, or context window, so this stays below the 85+ band.

Hacker News front page

Gemini 3.5 Flash

The title names Gemini 3.5 Flash, while the RSS body only includes a documentation link; the Hacker News item has 196 points and 179 comments, and the post does not disclose parameters, pricing, or context-window details.

Why it matters: HKR-H/R pass on an official Google Gemini 3.5 Flash release with HN traction; HKR-K fails because price, benchmarks, parameters, and context window are absent. That keeps it in the lower 78–84 band.

May 19Tuesday

AI HOT (Curated Pool)

Cursor releases Composer 2.5, calling it its strongest model yet

Cursor released Composer 2.5, claiming a 10x efficiency gain at comparable capability, with larger training scale, more complex reinforcement-learning environments, and a text-feedback mechanism.

Why it matters: Cursor Composer 2.5 is a substantive model update for a front-line AI coding tool, with HKR-H/K/R from the 10x efficiency and RL-training details. The single social-source summary lacks benchmarks, pricing, and reproducible tests, keeping it in the 78–84 band.

May 18Monday

Google DeepMind

Google DeepMind releases Gemini Omni Flash video model

Google DeepMind released Gemini Omni Flash, the first model in the Gemini Omni family. It combines image, audio, video and text inputs to generate high-quality video, and supports multi-turn editing in natural language.

Why it matters: Gemini Omni Flash folds video generation and conversational editing into one model, a shift in how multimodal creation gets accessed.

May 17Sunday

r/LocalLLaMA

MiroThinker-1.7 Open-Weight Deep Research Agent Based on Qwen3 MoE

MiroMindAI released the MiroThinker-1.7-deepresearch and mini APIs, with the mini version using 30B total parameters and 3B active parameters, weights on HuggingFace, and context management based on sliding window K=5 plus episode restarts.

Why it matters: HKR-H/K/R all pass, but the source is a Reddit thread and the lab is not top-tier. Open weights, MoE sizing, and context-management details clear featured, not same-day must-write.

May 16Saturday

Google DeepMind

Google DeepMind releases Gemini 3.5 Flash

Google DeepMind released the Gemini 3.5 model family, with the first model, Gemini 3.5 Flash, available the same day in the Gemini app, Google Search AI Mode, Google Antigravity, the Gemini API and Gemini Enterprise.

Why it matters: Google published 3.5 Flash's coding and agent benchmark scores and where it is available, enough to judge its place in long-horizon workflows.

May 15Friday

AI HOT (Curated Pool)

Granite Embedding Multilingual R2: Open Multilingual Embedding Model with 32K Context

IBM Granite released Granite Embedding Multilingual R2 on Hugging Face under Apache 2.0, with fewer than 100 million parameters, a 32K-token context length, and top same-scale retrieval performance on MTEB according to the post.

Why it matters: HKR-H/K/R pass: the 32K-context, sub-100M multilingual embedding model gives RAG builders a concrete open-source option. Impact is narrower than a frontier-model release, so it sits at the featured threshold.

May 12Tuesday

AI HOT (Curated Pool)

Thinking Machines Releases Native Multimodal Interaction Model for Real-Time Human-AI Collaboration

Thinking Machines released an interaction model that natively receives audio, video, and text input, processes foreground interaction at 200-millisecond intervals, and uses a background reasoning model for long-horizon planning and tool calls.

Why it matters: HKR-H/K/R all pass: this is more than a model notice, with a two-layer foreground/background interaction design. Pricing, access scope, and benchmarks are missing, so it sits at the lower end of 85-94.

May 11Monday

AI HOT (Curated Pool)

AntLingAGI Releases Trillion-Parameter Ring-2.6-1T Model

AntLingAGI released Ring-2.6-1T, a trillion-parameter thinking model available for free on OpenRouter until May 15, with adjustable thinking intensity, agent-oriented multi-step execution, tool calling, and tasks covering math logic and scientific research.

Why it matters: HKR-H/K/R all pass, but the post is thin: no benchmarks, pricing, architecture, or training details. Treat as a mid-weight model launch on OpenRouter, not a same-day must-write.

May 9Saturday

AI HOT (Curated Pool)

Baidu releases ERNIE 5.1 with compressed parameters and training cost

Baidu released ERNIE 5.1 with total parameters reduced to about one third of the original scale, active parameters to about one half, and pretraining cost to about 6% of same-scale models; the model is available on the ERNIE platform and Baidu AI Studio.

Why it matters: HKR-H/K/R all pass: Baidu ERNIE 5.1 is a domestic flagship-model release with concrete compression and 6% pretraining-cost claims. That puts it in the must-write band.

AI HOT (Curated Pool)

ERNIE 5.1 Released With Pretraining Cost at 6% of Comparable Models

Baidu released ERNIE 5.1, saying it builds on ERNIE 5.0 pretraining and improves search, reasoning, knowledge QA, creative writing, and agent capabilities, with pretraining cost at about 6% of comparable models.

Why it matters: Baidu released ERNIE 5.1 with a concrete “6% of reference pretraining cost” claim. HKR-H/K/R all pass, with a domestic flagship-model bump, but sparse technical detail keeps it below the 90s.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

May 5Tuesday

r/LocalLLaMA

Heretic 1.3 Released: Reproducible Models, Integrated Benchmarks, Lower Peak VRAM

Heretic 1.3 adds reproducible runs, integrated benchmarks, lower peak VRAM, and broader model support. The project claims 20,000 GitHub stars and 13 million model downloads. Reproduce directories capture PyTorch, GPU, driver, and accelerator details; benchmarks use lm-evaluation-harness for MMLU, EQ-Bench, GSM8K, and HellaSwag. The post names Qwen3.5 and Gemma 4 support, but does not disclose VRAM reduction figures.

Why it matters: HKR-K/R pass: 20k stars, 13M downloads, reproducibility metadata, and eval harness are concrete. HKR-H fails and VRAM reduction lacks numbers, so this sits at the featured threshold.

Apr 29Wednesday

r/LocalLLaMA

mistralai/Mistral-Medium-3.5-128B · Hugging Face

Mistral AI released Mistral Medium 3.5 128B on Hugging Face, with 128B dense parameters and a 256k context window. It supports text and image input, function calls, JSON output, and a Modified MIT License with exceptions for high-revenue firms. Reasoning effort is configurable as none or high per request.

Why it matters: HKR-H/K/R all pass for a major Mistral model release with concrete specs. It stays at 84 because benchmarks, pricing, and reproducible tests are not disclosed in the body.

Apr 25Saturday

MIT Technology Review · AI

Three reasons why DeepSeek’s new model matters

DeepSeek released a V4 preview with two versions: V4-Pro and V4-Flash. V4-Pro costs $1.74/M input tokens and $3.48/M output tokens; V4-Flash is about $0.14/$0.28, and both support 1M-token context. The key point is attention efficiency and open weights pressuring agentic coding costs.

Why it matters: HKR-H/K/R all pass: DeepSeek V4 is a domestic flagship release with 1M context, two price tiers, and open-weight cost pressure. The preview status keeps it below a full GPT/Claude major release, but it is same-day material.

Bloomberg Technology

China’s DeepSeek Unveils New Model a Year After Shock Launch

DeepSeek unveiled a new flagship AI model about one year after its open-source release jolted Silicon Valley. The title and RSS snippet confirm that timing; the post does not disclose the model name, size, pricing, benchmarks, or release terms. The key thing to watch is the missing launch detail, not the comeback framing.

Why it matters: A new DeepSeek flagship is newsworthy: HKR-H comes from the 'one year after the shock launch' hook, and HKR-R from the open-source and pricing rivalry it triggers. HKR-K fails because no model name, params, pricing, or benchmarks are disclosed, so this sits at the low end of the

Apr 24Friday

Bloomberg Technology

DeepSeek unveils flagship AI model a year after breakthrough

DeepSeek released preview versions of a new flagship AI model one year after its breakout. The RSS snippet calls it its most powerful open-source platform and frames it against OpenAI and Anthropic; the post does not disclose parameters, context length, benchmarks, or rollout timing. The actionable facts so far are limited to its preview status and open-source positioning.

Why it matters: A new DeepSeek flagship preview deserves real weight under the domestic-flagship rule, and Bloomberg adds source authority. HKR-H and HKR-R pass, but HKR-K fails because the story discloses no specs, context window, benchmarks, or release schedule, so this stays at the low end of

X · @dotey

DeepSeek releases and open-sources V4 preview; 1M context is standard across all services

DeepSeek released and open-sourced the V4 preview, making 1M context standard across all official services with no tier or price split. The post says V4-Pro and V4-Flash use token compression plus DSA sparse attention to cut compute and memory costs for 1M context; legacy APIs remain for 3 months and stop after July 24.

Why it matters: DeepSeek is a flagship Chinese model vendor, and this V4 preview is a substantive release with open source and 1M context made standard across official services. HKR-H/K/R all pass: the post includes mechanisms and a migration deadline, and the tier reset makes it a same-day P1.

Apr 23Thursday

Bloomberg Technology

Tencent unveils a major AI foundation model upgrade, testing its new OpenAI hire

Tencent announced a major upgrade to its AI foundation model. It is the company's first high-stakes AI test since hiring a top OpenAI researcher. The post does not disclose the model name, parameter count, benchmarks, or launch timing.

Why it matters: Bloomberg provides source authority, and the framing is strong: Tencent's model release is presented as the first test of its OpenAI hire, so HKR-H and HKR-R pass. HKR-K fails because the story does not disclose the model name, size, benchmarks, or launch timing, keeping it at a

Apr 21Tuesday

Latent Space

Moonshot Kimi K2.6 open-weight model refresh aims to catch Opus 4.6

Moonshot released Kimi K2.6, a 1T-parameter MoE with 32B active and 256K context. The post cites 58.6 on SWE-Bench Pro, 4,000+ tool calls, 12+ hour runs, and 300 parallel sub-agents. The key signal is long-horizon agent execution, not only open-model scores.

Why it matters: HKR-H/K/R all pass: Kimi K2.6 has a strong race narrative, concrete model and agent metrics, and direct relevance to open-model builders. The domestic flagship release signal lifts it into P1.

Apr 20Monday

QbitAI · WeChat

Sudo, valued above $2 billion, unveils embodied model Sudo R1 with zero real-robot data and ~98% first-try grasp success

Sudo unveiled embodied model Sudo R1 and says it achieved about 98% first-try grasp success in 200+ zero-shot tests with zero real-robot training data, nearing 100% within two attempts. The post says the 60-minute run covered 100+ unseen objects, including transparent, metallic, soft, and reflective items, using integrated world-model and reinforcement-learning training on a high-fidelity simulator. It also says Sudo is valued above $2 billion and is working with CATL, but the post does not disclose round size, benchmark protocol, or third-party validation.

Why it matters: Strong HKR-H/K/R: the zero-real-data, zero-shot, 98% claim is novel and concrete, and it hits robotics' data-cost nerve. Kept below 85 because the metrics are self-reported; funding amount, benchmark definition, and third-party validation are not disclosed.

Apr 16Thursday

36Kr (direct RSS)

Anthropic plans to release its Mythos model to UK banking institutions next week

Anthropic PBC plans to grant UK financial institutions early access to its Mythos model within the next week. The mechanism is the “Glass Wing” program for selected institutions; Anthropic says the model can identify and potentially exploit cybersecurity flaws, while the post does not disclose specs, pricing, or customer count. The key signal is controlled access, not a broad launch.

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 13Monday

Google DeepMind

Google DeepMind releases Gemini Robotics-ER 1.6

Google DeepMind released Gemini Robotics-ER 1.6, an upgrade to its reasoning-first robotics model. It strengthens spatial reasoning and multi-view understanding, and adds gauge-reading ability.

Why it matters: The post details the new model's changes in spatial reasoning, multi-view understanding and gauge reading, plus where it is available, so you can judge progress in high-level robot reasoning.

Apr 9Thursday

QbitAI · WeChat

Beyond MoE, Tencent introduces MoT: a 2B embodied model ranks first in 16 of 22 evaluations

Tencent Hunyuan and Robotics X released HY-Embodied-0.5; its MoT-2B uses 4B total params with 2B active and ranks first in 16 of 22 embodied evaluations. The post says it uses 100M+ embodied data, 600B+ pretraining tokens, 30M+ mid-training samples, plus visual latent tokens, bidirectional attention, RFT, RL, and online distillation. The key point is a rebuilt edge-oriented embodied stack, not a simple VLM fine-tune.

Why it matters: Strong on HKR-H/K/R: the headline has a real hook, the body includes concrete numbers and training mechanisms, and the edge-robotics angle lands with practitioners. I keep it at 83, not 85+, because this is a high-quality embodied-model release, not a broad same-day industry-def

X · @op7418

Meta releases Muse Spark model

Meta released the Muse Spark model with native multimodal reasoning, tool use, visual chain-of-thought, and multi-agent orchestration, but it is only available in the Meta AI app and is not open source for now. The snippet says its Contemplating mode coordinates multiple parallel agents for reasoning, and its Artificial Analysis score is below Gemini 3.1 Pro, GPT-5.4, and Claude Opus 4.6. The post does not disclose model size, pricing, or rollout timing.

Why it matters: A major-lab model launch plus the “poached team’s first output” angle lands HKR-H/K/R. The score stays near the featured floor because the post offers capability claims and relative benchmark placement only; params, pricing, rollout timing, and access scope are not disclosed.

Apr 3Friday

Google DeepMind

Google DeepMind releases the Gemma 4 open model family

Google DeepMind released Gemma 4, which it calls its most intelligent open model yet, aimed at advanced reasoning and agentic workflows under an Apache 2.0 license. The family comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. The 31B ranks 3rd among open models on the Arena AI text leaderboard, and the 26B ranks 6th.

Why it matters: Gemma 4 is Apache 2.0 and spans four sizes from on-device to workstation, so you can weigh deployment and fine-tuning options for open models.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Mar 17Tuesday

Mistral AI

Mistral releases Mistral Small 4, unifying reasoning, multimodal and coding

Mistral AI released Mistral Small 4, the first Mistral model to unify Magistral reasoning, Pixtral multimodal and Devstral coding-agent abilities in a single model. It ships under the Apache 2.0 license.

Why it matters: Merging reasoning, multimodal and coding agents into one open model is a direct test of what unified models do to deployment cost.

Feb 10Tuesday

36Kr (direct RSS)

Alibaba Qwen launches the new image generation foundation model Qwen-Image-2.0

Alibaba Qwen announced Qwen-Image-2.0, an image generation foundation model, and opened API invite testing on Alibaba Cloud Bailian. Developers can also try it for free in Qwen Chat; the post does not disclose model size, pricing, eval results, or a general release date.

Feb 6Friday

TechCrunch · AI

OpenAI launches new agentic coding model minutes after Anthropic releases its own

OpenAI launched an agentic coding model minutes after Anthropic released a similar one, and the model is meant to accelerate Codex, which OpenAI launched earlier this week. The RSS snippet gives only the timing and purpose; the post does not disclose the model name, benchmarks, pricing, context length, or availability. The signal is direct competition in agentic coding, not a substantiated performance claim.

Why it matters: Major-lab product news plus a minutes-apart Anthropic clash gives this HKR-H and HKR-R. The score stays in the low featured band because HKR-K is weak: the post lacks the model name, benchmarks, price, context window, and availability.

Feb 5Thursday

Mistral AI

Mistral releases Voxtral Transcribe 2 speech-to-text model family

Mistral released Voxtral Transcribe 2, a family of two speech-to-text models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live use.

Why it matters: The post gives latency, pricing and open-source licensing for both transcription models, enough to judge the options for real-time voice applications.

Jan 6Tuesday

NVIDIA Blog

NVIDIA unveils new open models, data and tools across agents, robotics, AVs and biomedicine

NVIDIA released open models, datasets and training tools spanning Nemotron, Cosmos, Alpamayo, Isaac GR00T and Clara, plus 10T language tokens, 500K robotics trajectories, 455K protein structures and 100TB of vehicle sensor data. Newly disclosed items include Nemotron Speech/RAG/Safety, Cosmos Reason 2, Transfer 2.5, Predict 2.5, GR00T N1.6 and Alpamayo 1; the key signal is that NVIDIA is opening the data stack across agents, physical AI, AVs and biomedicine.

Dec 17, 2025Wednesday

Mistral AI

Mistral releases OCR 3 with better forms and handwriting, plus Document AI Playground

Mistral released Mistral OCR 3, which wins 74% of head-to-head comparisons against Mistral OCR 2 on forms, scanned documents, complex tables and handwriting. Mistral says its accuracy beats enterprise document-processing tools and AI-native OCR options.

Why it matters: Mistral OCR 3's upgrades on forms, handwriting and complex tables, plus its $2 per 1,000 pages pricing, are useful when evaluating document parsing options.

Dec 9, 2025Tuesday

Mistral AI

Mistral releases Devstral 2 coding models and the Mistral Vibe CLI

Mistral AI released the Devstral 2 coding model family: the 123B Devstral 2 and the 24B Devstral Small 2, under a modified MIT license and Apache 2.0 respectively. Both are open source.

Why it matters: The post gives Devstral 2's SWE-bench scores, open-source licenses and deployment requirements, enough to judge the cost of running open coding models.

Dec 3, 2025Wednesday

Mistral AI

Mistral releases the Mistral 3 family, including 675B-parameter Mistral Large 3

Mistral AI released the Mistral 3 family: three dense models at 14B, 8B and 3B, plus Mistral Large 3, which uses a sparse MoE architecture with 41B active and 675B total parameters. All are open-sourced under Apache 2.0.

Why it matters: Mistral 3 ships an Apache 2.0 family from 3B to 675B in one release, a useful read on where open weights now stand for on-device and frontier capability.