Skip to content

#语音

0 today

Sep 25Friday

Google DeepMind

Google DeepMind releases Gemini 3.8 Live with Live Avatar

Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.

Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.

Jun 9Tuesday

Google DeepMind

Google DeepMind releases Gemini 3.5 Live Translate speech model

Google DeepMind released Gemini 3.5 Live Translate, an audio model for near-real-time speech-to-speech translation across more than 70 languages. It detects the language automatically and preserves the speaker's intonation, rhythm and pitch.

Why it matters: The original gives the model's language coverage, how the live translation works and the rollout pace across products, enough to judge where speech translation is usable.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

Apr 30Thursday

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

Apr 16Thursday

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 3Friday

X · @OpenAI

ChatGPT is now available in CarPlay

OpenAI is rolling out ChatGPT in CarPlay to iPhone users on iOS 26.4+ where CarPlay is supported. The post confirms voice mode is available in-car, but does not disclose regions, vehicle coverage, or feature limits. The key shift is distribution into the driving interface, not a new model launch.

Why it matters: This matters more as a distribution-surface shift than a model update. HKR-H and HKR-R pass on the CarPlay hook and assistant-entry competition; HKR-K stays limited because the post gives iOS 26.4+ rollout only, not regions, car support, or full feature bounds.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Feb 5Thursday

Mistral AI

Mistral releases Voxtral Transcribe 2 speech-to-text model family

Mistral released Voxtral Transcribe 2, a family of two speech-to-text models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live use.

Why it matters: The post gives latency, pricing and open-source licensing for both transcription models, enough to judge the options for real-time voice applications.

Sep 30, 2025Tuesday

OpenAI News

Sora 2 System Card

OpenAI published the Sora 2 System Card on September 30, 2025, and said the video-audio generation model will launch first via limited invites on sora.com and a standalone iOS app. The post confirms no video uploads and no image uploads with photorealistic people at launch; API timing, pricing, and benchmark scores are not disclosed.

Why it matters: This lands in the 78–84 band. HKR-H comes from the Sora 2 + iOS app hook; HKR-K from concrete launch limits and safety rules; HKR-R from competition and likeness-abuse nerves. It stays below P1 because price, eval scores, context details, and API timing are not disclosed.

OpenAI News

Sora 2 is here

OpenAI released Sora 2 on September 30, 2025 and launched a social iOS app called Sora built on the model. The post says it generates video with synced dialogue and sound effects, and its “characters” feature uses a one-time video and audio recording to verify identity and insert a real person’s likeness; pricing, generation limits, and rollout regions are not disclosed. The key shift is from model demo to a consumer app with a feed, teen limits, and parental controls.

Why it matters: This is a same-day write: OpenAI shipped a flagship video/audio model and attached it to a standalone app, so HKR-H/K/R all clear. The post gives real product facts like synced dialogue and sound effects, but missing price, duration caps, and rollout details keeps it below 90.

Aug 28, 2025Thursday

OpenAI News

Introducing gpt-realtime and Realtime API updates for production voice agents

OpenAI released the speech-to-speech model gpt-realtime and made the Realtime API generally available, adding remote MCP server support, image input, and SIP phone calling. The post reports 82.8% on Big Bench Audio versus 65.6% for the December 2024 model, and 30.5% on the audio MultiChallenge benchmark versus 20.6%. The key change is that tool access and phone connectivity now ship in the same production API.

Why it matters: This is a substantive OpenAI model + API release, not a minor refresh. HKR-H/K/R all pass: the release has a clear hook, hard benchmark deltas, and direct deployment impact for production voice agents, so it reaches p1.

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

Jul 15, 2025Tuesday

Mistral AI

Mistral releases Voxtral speech-understanding models in 24B and 3B

Mistral released Voxtral, a speech-understanding model in 24B and 3B versions, both open-sourced under Apache 2.0 and available via API. It supports a 32k token context, handling up to 30 minutes of transcription or 40 minutes of understanding, with built-in Q&A and summarization, multilingual recognition and voice function calling. It keeps the text abilities of Mistral Small 3.1.

Why it matters: Mistral open-sourced two speech-understanding models with 32k context and function calling, priced at less than half comparable APIs, which helps when picking a speech stack.

Mar 21, 2025Friday

OpenAI News

Early methods for studying affective use and emotional well-being on ChatGPT

OpenAI and MIT Media Lab studied affective use on ChatGPT with two tracks: nearly 40 million interactions in an observational analysis and a 4-week RCT with nearly 1,000 participants. The post says emotional engagement is rare overall and concentrated in a small subset of heavy Advanced Voice Mode users; the provided body does not fully disclose all quantitative well-being results. Watch subgroup effects, not platform averages.

Mar 20, 2025Thursday

OpenAI News

Introducing next-generation audio models in the API

OpenAI released three API audio models on March 20, 2025: gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-mini-tts. The post says the STT models beat Whisper v2 and v3 on FLEURS and other benchmarks across 100+ languages, while the TTS model adds style control but stays limited to monitored preset synthetic voices. The key shift is controllable TTS plus lower WER; the post does not disclose pricing or latency figures.

Why it matters: OpenAI shipped 3 API audio models with concrete benchmark and mechanism details, so HKR-H/K/R all pass and it clears featured. I kept it at 84, not 85+, because price, latency, and a fuller benchmark table are not disclosed.

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

Aug 8, 2024Thursday

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.

Sep 25, 2023Monday

OpenAI News

ChatGPT can now see, hear, and speak

OpenAI says ChatGPT now supports seeing, hearing, and speaking. The post body is empty, so it does not disclose model versions, rollout timing, regional limits, pricing, or API scope. The real watchpoints are voice latency, vision limits, and access paths.

Why it matters: This is a substantive OpenAI product update: the title confirms vision input, voice input, and speech output for ChatGPT, so HKR-H/K/R all pass. The copy provided here omits tiers, rollout scope, latency, and pricing, which keeps it at 88 rather than the top of the band.