Skip to content

Voice & audio

AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

116 picksRelated topicsMultimodalAI videoProduct updates

Latest picks

101–116 of 116

Apr 3Friday

X · @OpenAI

ChatGPT is now available in CarPlay

OpenAI is rolling out ChatGPT in CarPlay to iPhone users on iOS 26.4+ where CarPlay is supported. The post confirms voice mode is available in-car, but does not disclose regions, vehicle coverage, or feature limits. The key shift is distribution into the driving interface, not a new model launch.

Why it matters: This matters more as a distribution-surface shift than a model update. HKR-H and HKR-R pass on the CarPlay hook and assistant-entry competition; HKR-K stays limited because the post gives iOS 26.4+ rollout only, not regions, car support, or full feature bounds.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Feb 14Saturday

MIT Technology Review · AI

ALS stole this musician’s voice. AI let him sing again.

Patrick Darling, 32, returned to the stage on February 11 in London after two years without singing, using an AI voice clone rebuilt from old recordings. The post says speech cloning typically needs about 10 minutes of clean audio; his singing clone was built from noisy phone clips and kitchen recordings, then refined with Eleven Music over about six weeks. The practical signal is access, not sentiment: ElevenLabs offers the tools free to people who lost their voices to ALS and similar conditions, but the post does not disclose model details.

Why it matters: HKR-H/K/R all land: the hook is strong, the story gives concrete reproducible details, and the use case hits accessibility plus voice-rights nerves. Still, this is a strong application story, not a major model, product, or research release, so it stays in low featured.

Feb 9Monday

36Kr (direct RSS)

Voice Ask is live: why is Xiaohongshu pushing search-by-question?

Xiaohongshu fully launched Voice Ask on Jan. 27, letting users long-press to speak on the search page and get structured answers distilled from in-app user experience posts. The post says it can handle 3-minute spoken queries, foreign languages, and dialects, but does not disclose the model, ASR stack, latency, or accuracy. The real shift is from 3-4 character keyword search to longer spoken questions, widening search intent capture and scenario coverage.

36Kr (direct RSS)

Former Baichuan co-founder Jiao Ke bets on AI audio to build AI hosts

Jiao Ke said Laifu Radio now has 15 Chinese AI hosts and 2 English ones, and raised over $10 million across two rounds by H2 2025. He said users average about 30 minutes per day, AI can prepare timely audio in under an hour, and the team treats DTU plus long-memory infra as the key moat. The real bet is not an AI podcast tool but interactive AI hosts that remember user preferences; the post also says it is working with some automakers on in-car personalized AI radio.

Why it matters: HKR-H lands because the story reframes audio AI as persistent hosts, not a podcast tool. HKR-K is strong on numbers and mechanism; HKR-R lands via memory plus in-car distribution. Early-stage company scope keeps it at featured, not p1.

Feb 5Thursday

Mistral AI

Mistral releases Voxtral Transcribe 2 speech-to-text model family

Mistral released Voxtral Transcribe 2, a family of two speech-to-text models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live use.

Why it matters: The post gives latency, pricing and open-source licensing for both transcription models, enough to judge the options for real-time voice applications.

Sep 30, 2025Tuesday

OpenAI News

Sora 2 System Card

OpenAI published the Sora 2 System Card on September 30, 2025, and said the video-audio generation model will launch first via limited invites on sora.com and a standalone iOS app. The post confirms no video uploads and no image uploads with photorealistic people at launch; API timing, pricing, and benchmark scores are not disclosed.

Why it matters: This lands in the 78–84 band. HKR-H comes from the Sora 2 + iOS app hook; HKR-K from concrete launch limits and safety rules; HKR-R from competition and likeness-abuse nerves. It stays below P1 because price, eval scores, context details, and API timing are not disclosed.

OpenAI News

Sora 2 is here

OpenAI released Sora 2 on September 30, 2025 and launched a social iOS app called Sora built on the model. The post says it generates video with synced dialogue and sound effects, and its “characters” feature uses a one-time video and audio recording to verify identity and insert a real person’s likeness; pricing, generation limits, and rollout regions are not disclosed. The key shift is from model demo to a consumer app with a feed, teen limits, and parental controls.

Why it matters: This is a same-day write: OpenAI shipped a flagship video/audio model and attached it to a standalone app, so HKR-H/K/R all clear. The post gives real product facts like synced dialogue and sound effects, but missing price, duration caps, and rollout details keeps it below 90.

Aug 28, 2025Thursday

OpenAI News

Introducing gpt-realtime and Realtime API updates for production voice agents

OpenAI released the speech-to-speech model gpt-realtime and made the Realtime API generally available, adding remote MCP server support, image input, and SIP phone calling. The post reports 82.8% on Big Bench Audio versus 65.6% for the December 2024 model, and 30.5% on the audio MultiChallenge benchmark versus 20.6%. The key change is that tool access and phone connectivity now ship in the same production API.

Why it matters: This is a substantive OpenAI model + API release, not a minor refresh. HKR-H/K/R all pass: the release has a clear hook, hard benchmark deltas, and direct deployment impact for production voice agents, so it reaches p1.

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

Jul 15, 2025Tuesday

Mistral AI

Mistral releases Voxtral speech-understanding models in 24B and 3B

Mistral released Voxtral, a speech-understanding model in 24B and 3B versions, both open-sourced under Apache 2.0 and available via API. It supports a 32k token context, handling up to 30 minutes of transcription or 40 minutes of understanding, with built-in Q&A and summarization, multilingual recognition and voice function calling. It keeps the text abilities of Mistral Small 3.1.

Why it matters: Mistral open-sourced two speech-understanding models with 32k context and function calling, priced at less than half comparable APIs, which helps when picking a speech stack.

Mar 21, 2025Friday

OpenAI News

Early methods for studying affective use and emotional well-being on ChatGPT

OpenAI and MIT Media Lab studied affective use on ChatGPT with two tracks: nearly 40 million interactions in an observational analysis and a 4-week RCT with nearly 1,000 participants. The post says emotional engagement is rare overall and concentrated in a small subset of heavy Advanced Voice Mode users; the provided body does not fully disclose all quantitative well-being results. Watch subgroup effects, not platform averages.

Mar 20, 2025Thursday

OpenAI News

Introducing next-generation audio models in the API

OpenAI released three API audio models on March 20, 2025: gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-mini-tts. The post says the STT models beat Whisper v2 and v3 on FLEURS and other benchmarks across 100+ languages, while the TTS model adds style control but stays limited to monitored preset synthetic voices. The key shift is controllable TTS plus lower WER; the post does not disclose pricing or latency figures.

Why it matters: OpenAI shipped 3 API audio models with concrete benchmark and mechanism details, so HKR-H/K/R all pass and it clears featured. I kept it at 84, not 85+, because price, latency, and a fuller benchmark table are not disclosed.

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

Aug 8, 2024Thursday

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.

Sep 25, 2023Monday

OpenAI News

ChatGPT can now see, hear, and speak

OpenAI says ChatGPT now supports seeing, hearing, and speaking. The post body is empty, so it does not disclose model versions, rollout timing, regional limits, pricing, or API scope. The real watchpoints are voice latency, vision limits, and access paths.

Why it matters: This is a substantive OpenAI product update: the title confirms vision input, voice input, and speech output for ChatGPT, so HKR-H/K/R all pass. The copy provided here omits tiers, rollout scope, latency, and pricing, which keeps it at 88 rather than the top of the band.