Skip to content

#语音

0 today

May 8Friday

AI HOT (Curated Pool)

Apple's First AI Wearable: Camera-Equipped AirPods Enter DVT Stage

Apple’s camera-equipped AirPods have entered DVT, with launch possible in September. Each earbud uses a low-res camera for visual Q&A with the upgraded Siri. The post cites Google Gemini support and a data-upload indicator light.

Why it matters: HKR-H/K/R all pass, but this is an unconfirmed hardware rumor, not an Apple launch. DVT status, camera design, and Gemini dependency keep it in the low featured band.

AI HOT (Curated Pool)

OpenAI launches official openai-cli for terminal API calls

OpenAI open-sourced openai-cli for direct API calls from the terminal. The Apache 2.0 tool installs via Homebrew or Go and covers Responses API, structured output, image editing, transcription, and key config. The key detail is Agent workflows using cloud tools like web search and code interpreter.

Why it matters: HKR-H/K/R all pass: official OpenAI terminal tooling is clickable, with concrete install/license/API details and workflow resonance. It is still a developer tooling update, not a model or major capability release, so 76 fits the featured threshold.

The Verge · AI

Apple’s AirPods with cameras for AI are reportedly close to production

Mark Gurman says Apple’s camera-equipped AirPods are in DVT, one step before PVT. Testers are using prototypes; the cameras capture low-resolution visual input, not photos or video, for Siri queries like ingredient prompts.

Why it matters: HKR-H/K/R all pass: Gurman/The Verge provides a concrete DVT-stage Apple AI hardware update. It is still pre-production, not a launch, so it stays in the 72–77 band.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.

May 5Tuesday

r/LocalLLaMA

vibevoice.cpp: Microsoft VibeVoice ported to ggml/C++ with no Python at inference

LocalAI released vibevoice.cpp, a ggml/C++ port of Microsoft VibeVoice for CPU, CUDA, Metal, and Vulkan inference. TTS uses a 30s reference clip for 24kHz cloned speech; ASR uses a 7B model with diarized JSON and was tested on 17min audio. The key constraint is memory: 17min CPU Q8_0 peaks near 26GB, with no streaming output yet.

Why it matters: HKR-H/K/R all pass: a practical open-source VibeVoice C++ port with concrete runtime numbers. Reddit-source scope and niche audio deployment keep it in the 72–77 featured band, not same-day must-write.

May 4Monday

r/LocalLLaMA

Gemma 4 E2B runs well on an 8GB Android phone, powering a private voice notes app

A Reddit user ran Gemma 4 E2B locally on an 8GB OnePlus CE 5 and built a private voice notes app. Whisper Small 244MB transcribes, Gemma 4 E2B 2.4GB splits and tags, and a 10-15s note takes 12-15s end to end. Search uses query expansion, FTS lanes, RRF, and optional Gemma top-K reranking with a 15s fallback.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person build, not an official Google release. Concrete hardware, latency, model size, and retrieval details place it near the top of the tutorial band.

May 2Saturday

Hacker News front page

Spotify Adds 'Verified' Badges to Distinguish Human Artists from AI

Spotify added 'Verified' badges for human artists to distinguish them from AI, per the title. The RSS snippet does not disclose the verification process, rollout scope, timing, or review criteria.

Why it matters: HKR-H and HKR-R are strong: human-vs-AI artist labeling is clickable and identity-charged. HKR-K is thin because only the badge fact is disclosed; no audit mechanism or rollout scope. Mid-weight product update, not P1.

May 1Friday

QbitAI · WeChat

He Used AI to Run a Music Festival About Not Doing a PhD

Bilibili creator Huntunpi Qiezong made 42 AI-generated “Don’t Do a PhD” songs, passing 50 million views. One track took over 100 generations, racing Suno, MiniMax Music, HeartMuLa, and ACE-Step. The key signal is the human curation cost in AI music workflows.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post gives concrete counts and workflow details, and the human curation cost resonates with creators. This is a strong case study, not a model or platform release, so it fits 78–84.

Bloomberg Technology

The Audio Industry Is Grappling with the Rise of ‘Podslop’

Podcast Index says 39% of new podcasts over nine days were likely AI-generated. The title cites industry concern over “podslop”; the post does not disclose detection methods, sample size, or platform split.

Why it matters: HKR-H/K/R pass, but the post lacks detection method, sample size, and platform split. Bloomberg plus Podcast Index’s 39% claim clears featured, not a major industry event.

Apr 30Thursday

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

r/LocalLLaMA

Building a fully local PDF-to-audiobook workflow with Kokoro 82M, Qwen and llama.cpp

Reddit user purellmagents shared a local PDF-to-audiobook workflow using Kokoro 82M, Qwen 3.5 0.8B/2B, and llama.cpp. The Tauri 2.0 app runs on an M1 Mac, reads 15 initial sentences, then prepares the next 15. The hard parts are PDF-text alignment, code snippets, tables, and first-generation latency.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal workflow, not a model or platform release. Specific components and the 15-sentence pipeline keep it at the low featured band.

Apr 29Wednesday

Xinzhiyuan · WeChat

Google Translate Turns 20 as Pichai Highlights Four AI Generations

Google Translate turned 20 on April 28, and Pichai said it now has 1B monthly users. The post traces four AI phases: SMT, GNMT, PaLM 2, and Gemini 2.5 Flash Native Audio, including 110 languages added in 2024. The key shift is native speech-to-speech translation that preserves intonation, pacing, and pitch.

Why it matters: HKR-H/K/R all pass, but the core event is a Google Translate anniversary and architecture recap, not a clear launch. The 1B MAU, 110-language expansion, and native speech-to-speech detail justify featured at the 72–77 band.

X · @dotey

Microsoft VibeVoice-ASR tested on Mac for a one-hour podcast

Simon Willison ran 4-bit VibeVoice-ASR on an M5 Max MacBook Pro and transcribed a one-hour podcast in 8m45s. The 9B MIT-licensed model supports 60-minute audio, 50+ languages, and structured speaker output. Memory is the constraint: prefill peaked at 61.5GB, making 32GB laptops impractical.

Why it matters: HKR-H/K/R all pass: Simon Willison’s local test gives speed, parameter size, and memory peak that practitioners can act on. It is a single benchmark, not a fresh model launch, so it stays at the featured threshold.

Apr 28Tuesday

QbitAI · WeChat

ModelBest Releases MiniCPM-o 4.5 Technical Report for Consumer-GPU Deployment

ModelBest, OpenBMB, Tsinghua THUNLP and THUMAI released the MiniCPM-o 4.5 technical report, covering a roughly 9B-parameter model. It supports video, audio and text streams; a 12GB RTX 5070 runs full-duplex mode at RTF 0.4. The key mechanism is Omni-Flow: a unified timeline with time-division multiplexing, without external VAD.

Why it matters: HKR-H/K/R all pass: a 9B omni model runs full-duplex on a 12GB RTX 5070 with RTF 0.4, using Omni-Flow timeline alignment. It is below a frontier-lab flagship release, so 78–84 fits.

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

Apr 25Saturday

Hacker News front page

Google Flow Music

Google Flow Music launched a web creation entry with six sections: songs, playlists, Spaces, videos, projects, and Turntable. The page says Producer creates full songs with Lyria 3, and AI music videos use Veo. Pricing, regions, model specs, and rights terms are not disclosed.

Why it matters: HKR-H/K/R pass: a Google AI music web product tying Lyria 3 and Veo is clickable, concrete, and competitive. Score stays in 72–77 because price, regions, rights, and model specs are not disclosed.

Apr 24Friday

Hacker News front page

Refuse to let your doctor record you

Emily M. Bender and Decca Muldowney give 9 reasons to refuse AI medical scribes. The tools record visits and draft chart notes, raising privacy, consent, automation-bias, and speech-recognition disparity risks. The key concern is clinics converting saved time into more visits.

Why it matters: HKR-H/K/R all pass: the title has a sharp healthcare-AI hook, the post explains the audio-to-chart-note mechanism and 9 risk areas, and privacy/consent will travel. It is commentary without hard data, so it stays in the 72–77 band.

Apr 23Thursday

The Verge · AI

Google Meet will take AI notes for in-person meetings too

Google expanded Gemini notetaking to in-person meetings and added support for Zoom and Microsoft Teams. The post confirms summaries and transcripts; in-person support had previously been limited to Android alpha users. Google also says it works for impromptu meetings outside meeting rooms, which matters because the recorder is no longer confined to native Meet calls.

Apr 20Monday

Hacker News front page

Deezer says 44% of songs uploaded to its platform daily are AI-generated

Deezer says 44% of songs uploaded to its platform each day are AI-generated, with the headline disclosing the 44% share. The RSS snippet does not disclose the measurement period, detection method, sample size, or any enforcement policy.

Why it matters: This clears HKR-H/K/R on a striking platform-level stat and strong resonance around AI-content flooding and rights. It stays at 76 because the claim is a single company disclosure; detection method, timeframe, and enforcement details are not disclosed.

Apr 17Friday

X · @dotey

browser-use open-sources video-use, a Claude Code skill that turns raw camera footage into edited videos

browser-use released video-use, a Claude Code skill that turns raw footage into a final.mp4 automatically. It converts footage into ElevenLabs word-level timestamp transcripts, shrinking one asset to about 12KB; the post says feeding frames directly would cost about 45 million tokens. The key detail is the structured editing pipeline: the model mostly reads text, uses timeline images only at uncertain cuts, and runs up to 3 self-check repair passes after rendering.

Why it matters: Strong HKR-H/K/R: the result is instantly clickable, and the post includes a concrete text-first editing architecture with 12KB vs about 45M-token economics. Kept below higher bands because this is a builder-facing Claude Code skill, not a platform-level release.

Apr 16Thursday

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

Apr 8Wednesday

QbitAI · WeChat

Xiaomi unveils two AI audio frameworks: Any2Speech and Midasheng-audio-generate

Xiaomi's large-model application team introduced Xiaomi Any2Speech and Midasheng-audio-generate. Any2Speech generates up to about 10 minutes per inference, while the other model turns one text prompt into mixed audio with speech, music, and ambient sound. The post names GST labeling, dual-path planning with dimension dropout, Flow Matching, and five-field structured labels; benchmark scores, training scale, and commercial terms are not disclosed.

Why it matters: Xiaomi released two audio-generation frameworks with a clear hook and concrete mechanisms, so HKR-H and HKR-K pass. HKR-R is weaker because benchmark results, training data scale, open-source status, and commercial terms are not disclosed, so this sits at the low end of featured.

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.

Apr 3Friday

X · @OpenAI

ChatGPT is now available in CarPlay

OpenAI is rolling out ChatGPT in CarPlay to iPhone users on iOS 26.4+ where CarPlay is supported. The post confirms voice mode is available in-car, but does not disclose regions, vehicle coverage, or feature limits. The key shift is distribution into the driving interface, not a new model launch.

Why it matters: This matters more as a distribution-surface shift than a model update. HKR-H and HKR-R pass on the CarPlay hook and assistant-entry competition; HKR-K stays limited because the post gives iOS 26.4+ rollout only, not regions, car support, or full feature bounds.

Mar 24Tuesday

Mistral AI

Mistral AI releases Voxtral TTS, a 4B-parameter speech model

Mistral AI released Voxtral TTS, its first text-to-speech model. It has 4B parameters and supports nine languages: English, French, German, Spanish, Dutch, Portuguese, Italian, Hindi and Arabic. It handles emotional expression and zero-shot cross-lingual voice adaptation.

Why it matters: The 4B size, nine languages, 70ms latency and pricing give readers a basis for judging cost and model choice in enterprise voice agents.

Feb 14Saturday

MIT Technology Review · AI

ALS stole this musician’s voice. AI let him sing again.

Patrick Darling, 32, returned to the stage on February 11 in London after two years without singing, using an AI voice clone rebuilt from old recordings. The post says speech cloning typically needs about 10 minutes of clean audio; his singing clone was built from noisy phone clips and kitchen recordings, then refined with Eleven Music over about six weeks. The practical signal is access, not sentiment: ElevenLabs offers the tools free to people who lost their voices to ALS and similar conditions, but the post does not disclose model details.

Why it matters: HKR-H/K/R all land: the hook is strong, the story gives concrete reproducible details, and the use case hits accessibility plus voice-rights nerves. Still, this is a strong application story, not a major model, product, or research release, so it stays in low featured.

Feb 9Monday

36Kr (direct RSS)

Voice Ask is live: why is Xiaohongshu pushing search-by-question?

Xiaohongshu fully launched Voice Ask on Jan. 27, letting users long-press to speak on the search page and get structured answers distilled from in-app user experience posts. The post says it can handle 3-minute spoken queries, foreign languages, and dialects, but does not disclose the model, ASR stack, latency, or accuracy. The real shift is from 3-4 character keyword search to longer spoken questions, widening search intent capture and scenario coverage.

36Kr (direct RSS)

Former Baichuan co-founder Jiao Ke bets on AI audio to build AI hosts

Jiao Ke said Laifu Radio now has 15 Chinese AI hosts and 2 English ones, and raised over $10 million across two rounds by H2 2025. He said users average about 30 minutes per day, AI can prepare timely audio in under an hour, and the team treats DTU plus long-memory infra as the key moat. The real bet is not an AI podcast tool but interactive AI hosts that remember user preferences; the post also says it is working with some automakers on in-car personalized AI radio.

Why it matters: HKR-H lands because the story reframes audio AI as persistent hosts, not a podcast tool. HKR-K is strong on numbers and mechanism; HKR-R lands via memory plus in-car distribution. Early-stage company scope keeps it at featured, not p1.

Feb 5Thursday

Mistral AI

Mistral releases Voxtral Transcribe 2 speech-to-text model family

Mistral released Voxtral Transcribe 2, a family of two speech-to-text models: Voxtral Mini Transcribe V2 for batch transcription and Voxtral Realtime for live use.

Why it matters: The post gives latency, pricing and open-source licensing for both transcription models, enough to judge the options for real-time voice applications.

Sep 30, 2025Tuesday

OpenAI News

Sora 2 System Card

OpenAI published the Sora 2 System Card on September 30, 2025, and said the video-audio generation model will launch first via limited invites on sora.com and a standalone iOS app. The post confirms no video uploads and no image uploads with photorealistic people at launch; API timing, pricing, and benchmark scores are not disclosed.

Why it matters: This lands in the 78–84 band. HKR-H comes from the Sora 2 + iOS app hook; HKR-K from concrete launch limits and safety rules; HKR-R from competition and likeness-abuse nerves. It stays below P1 because price, eval scores, context details, and API timing are not disclosed.

OpenAI News

Sora 2 is here

OpenAI released Sora 2 on September 30, 2025 and launched a social iOS app called Sora built on the model. The post says it generates video with synced dialogue and sound effects, and its “characters” feature uses a one-time video and audio recording to verify identity and insert a real person’s likeness; pricing, generation limits, and rollout regions are not disclosed. The key shift is from model demo to a consumer app with a feed, teen limits, and parental controls.

Why it matters: This is a same-day write: OpenAI shipped a flagship video/audio model and attached it to a standalone app, so HKR-H/K/R all clear. The post gives real product facts like synced dialogue and sound effects, but missing price, duration caps, and rollout details keeps it below 90.

Aug 28, 2025Thursday

OpenAI News

Introducing gpt-realtime and Realtime API updates for production voice agents

OpenAI released the speech-to-speech model gpt-realtime and made the Realtime API generally available, adding remote MCP server support, image input, and SIP phone calling. The post reports 82.8% on Big Bench Audio versus 65.6% for the December 2024 model, and 30.5% on the audio MultiChallenge benchmark versus 20.6%. The key change is that tool access and phone connectivity now ship in the same production API.

Why it matters: This is a substantive OpenAI model + API release, not a minor refresh. HKR-H/K/R all pass: the release has a clear hook, hard benchmark deltas, and direct deployment impact for production voice agents, so it reaches p1.

Jul 17, 2025Thursday

Mistral AI

Mistral adds Deep Research, voice mode and more to Le Chat

Mistral rolled out a batch of new Le Chat features: a preview Deep Research mode, a voice mode powered by the new Voxtral speech model, a multilingual thinking mode backed by the Magistral reasoning model, Projects for organizing conversations, and advanced image editing built with Black Forest Labs.

Why it matters: Mistral announced five Le Chat features at once, so readers can see how its research, voice and image-editing abilities fit together.

Jul 15, 2025Tuesday

Mistral AI

Mistral releases Voxtral speech-understanding models in 24B and 3B

Mistral released Voxtral, a speech-understanding model in 24B and 3B versions, both open-sourced under Apache 2.0 and available via API. It supports a 32k token context, handling up to 30 minutes of transcription or 40 minutes of understanding, with built-in Q&A and summarization, multilingual recognition and voice function calling. It keeps the text abilities of Mistral Small 3.1.

Why it matters: Mistral open-sourced two speech-understanding models with 32k context and function calling, priced at less than half comparable APIs, which helps when picking a speech stack.

Mar 21, 2025Friday

OpenAI News

Early methods for studying affective use and emotional well-being on ChatGPT

OpenAI and MIT Media Lab studied affective use on ChatGPT with two tracks: nearly 40 million interactions in an observational analysis and a 4-week RCT with nearly 1,000 participants. The post says emotional engagement is rare overall and concentrated in a small subset of heavy Advanced Voice Mode users; the provided body does not fully disclose all quantitative well-being results. Watch subgroup effects, not platform averages.

Mar 20, 2025Thursday

OpenAI News

Introducing next-generation audio models in the API

OpenAI released three API audio models on March 20, 2025: gpt-4o-transcribe, gpt-4o-mini-transcribe, and gpt-4o-mini-tts. The post says the STT models beat Whisper v2 and v3 on FLEURS and other benchmarks across 100+ languages, while the TTS model adds style control but stays limited to monitored preset synthetic voices. The key shift is controllable TTS plus lower WER; the post does not disclose pricing or latency figures.

Why it matters: OpenAI shipped 3 API audio models with concrete benchmark and mechanism details, so HKR-H/K/R all pass and it clears featured. I kept it at 84, not 85+, because price, latency, and a fuller benchmark table are not disclosed.

Oct 1, 2024Tuesday

OpenAI News

Introducing the Realtime API

OpenAI launched a public beta of the Realtime API on Oct. 1, 2024 for all paid developers, using a persistent WebSocket to stream low-latency speech-to-speech interactions with GPT-4o. It supports function calling and interruption handling, priced at $5/1M text input tokens and $100/1M audio input tokens; the post also says audio I/O for Chat Completions would arrive in the following weeks.

Why it matters: OpenAI moved voice apps from stitched ASR+TTS calls to a persistent GPT-4o session, with function calling, interruption handling, and published audio/token pricing. HKR-H/K/R all pass, so this is a same-day must-write developer platform update and clears p1.

Aug 8, 2024Thursday

OpenAI News

GPT-4o System Card

OpenAI published the GPT-4o System Card on August 8, 2024, reporting 3 of 4 Preparedness categories as low risk and persuasion as borderline medium. The post says GPT-4o accepts text, audio, image, and video inputs, responds to audio in as little as 232 ms with a 320 ms average, and is 50% cheaper than GPT-4 Turbo in the API. The key issue for practitioners is voice safety: the card names unauthorized voice generation, speaker identification, and sensitive trait attribution, and says only models with post-mitigation scores at medium or below can be deployed.

Why it matters: This is not a routine post: it adds concrete preparedness ratings, 232ms voice latency, and a clear deployment threshold. HKR-H/K/R all pass, but it is a safety disclosure rather than a new model or major launch, so it lands as featured, not p1.

Sep 25, 2023Monday

OpenAI News

ChatGPT can now see, hear, and speak

OpenAI says ChatGPT now supports seeing, hearing, and speaking. The post body is empty, so it does not disclose model versions, rollout timing, regional limits, pricing, or API scope. The real watchpoints are voice latency, vision limits, and access paths.

Why it matters: This is a substantive OpenAI product update: the title confirms vision input, voice input, and speech output for ChatGPT, so HKR-H/K/R all pass. The copy provided here omits tiers, rollout scope, latency, and pricing, which keeps it at 88 rather than the top of the band.