Skip to content

Voice & audio

AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

116 picksRelated topicsMultimodalAI videoProduct updates

Latest picks

81–100 of 116

May 5Tuesday

r/LocalLLaMA

vibevoice.cpp: Microsoft VibeVoice ported to ggml/C++ with no Python at inference

LocalAI released vibevoice.cpp, a ggml/C++ port of Microsoft VibeVoice for CPU, CUDA, Metal, and Vulkan inference. TTS uses a 30s reference clip for 24kHz cloned speech; ASR uses a 7B model with diarized JSON and was tested on 17min audio. The key constraint is memory: 17min CPU Q8_0 peaks near 26GB, with no streaming output yet.

Why it matters: HKR-H/K/R all pass: a practical open-source VibeVoice C++ port with concrete runtime numbers. Reddit-source scope and niche audio deployment keep it in the 72–77 featured band, not same-day must-write.

May 4Monday

r/LocalLLaMA

Gemma 4 E2B runs well on an 8GB Android phone, powering a private voice notes app

A Reddit user ran Gemma 4 E2B locally on an 8GB OnePlus CE 5 and built a private voice notes app. Whisper Small 244MB transcribes, Gemma 4 E2B 2.4GB splits and tags, and a 10-15s note takes 12-15s end to end. Search uses query expansion, FTS lanes, RRF, and optional Gemma top-K reranking with a 15s fallback.

Why it matters: HKR-H/K/R all pass, but this is a Reddit first-person build, not an official Google release. Concrete hardware, latency, model size, and retrieval details place it near the top of the tutorial band.

May 2Saturday

Hacker News front page

Spotify Adds 'Verified' Badges to Distinguish Human Artists from AI

Spotify added 'Verified' badges for human artists to distinguish them from AI, per the title. The RSS snippet does not disclose the verification process, rollout scope, timing, or review criteria.

Why it matters: HKR-H and HKR-R are strong: human-vs-AI artist labeling is clickable and identity-charged. HKR-K is thin because only the badge fact is disclosed; no audit mechanism or rollout scope. Mid-weight product update, not P1.

May 1Friday

QbitAI · WeChat

He Used AI to Run a Music Festival About Not Doing a PhD

Bilibili creator Huntunpi Qiezong made 42 AI-generated “Don’t Do a PhD” songs, passing 50 million views. One track took over 100 generations, racing Suno, MiniMax Music, HeartMuLa, and ACE-Step. The key signal is the human curation cost in AI music workflows.

Why it matters: HKR-H/K/R all pass: the hook is unusual, the post gives concrete counts and workflow details, and the human curation cost resonates with creators. This is a strong case study, not a model or platform release, so it fits 78–84.

Bloomberg Technology

The Audio Industry Is Grappling with the Rise of ‘Podslop’

Podcast Index says 39% of new podcasts over nine days were likely AI-generated. The title cites industry concern over “podslop”; the post does not disclose detection methods, sample size, or platform split.

Why it matters: HKR-H/K/R pass, but the post lacks detection method, sample size, and platform split. Bloomberg plus Podcast Index’s 39% claim clears featured, not a major industry event.

Apr 30Thursday

Google DeepMind

Google DeepMind announces AI co-clinician medical research program

Google DeepMind announced an AI co-clinician research program, exploring how AI agents can assist patient care under a doctor's clinical supervision. In a blinded evaluation of 98 real primary care queries, the system made no critical errors in 97 cases, and doctors preferred its answers over mainstream evidence synthesis tools. On 140 consultation skills, it matched or beat primary care physicians on 68, but expert physicians were still better overall at spotting red flags and key physical exams.

Why it matters: Google DeepMind published its AI co-clinician research program and a multimodal consultation evaluation, showing where medical agents' abilities currently end.

r/LocalLLaMA

Building a fully local PDF-to-audiobook workflow with Kokoro 82M, Qwen and llama.cpp

Reddit user purellmagents shared a local PDF-to-audiobook workflow using Kokoro 82M, Qwen 3.5 0.8B/2B, and llama.cpp. The Tauri 2.0 app runs on an M1 Mac, reads 15 initial sentences, then prepares the next 15. The hard parts are PDF-text alignment, code snippets, tables, and first-generation latency.

Why it matters: HKR-H/K/R all pass, but this is a Reddit personal workflow, not a model or platform release. Specific components and the 15-sentence pipeline keep it at the low featured band.

Apr 29Wednesday

Xinzhiyuan · WeChat

Google Translate Turns 20 as Pichai Highlights Four AI Generations

Google Translate turned 20 on April 28, and Pichai said it now has 1B monthly users. The post traces four AI phases: SMT, GNMT, PaLM 2, and Gemini 2.5 Flash Native Audio, including 110 languages added in 2024. The key shift is native speech-to-speech translation that preserves intonation, pacing, and pitch.

Why it matters: HKR-H/K/R all pass, but the core event is a Google Translate anniversary and architecture recap, not a clear launch. The 1B MAU, 110-language expansion, and native speech-to-speech detail justify featured at the 72–77 band.

X · @dotey

Microsoft VibeVoice-ASR tested on Mac for a one-hour podcast

Simon Willison ran 4-bit VibeVoice-ASR on an M5 Max MacBook Pro and transcribed a one-hour podcast in 8m45s. The 9B MIT-licensed model supports 60-minute audio, 50+ languages, and structured speaker output. Memory is the constraint: prefill peaked at 61.5GB, making 32GB laptops impractical.

Why it matters: HKR-H/K/R all pass: Simon Willison’s local test gives speed, parameter size, and memory peak that practitioners can act on. It is a single benchmark, not a fresh model launch, so it stays at the featured threshold.

Apr 28Tuesday

QbitAI · WeChat

ModelBest Releases MiniCPM-o 4.5 Technical Report for Consumer-GPU Deployment

ModelBest, OpenBMB, Tsinghua THUNLP and THUMAI released the MiniCPM-o 4.5 technical report, covering a roughly 9B-parameter model. It supports video, audio and text streams; a 12GB RTX 5070 runs full-duplex mode at RTF 0.4. The key mechanism is Omni-Flow: a unified timeline with time-division multiplexing, without external VAD.

Why it matters: HKR-H/K/R all pass: a 9B omni model runs full-duplex on a 12GB RTX 5070 with RTF 0.4, using Omni-Flow timeline alignment. It is below a frontier-lab flagship release, so 78–84 fits.

QbitAI · WeChat

Xiaomi open-sources MiMo-V2.5 series; Pro builds a macOS-like desktop in 4 hours

Xiaomi open-sourced MiMo-V2.5 weights, covering Pro Agent, multimodal base, TTS, and ASR models. MiMo-V2.5-Pro built a 54-app macOS-like desktop in 4 hours without human takeover; it scored 233/233 on SysY with 672 tool calls in 4.3 hours. Key details for practitioners are the 1M context, 100T-token program, and free Agent-framework access.

Why it matters: HKR-H/K/R all pass: Xiaomi open-sourced MiMo-V2.5 weights with concrete agent and coding-task numbers. Domestic flagship model release bump puts it in the must-write same-day band.

Apr 25Saturday

Hacker News front page

Google Flow Music

Google Flow Music launched a web creation entry with six sections: songs, playlists, Spaces, videos, projects, and Turntable. The page says Producer creates full songs with Lyria 3, and AI music videos use Veo. Pricing, regions, model specs, and rights terms are not disclosed.

Why it matters: HKR-H/K/R pass: a Google AI music web product tying Lyria 3 and Veo is clickable, concrete, and competitive. Score stays in 72–77 because price, regions, rights, and model specs are not disclosed.

Apr 24Friday

Hacker News front page

Refuse to let your doctor record you

Emily M. Bender and Decca Muldowney give 9 reasons to refuse AI medical scribes. The tools record visits and draft chart notes, raising privacy, consent, automation-bias, and speech-recognition disparity risks. The key concern is clinics converting saved time into more visits.

Why it matters: HKR-H/K/R all pass: the title has a sharp healthcare-AI hook, the post explains the audio-to-chart-note mechanism and 9 risk areas, and privacy/consent will travel. It is commentary without hard data, so it stays in the 72–77 band.

Apr 23Thursday

The Verge · AI

Google Meet will take AI notes for in-person meetings too

Google expanded Gemini notetaking to in-person meetings and added support for Zoom and Microsoft Teams. The post confirms summaries and transcripts; in-person support had previously been limited to Android alpha users. Google also says it works for impromptu meetings outside meeting rooms, which matters because the recorder is no longer confined to native Meet calls.

Apr 20Monday

Hacker News front page

Deezer says 44% of songs uploaded to its platform daily are AI-generated

Deezer says 44% of songs uploaded to its platform each day are AI-generated, with the headline disclosing the 44% share. The RSS snippet does not disclose the measurement period, detection method, sample size, or any enforcement policy.

Why it matters: This clears HKR-H/K/R on a striking platform-level stat and strong resonance around AI-content flooding and rights. It stays at 76 because the claim is a single company disclosure; detection method, timeframe, and enforcement details are not disclosed.

Apr 17Friday

X · @dotey

browser-use open-sources video-use, a Claude Code skill that turns raw camera footage into edited videos

browser-use released video-use, a Claude Code skill that turns raw footage into a final.mp4 automatically. It converts footage into ElevenLabs word-level timestamp transcripts, shrinking one asset to about 12KB; the post says feeding frames directly would cost about 45 million tokens. The key detail is the structured editing pipeline: the model mostly reads text, uses timeline images only at uncertain cuts, and runs up to 3 self-check repair passes after rendering.

Why it matters: Strong HKR-H/K/R: the result is instantly clickable, and the post includes a concrete text-first editing architecture with 12KB vs about 45M-token economics. Kept below higher bands because this is a builder-facing Claude Code skill, not a platform-level release.

Apr 16Thursday

Google DeepMind

Google DeepMind releases Gemini 3.1 Flash TTS

Google DeepMind released Gemini 3.1 Flash TTS, a text-to-speech model built around controllability and expressiveness. It is in preview on the Gemini API, Google AI Studio, Vertex AI and Google Vids.

Why it matters: The post covers the new model's audio-tag controls, Elo scores and preview entry points, so you can judge how controllable speech generation has become.

Apr 14Tuesday

最佳拍档 (BestPartners)

Global GPU shortage worsens: H100 rental prices rose nearly 40% in five months

SemiAnalysis says Nvidia H100 one-year rental pricing rose from $1.70 to $2.35 per GPU-hour between Oct 2025 and Mar 2026, up nearly 40% in five months. The post attributes this to Anthropic-driven demand, multi-agent and media generation workloads, and memory cost spikes, with LPDDR5 and DDR5 contract prices up about 4x and 5x year over year; much new capacity is already prebooked. The key variable is the supply gap, not Blackwell refreshes alone.

Why it matters: Strong HKR-H/K/R: the story has a sharp price-shock hook, concrete market data, and clear resonance with compute-cost anxiety. It stays below P1 because this is a secondary video synthesis of a SemiAnalysis report, not a primary company or product announcement.

Apr 8Wednesday

QbitAI · WeChat

Xiaomi unveils two AI audio frameworks: Any2Speech and Midasheng-audio-generate

Xiaomi's large-model application team introduced Xiaomi Any2Speech and Midasheng-audio-generate. Any2Speech generates up to about 10 minutes per inference, while the other model turns one text prompt into mixed audio with speech, music, and ambient sound. The post names GST labeling, dual-path planning with dimension dropout, Flow Matching, and five-field structured labels; benchmark scores, training scale, and commercial terms are not disclosed.

Why it matters: Xiaomi released two audio-generation frameworks with a clear hook and concrete mechanisms, so HKR-H and HKR-K pass. HKR-R is weaker because benchmark results, training data scale, open-source status, and commercial terms are not disclosed, so this sits at the low end of featured.

QbitAI · WeChat

Free open-source 2B Chinese speech model reproduces Mangzhuang Ren with high-speed tonguetwisters

ModelBest, OpenBMB, and Tsinghua University released VoxCPM 2, a 2B open speech model that supports 9 Chinese dialects, 30 foreign languages, and 48kHz audio. The post says generation often finishes within 1 second, recommends reference audio of at least 5 seconds, and supports denoising, LoRA, and full fine-tuning; the key detail is its tokenizer-free diffusion autoregressive continuous representation design.

Why it matters: This is a substantive open-source speech release, not a thin demo: the post gives 2B, 48kHz, 9 Chinese dialects, 30 languages, ref audio ≥5s, and a tokenizer-free route. HKR-H/K/R all pass, but the event is not large enough for a must-write P1.