Skip to content

Voice & audio

AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

116 picksRelated topicsMultimodalAI videoProduct updates

Latest picks

21–40 of 116

Jun 10Wednesday

AI HOT (Curated Pool)

Google Gemini 3.5 Live Translate enters public preview with 70+ languages

Google released Gemini 3.5 Live Translate in public preview through the Gemini API, offering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs.

Why it matters: HKR-H/K/R all pass: Google’s speech-to-speech translation API has a clear developer hook and concrete scale numbers. Single X-source detail and missing price, latency benchmarks, and regions keep it at 78.

Jun 9Tuesday

AI HOT (Curated Pool)

Google Releases Gemini 3.5 Live Translate for Real-Time Speech Translation

Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.

Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.

Google DeepMind

Google DeepMind releases Gemini 3.5 Live Translate speech model

Google DeepMind released Gemini 3.5 Live Translate, an audio model for near-real-time speech-to-speech translation across more than 70 languages. It detects the language automatically and preserves the speaker's intonation, rhythm and pitch.

Why it matters: The original gives the model's language coverage, how the live translation works and the rollout pace across products, enough to judge where speech translation is usable.

AI HOT (Curated Pool)

Google DeepMind Releases Gemma 4 12B, a Unified Encoder-Free Multimodal Model

Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.

Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.

Jun 8Monday

AI HOT (Curated Pool)

VoxCPM2 technical report released

OpenBMB released the VoxCPM2 technical report, covering a 2B-parameter speech generation model trained on more than 2 million hours of multilingual speech data, with support for 30 languages and 9 Chinese dialects.

Why it matters: HKR-H/K/R pass via the 2B size, 2M+ training hours, and dialect coverage; the score stays at the low end of 78–84 because the post lacks benchmarks, license terms, and adoption data.

Jun 6Saturday

r/LocalLLaMA

Big week for open AI, with 25+ notable open-weight drops across every modality

Victor M summarized 25+ open-weight model releases in one week, including NVIDIA Nemotron 3 Ultra, a 550B hybrid Mamba-MoE with 55B active parameters and a 1M-token context window.

Why it matters: HKR-H/K/R all pass: the story combines a 25+ open-weight wave with NVIDIA’s 550B, 1M-context Nemotron. Reddit/X sourcing keeps it in the 78-84 band, not p1.

r/LocalLLaMA

dots.tts 2B SOTA TTS from RedNote

RedNote released dots.tts, a 2B-parameter open-source TTS model under Apache 2.0. It uses a fully continuous architecture, supports 48 kHz synthesis and zero-shot voice cloning, and maps text directly to speech without a phoneme pipeline.

Why it matters: HKR-H/K/R pass, but the source is a Reddit summary and the SOTA claim lacks benchmark names or scores. Apache 2.0, 2B params, 48 kHz, and a no-phoneme pipeline justify low featured.

AI HOT (Curated Pool)

Google AI weekly product updates: Nano Banana 2, Co-Scientist, dreambeans, Gemma 4, and more

Google AI announced six updates: Nano Banana 2 is generally available, Gemma 4 12B can run fully offline on laptops, and Magenta RealTime 2 is open source.

Why it matters: HKR-H/K/R all pass: the post bundles six Google AI updates with concrete local and open-source hooks. Lacking benchmarks, licensing, and pricing keeps it below the 78+ good-quality band.

Jun 5Friday

AI HOT (Curated Pool)

Google Magenta RealTime 2 (MRT2) real-time music model released

Google AI for Developers released the open-weight Magenta RealTime 2 music model, supporting MIDI, live text prompts, and gestures, with native MacBook latency under 200 ms.

Why it matters: HKR-H/K/R all pass: Google Magenta MRT2 has a concrete real-time audio hook, open weights, and sub-200ms local latency. It is strong for creative-AI builders, but narrower than a general foundation-model release.

AI HOT (Curated Pool)

Boson AI and LMSYS Release Higgs Audio v3 TTS End-to-End Service Based on SGLang-Omni

Boson AI and LMSYS released the Higgs Audio v3 TTS service with about 4B parameters, a Qwen3-4B backbone, support for 100 languages, streaming synthesis, and text tags for controlling 20+ emotions plus style, rhythm, and sound effects.

Why it matters: HKR-H and HKR-K pass via the 4B/100-language/streaming TTS hook. HKR-R is weaker because the post lacks latency, pricing, and release-form details, so this sits at the lower featured band.

Jun 4Thursday

Synced · WeChat

Google releases Gemma 4 12B for 16GB laptops

Google released Gemma 4 12B, a medium-size model that runs locally with 16GB VRAM or unified memory. It uses an encoder-free multimodal architecture, supports native audio input, ships under Apache 2.0, and includes an MTP draft model for lower latency.

Why it matters: Google’s Gemma 4 12B has clear HKR-H/K/R: 16GB local running, 12B scale, and Apache 2.0 licensing. It is a strong open-model update, not a must-write foundation-model launch.

Synced · WeChat

Office Whispering Is Turning Typing Into an Old Skill

AI dictation tools are moving into developer and office workflows, with Wispr Flow reporting over 2.5 million global downloads, 70% 12-month retention, and 100x annual user growth, while OpenAI’s gpt-4o-transcribe reached a 2.5% word error rate in a third-party evaluation cited by the article.

Why it matters: HKR-H/K/R all pass, but this is a data-backed workflow trend piece, not a model launch or platform update. It sits at the lower featured threshold.

AI HOT (Curated Pool)

Miso One Open-Sources Voice Model: 8B Parameters, 110ms Latency, One-Shot Voice Cloning

Miso One released an 8B-parameter open-weight TTS model with one-shot voice cloning from a short sample, 110ms inference latency, GitHub self-hosting without an API, and local audio data handling; the post says API access is coming but does not disclose pricing or launch timing.

Why it matters: HKR-H/K/R all pass, but this is a single X-sourced launch with no benchmark suite, license detail, or third-party reproduction. The 8B, 110ms, self-hosted open TTS facts clear featured, not higher.

Jun 3Wednesday

AI HOT (Curated Pool)

Suno raises $400 million Series D

Suno raised a $400 million Series D at a $5.4 billion valuation; the RSS snippet does not disclose the lead investor, participating investors, or planned use of proceeds.

Why it matters: HKR-H/K/R pass on Suno’s $400M Series D and $5.4B valuation, a clear AI-audio funding signal. Lead investor, participants, and use of funds are undisclosed, keeping it below the must-write band.

AI HOT (Curated Pool)

Grok Becomes Vapi's Default Voice Engine

xAI partnered with Vapi to make Grok the default engine for 12 core voices, covering more than 2.5 million voice agents, and Grok Voice ranked first in Vapi’s independent blind test.

Why it matters: HKR-H/K/R all pass: the default-engine switch has scale, numbers, and voice-agent market resonance. Single-source partnership news lacks test methodology, pricing, and migration data, so it stays in the mid product-update band.

Jun 2Tuesday

AI HOT (Curated Pool)

Gemini Omni Supports Creating Personal Digital Avatars

Gemini App says Gemini Omni can add users to video creation by generating a digital avatar that resembles their appearance and voice; the post does not disclose rollout scope, pricing, or safety mechanisms.

Why it matters: HKR-H/K/R all pass: the official Gemini App post has a strong multimodal avatar hook. Scope, pricing, consent, and safety controls are not disclosed, keeping it in the mid-weight product-update band.

Jun 1Monday

The Verge · AI

AI is blowing up music. How should the Grammys handle it?

Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.

Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 30Saturday

AI HOT (Curated Pool)

OpenAI launches real-time translation model with 70+ input languages

OpenAI launched gpt-realtime-translate, a speech translation model that accepts 70+ input languages and outputs speech in 13 target languages; the post says the feature is running on smart glasses.

Why it matters: HKR-H/K/R all pass: OpenAI has a concrete realtime-translation model with numbers and a wearable demo. Missing latency, pricing, and API availability keep it below P1.

May 29Friday

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.