Skip to content

Voice & audio

AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

116 picksRelated topicsMultimodalAI videoProduct updates

Latest picks

41–60 of 116

May 27Wednesday

TechCrunch · AI

ElevenLabs’ new music-generation model can switch genres mid-track

ElevenLabs introduced a music-generation model that can regenerate one section of a song without changing the rest of the track; the post does not disclose launch timing, pricing, or model parameters.

Why it matters: HKR-H/K pass because ElevenLabs adds segment-level regeneration and mid-track genre switching. Price, launch timing, and model specs are not disclosed, so the industry impact stays mid-tier.

AI HOT (Curated Pool)

Reachy Mini enables fully local voice interaction

Reachy Mini implements local voice interaction through the speech-to-speech library, using a cascaded pipeline with a Realtime API-compatible WebSocket interface and default components including Silero VAD, Parakeet-TDT, and Qwen3-TTS.

Why it matters: HKR-H/K/R all pass: the post has a clear local-robot voice hook, concrete stack details, and edge-agent resonance. Scope stays limited to Reachy Mini voice interaction, so it sits at the featured threshold.

AI HOT (Curated Pool)

MiMo 2.5 Pro Gets Major Price Cut, Matching DeepSeek V4 Pro

Xiaomi permanently cut MiMo-V2.5 API prices by up to 99%, matched DeepSeek V4 Pro pricing, increased same-price token allowances by 5–8x, reset existing user quotas in full, and set the new pricing to take effect on May 26.

Why it matters: HKR-H/K/R all pass: the 99% cut creates a price-war hook, the post gives 5-8x token economics, and API cost pressure resonates. It remains a pricing update, not a model or capability release, so it stays below the 78+ band.

May 26Tuesday

Financial Times · Technology

Spotify chief defends AI-generated music

Spotify struck a deal with Universal that allows subscribers to create “controlled” covers and remixes; the post does not disclose the licensing scope, revenue split, or launch timing.

Why it matters: FT sourcing and a Spotify-Universal licensing frame clear HKR-H/K/R. Scope, revenue split, and launch timing are not disclosed, so this stays at the featured threshold for a mid-weight product/partnership update.

May 24Sunday

AI HOT (Curated Pool)

StepAudio 2.5 Realtime Voice Released with Paralinguistic Awareness and Persona Interaction

StepFun released StepAudio 2.5 Realtime with Chinese and English real-time voice support, API-based custom personas, more than 10,000 native persona options, millions of composable traits, and 5 built-in preset personas.

Why it matters: HKR-H/K/R all pass, but the source is an official X post and lacks latency, pricing, benchmarks, and rollout scope. This fits the low featured band for a mid-weight product update.

May 23Saturday

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

TechCrunch · AI

AI is being used to resurrect the voices of dead pilots

People used AI to reconstruct voices from spectrogram images of cockpit recordings, and the NTSB temporarily blocked access to its docket system; the post does not disclose the model, case count, or duration of the access block.

Why it matters: HKR-H/K/R all pass: the headline has a sharp ethics hook, and the summary gives a mechanism plus NTSB action. Scope is narrower than a major product or policy event, so it fits the 72–77 featured band.

Hacker News front page

NTSB pulls docket after AI recreates dead pilots' voices

NTSB pulled an accident docket after AI users recreated dead pilots’ voices; the post only includes RSS and Hacker News metadata with 20 points and 17 comments, and does not disclose the docket number, audio source, or removal conditions.

Why it matters: HKR-H and HKR-R are strong: cloned voices of dead pilots forced an NTSB docket pull. HKR-K is real but thin because docket ID, audio source, and removal terms are not disclosed.

May 22Friday

AI HOT (Curated Pool)

NetEase Youdao Open-Sources Ziyue 4 Multimodal and Text-to-Speech Models

NetEase Youdao open-sourced its Ziyue 4.0 multimodal and text-to-speech models, with the 27B multimodal model reporting 81.4% accuracy on Chinese math reasoning tasks and the speech model supporting 14 languages.

Why it matters: HKR-H/K/R pass: the story has a concrete open-source hook, specific model numbers, and practitioner relevance. NetEase Youdao is not a frontier lab, so it stays below the 78+ good-quality band.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

TechCrunch · AI

Spotify and Universal Music Strike Deal Allowing Fan-Made AI Covers and Remixes

Spotify is partnering with Universal Music Group to let Premium subscribers create AI-generated covers and remixes; the RSS snippet says participating artists receive a revenue share, but the post does not disclose the percentage or launch terms.

Why it matters: HKR-H/K/R all pass: Spotify and UMG create an authorized path for AI covers/remixes, with Premium access and artist revenue share disclosed. Split details are missing, so it stays at featured threshold, not p1.

May 21Thursday

The Verge · AI

Spotify is launching AI-generated remixes

Spotify and UMG announced a licensing deal that lets Premium subscribers pay for AI-generated remixes and covers of streaming songs; artists can opt out, while participating artists collect royalties from these AI remixes.

Why it matters: HKR-H/K/R all pass: Spotify and UMG add licensed AI covers/remixes with paid use, opt-out, and royalties. It is consumer audio, not a model or dev-tool release, so it sits just above the featured threshold.

Financial Times · Technology

Spotify targets high-spending superfans with AI-generated music

Spotify and Universal Music Group struck a licensing deal for a paid AI-generated music add-on inside Spotify’s app, targeting high-spending superfans; the RSS snippet does not disclose pricing, launch timing, supported markets, or model details.

Why it matters: HKR-H/K/R all pass: Spotify-UMG licensing turns AI music into a paid in-app product, not just a demo. Pricing, launch date, and revenue split are not disclosed, so this stays below must-write range.

May 20Wednesday

AI HOT (Curated Pool)

Stability AI Launches Stability Audio 3.0 for Songs Up to 6 Minutes

Stability AI launched the Stability Audio 3.0 audio generation model family with four sizes ranging from 459 million to 2.7 billion parameters; the small model targets on-device use and generates audio under 2 minutes locally, while medium and large models support full music creation beyond 6 minutes and 20 seconds.

Why it matters: HKR-H/K pass because Stability AI gives concrete duration and model-size details. HKR-R is weak: no benchmarks, licensing, pricing, or access terms are disclosed, so this sits at the featured threshold.

TechCrunch · AI

Stability AI releases a new audio model that can create 6-minute songs

Stability AI released Stability Audio 3.0 small; the title says it can create six-minute songs, while the RSS snippet only discloses that the small model can run on-device and generate two-minute tracks.

Why it matters: Mid-weight product update: HKR-H comes from the 6-minute-song hook, HKR-K from on-device use and 2-minute tracks. The title/body duration mismatch keeps it at the low featured threshold.

AI HOT (Curated Pool)

Gemini Omni Supports Video Creation With Personal Likeness and Voice

Gemini Omni lets users create digital-avatar videos using their personal likeness and voice, and the avatar can generate videos without uploading an image each time; the post does not disclose pricing, regions, or launch timing.

Why it matters: HKR-H/K/R pass: personal avatar video is clicky, reusable identity is a concrete mechanism, and voice/likeness raises creator and safety stakes. Price, regions, and launch timing are not disclosed, keeping it near the featured floor.

TechCrunch · AI

You Can Now Talk to Your Gmail Inbox, as Seen at Google I/O 2026

Google expanded Gmail’s AI Inbox with conversational voice search, letting users ask Gemini to find details buried in email. The RSS snippet does not disclose rollout scope, supported languages, pricing, latency, or the retrieval mechanism behind Gmail search.

Why it matters: HKR-H/K pass: a Google-scale Gmail voice inbox feature is concrete and clickable. HKR-R is weak because rollout, language support, pricing, and retrieval mechanics are not disclosed.

TechCrunch · AI

Google takes a page from Meta, announces audio-powered smart glasses at I/O 2026

Google announced “audio glasses” at I/O 2026, letting users issue voice commands across its apps and services, including Gemini; the RSS snippet does not disclose price, launch timing, or hardware specifications.

Why it matters: HKR-H/K/R pass: Google announced Gemini-linked audio glasses at I/O 2026, a credible AI-hardware platform move. Missing price, launch date, and specs keep it in the low featured band.

The Verge · AI

Gmail is going to start talking to you

Google is launching Gmail Live for Gmail, letting users tap a search-bar icon and ask voice questions about inbox content; a press demo retrieved school event dates, locations, and an upcoming Detroit trip from the employee’s email.

Why it matters: HKR-H/K/R pass: Gmail Live adds voice email queries inside a mass-market Google surface. The post gives demo cases, but no launch date, pricing, or model details, so it stays at the lower featured band.

TechCrunch · AI

Google's Gemini Omni turns images, audio, and text into video

Google's Gemini Omni generates and edits video through conversation, using text, images, audio, and video as inputs, with Omni Flash named as the starting version; the RSS snippet says the model reasons across modalities, but the post does not disclose launch date, pricing, context limits, benchmarks, or API availability.

Why it matters: Google-scale Gemini multimodal video update clears HKR-H/K/R: Omni Flash, chat-based editing, and four input types are concrete. Pricing and rollout are not disclosed, so it sits in the lower must-write band.