Skip to content

#语音

0 today

Jun 1Monday

The Verge · AI

AI is blowing up music. How should the Grammys handle it?

Deezer reports that more than 50,000 AI-generated songs are uploaded each day, while Recording Academy CEO Harvey Mason Jr. says AI is now present in every recent music session he has attended and Grammy rules still bar AI music from the industry’s highest honors.

Why it matters: HKR-H/K/R all pass, but this is a podcast-style policy discussion rather than a model, product, or binding regulation story. The concrete signal is the 50,000/day Deezer figure plus the Grammy eligibility conflict.

r/LocalLLaMA

I ported NVIDIA Parakeet speech-to-text to ggml: same output as NeMo, faster, GGUF-quantized, no Python

mudler_it ported NVIDIA Parakeet speech-to-text models to C++/ggml with no Python or PyTorch, reporting byte-for-byte NeMo parity on f32/f16, up to about 5x GPU speedups on larger TDT and hybrid models, and GGUF quantization across f16, q8_0, q6_k, q5_k, and q4_k.

Why it matters: HKR-H/K/R all pass: the port has a concrete local-inference hook, byte-parity and speed claims, and clear practitioner resonance. Source scope keeps it at the low featured band, not P1.

May 30Saturday

AI HOT (Curated Pool)

OpenAI launches real-time translation model with 70+ input languages

OpenAI launched gpt-realtime-translate, a speech translation model that accepts 70+ input languages and outputs speech in 13 target languages; the post says the feature is running on smart glasses.

Why it matters: HKR-H/K/R all pass: OpenAI has a concrete realtime-translation model with numbers and a wearable demo. Missing latency, pricing, and API availability keep it below P1.

May 29Friday

AI HOT (Curated Pool)

Xiaomi Open-Sources Controllable Video Foley Model ControlFoley

Xiaomi’s large model application team open-sourced ControlFoley, a controllable video Foley model supporting three tasks: text-guided video dubbing, text-controlled video dubbing, and reference-audio-controlled video dubbing, with code, model weights, and an online demo released.

Why it matters: ControlFoley clears HKR-H/K/R with controllable video Foley plus code, weights, and demo. It is a useful multimodal-audio release from Xiaomi, but not a flagship foundation-model launch, so it sits near the featured threshold.

May 27Wednesday

TechCrunch · AI

ElevenLabs’ new music-generation model can switch genres mid-track

ElevenLabs introduced a music-generation model that can regenerate one section of a song without changing the rest of the track; the post does not disclose launch timing, pricing, or model parameters.

Why it matters: HKR-H/K pass because ElevenLabs adds segment-level regeneration and mid-track genre switching. Price, launch timing, and model specs are not disclosed, so the industry impact stays mid-tier.

AI HOT (Curated Pool)

Reachy Mini enables fully local voice interaction

Reachy Mini implements local voice interaction through the speech-to-speech library, using a cascaded pipeline with a Realtime API-compatible WebSocket interface and default components including Silero VAD, Parakeet-TDT, and Qwen3-TTS.

Why it matters: HKR-H/K/R all pass: the post has a clear local-robot voice hook, concrete stack details, and edge-agent resonance. Scope stays limited to Reachy Mini voice interaction, so it sits at the featured threshold.

AI HOT (Curated Pool)

MiMo 2.5 Pro Gets Major Price Cut, Matching DeepSeek V4 Pro

Xiaomi permanently cut MiMo-V2.5 API prices by up to 99%, matched DeepSeek V4 Pro pricing, increased same-price token allowances by 5–8x, reset existing user quotas in full, and set the new pricing to take effect on May 26.

Why it matters: HKR-H/K/R all pass: the 99% cut creates a price-war hook, the post gives 5-8x token economics, and API cost pressure resonates. It remains a pricing update, not a model or capability release, so it stays below the 78+ band.

May 26Tuesday

Financial Times · Technology

Spotify chief defends AI-generated music

Spotify struck a deal with Universal that allows subscribers to create “controlled” covers and remixes; the post does not disclose the licensing scope, revenue split, or launch timing.

Why it matters: FT sourcing and a Spotify-Universal licensing frame clear HKR-H/K/R. Scope, revenue split, and launch timing are not disclosed, so this stays at the featured threshold for a mid-weight product/partnership update.

May 24Sunday

AI HOT (Curated Pool)

StepAudio 2.5 Realtime Voice Released with Paralinguistic Awareness and Persona Interaction

StepFun released StepAudio 2.5 Realtime with Chinese and English real-time voice support, API-based custom personas, more than 10,000 native persona options, millions of composable traits, and 5 built-in preset personas.

Why it matters: HKR-H/K/R all pass, but the source is an official X post and lacks latency, pricing, benchmarks, and rollout scope. This fits the low featured band for a mid-weight product update.

May 23Saturday

r/LocalLLaMA

meituan-longcat/LongCat-Video-Avatar-1.5 on Hugging Face

Meituan LongCat released LongCat-Video-Avatar-1.5 on Hugging Face, supporting AT2V, ATI2V, and video continuation while replacing Wav2Vec2 with Whisper-Large and using DMD2 distillation to reduce inference to 8 NFE; the model weights are released under the MIT License.

Why it matters: HKR-H/K/R all pass: open MIT video-avatar weights plus 8 NFE inference give local multimodal builders real signal. This is a mid-weight open-source model update, not an 85+ same-day industry event.

TechCrunch · AI

AI is being used to resurrect the voices of dead pilots

People used AI to reconstruct voices from spectrogram images of cockpit recordings, and the NTSB temporarily blocked access to its docket system; the post does not disclose the model, case count, or duration of the access block.

Why it matters: HKR-H/K/R all pass: the headline has a sharp ethics hook, and the summary gives a mechanism plus NTSB action. Scope is narrower than a major product or policy event, so it fits the 72–77 featured band.

Hacker News front page

NTSB pulls docket after AI recreates dead pilots' voices

NTSB pulled an accident docket after AI users recreated dead pilots’ voices; the post only includes RSS and Hacker News metadata with 20 points and 17 comments, and does not disclose the docket number, audio source, or removal conditions.

Why it matters: HKR-H and HKR-R are strong: cloned voices of dead pilots forced an NTSB docket pull. HKR-K is real but thin because docket ID, audio source, and removal terms are not disclosed.

May 22Friday

AI HOT (Curated Pool)

NetEase Youdao Open-Sources Ziyue 4 Multimodal and Text-to-Speech Models

NetEase Youdao open-sourced its Ziyue 4.0 multimodal and text-to-speech models, with the 27B multimodal model reporting 81.4% accuracy on Chinese math reasoning tasks and the speech model supporting 14 languages.

Why it matters: HKR-H/K/R pass: the story has a concrete open-source hook, specific model numbers, and practitioner relevance. NetEase Youdao is not a frontier lab, so it stays below the 78+ good-quality band.

AI HOT (Curated Pool)

Zhipu releases GLM-5.1-highspeed, claiming a large-model API speed record

Zhipu released the GLM-5.1-highspeed API to selected enterprise customers on May 22, with a claimed output speed of 400 tokens/s, built by the GLM team and TileRT team through system-level optimization.

Why it matters: HKR-H/K/R all pass: Zhipu’s GLM-5.1 high-speed API has a concrete 400 tokens/s claim and domestic flagship-model relevance. Test setup, pricing, and availability are not disclosed, so it stays in the 78–84 band.

TechCrunch · AI

Spotify and Universal Music Strike Deal Allowing Fan-Made AI Covers and Remixes

Spotify is partnering with Universal Music Group to let Premium subscribers create AI-generated covers and remixes; the RSS snippet says participating artists receive a revenue share, but the post does not disclose the percentage or launch terms.

Why it matters: HKR-H/K/R all pass: Spotify and UMG create an authorized path for AI covers/remixes, with Premium access and artist revenue share disclosed. Split details are missing, so it stays at featured threshold, not p1.

May 21Thursday

The Verge · AI

Spotify is launching AI-generated remixes

Spotify and UMG announced a licensing deal that lets Premium subscribers pay for AI-generated remixes and covers of streaming songs; artists can opt out, while participating artists collect royalties from these AI remixes.

Why it matters: HKR-H/K/R all pass: Spotify and UMG add licensed AI covers/remixes with paid use, opt-out, and royalties. It is consumer audio, not a model or dev-tool release, so it sits just above the featured threshold.

Financial Times · Technology

Spotify targets high-spending superfans with AI-generated music

Spotify and Universal Music Group struck a licensing deal for a paid AI-generated music add-on inside Spotify’s app, targeting high-spending superfans; the RSS snippet does not disclose pricing, launch timing, supported markets, or model details.

Why it matters: HKR-H/K/R all pass: Spotify-UMG licensing turns AI music into a paid in-app product, not just a demo. Pricing, launch date, and revenue split are not disclosed, so this stays below must-write range.

May 20Wednesday

AI HOT (Curated Pool)

Stability AI Launches Stability Audio 3.0 for Songs Up to 6 Minutes

Stability AI launched the Stability Audio 3.0 audio generation model family with four sizes ranging from 459 million to 2.7 billion parameters; the small model targets on-device use and generates audio under 2 minutes locally, while medium and large models support full music creation beyond 6 minutes and 20 seconds.

Why it matters: HKR-H/K pass because Stability AI gives concrete duration and model-size details. HKR-R is weak: no benchmarks, licensing, pricing, or access terms are disclosed, so this sits at the featured threshold.

TechCrunch · AI

Stability AI releases a new audio model that can create 6-minute songs

Stability AI released Stability Audio 3.0 small; the title says it can create six-minute songs, while the RSS snippet only discloses that the small model can run on-device and generate two-minute tracks.

Why it matters: Mid-weight product update: HKR-H comes from the 6-minute-song hook, HKR-K from on-device use and 2-minute tracks. The title/body duration mismatch keeps it at the low featured threshold.

AI HOT (Curated Pool)

Gemini Omni Supports Video Creation With Personal Likeness and Voice

Gemini Omni lets users create digital-avatar videos using their personal likeness and voice, and the avatar can generate videos without uploading an image each time; the post does not disclose pricing, regions, or launch timing.

Why it matters: HKR-H/K/R pass: personal avatar video is clicky, reusable identity is a concrete mechanism, and voice/likeness raises creator and safety stakes. Price, regions, and launch timing are not disclosed, keeping it near the featured floor.

TechCrunch · AI

You Can Now Talk to Your Gmail Inbox, as Seen at Google I/O 2026

Google expanded Gmail’s AI Inbox with conversational voice search, letting users ask Gemini to find details buried in email. The RSS snippet does not disclose rollout scope, supported languages, pricing, latency, or the retrieval mechanism behind Gmail search.

Why it matters: HKR-H/K pass: a Google-scale Gmail voice inbox feature is concrete and clickable. HKR-R is weak because rollout, language support, pricing, and retrieval mechanics are not disclosed.

TechCrunch · AI

Google takes a page from Meta, announces audio-powered smart glasses at I/O 2026

Google announced “audio glasses” at I/O 2026, letting users issue voice commands across its apps and services, including Gemini; the RSS snippet does not disclose price, launch timing, or hardware specifications.

Why it matters: HKR-H/K/R pass: Google announced Gemini-linked audio glasses at I/O 2026, a credible AI-hardware platform move. Missing price, launch date, and specs keep it in the low featured band.

The Verge · AI

Gmail is going to start talking to you

Google is launching Gmail Live for Gmail, letting users tap a search-bar icon and ask voice questions about inbox content; a press demo retrieved school event dates, locations, and an upcoming Detroit trip from the employee’s email.

Why it matters: HKR-H/K/R pass: Gmail Live adds voice email queries inside a mass-market Google surface. The post gives demo cases, but no launch date, pricing, or model details, so it stays at the lower featured band.

TechCrunch · AI

Google's Gemini Omni turns images, audio, and text into video

Google's Gemini Omni generates and edits video through conversation, using text, images, audio, and video as inputs, with Omni Flash named as the starting version; the RSS snippet says the model reasons across modalities, but the post does not disclose launch date, pricing, context limits, benchmarks, or API availability.

Why it matters: Google-scale Gemini multimodal video update clears HKR-H/K/R: Omni Flash, chat-based editing, and four input types are concrete. Pricing and rollout are not disclosed, so it sits in the lower must-write band.

AI HOT (Curated Pool)

Google launches Antigravity 2.0 platform, builds an OS in 12 hours

Google announced Antigravity 2.0 at I/O and demonstrated an agent building a runnable operating system from scratch in 12 hours, using 93 parallel sub-agents, more than 15,000 model calls, and 2.6 billion tokens, with API costs under $1,000.

Why it matters: HKR-H/K/R all pass: a Google I/O agent-platform release with concrete demo metrics. The post lacks availability, pricing, and replication details, so it lands in the lower 85–94 band.

AI HOT (Curated Pool)

Google releases Gemini Omni for any-input-to-any-output generation and conversational video editing

Google released Gemini Omni and Omni Flash at I/O 2026, supporting text, image, audio, and video inputs and outputs, with conversational video editing; Omni Flash is available in Gemini App, Google Flow, and YouTube Shorts, while the post does not disclose the API launch date.

Why it matters: HKR-H/K/R all pass: Google announced Gemini Omni and Omni Flash as a major multimodal update at I/O. API timing, pricing, and benchmarks are not disclosed, so it stays below 90.

r/LocalLLaMA

Floor for local meeting summarization on a 6GB GPU: Qwen3.5 0.8B works in 57s, Granite 4 350M hallucinates

The author tested VoiceFlow 1.6.0 on an RTX 3060 Laptop 6GB, where Qwen3.5 0.8B summarized a 4-minute meeting in 57 seconds with 16K context, while Granite 4 350M returned summaries in 0.6-2.8 seconds but fabricated Binance and Star Trek content.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the test reports hardware/context/timing, and local meeting summarization hits privacy and cost nerves. Single Reddit experiment limits authority, so 73 featured.

May 19Tuesday

r/LocalLLaMA

21 GPUs benchmarked running a small TTS model, with 5GB peak VRAM

A Reddit user rented 21 GPUs on vast.ai to benchmark OmniVoice, a small TTS model with about 5GB peak VRAM, using xRT as the audio generation speed metric and averaging 3 voice-cloning runs with reference audio.

Why it matters: HKR-H/K/R pass: a 21-GPU TTS benchmark with 5GB peak VRAM and 3-run xRT averaging is useful to local-inference builders. Scope is niche, so it sits at the low featured band.

May 15Friday

AI HOT (Curated Pool)

Connect Grok to the Hermes Agent

xAI connects Grok subscription accounts to Nous Research’s open-source Hermes Agent across all subscription tiers, letting users run Grok 4.3 text chat and reasoning, generate spoken replies with text-to-speech, create images and videos with Grok Imagine, and connect the agent to WhatsApp or Discord.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight xAI product integration with an open-source agent, not a flagship model release. Featured fits; it does not clear the 85+ same-day bar.

AI HOT (Curated Pool)

Accelerating On-Device AI: Arm and Google AI Edge Optimization Practices

Arm SME2 and Google AI Edge integrate with LiteRT, XNNPACK, and KleidiAI to optimize Stability AI’s stable-audio-open-small, delivering over 2x faster audio generation and 4x lower memory use on Arm-based mobile devices and laptops.

Why it matters: HKR-H/K/R pass via concrete 2x speed and 4x memory gains, plus an edge-deployment cost hook. Scope stays narrow to one audio model on Arm devices, so it lands at the featured threshold.

May 13Wednesday

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

TechCrunch · AI

Google adds Gemini-powered dictation to Gboard, which could be bad news for dictation startups

Google adds Gemini-powered dictation to Gboard for an initial launch on Samsung Galaxy and Google Pixel phones; the post does not disclose supported languages, pricing, offline behavior, or a rollout date.

Why it matters: HKR-H/K/R all pass: Gemini dictation lands inside Gboard with Galaxy and Pixel named, and the platform-bundling angle matters to AI startups. Missing language, pricing, offline mode, and timing keep it at the low featured band.

May 12Tuesday

TechCrunch · AI

AI Voice Startup Vapi Hits $500M Valuation After Winning Amazon Ring Over 40 Rivals

Vapi reached a $500 million valuation after Amazon Ring chose its AI voice platform over 40 rivals, and the RSS snippet says its enterprise business has grown tenfold since early 2025 as companies move support and sales calls to AI agents.

Why it matters: HKR-H/K/R all pass: Amazon Ring’s 40-rival selection and Vapi’s $500M valuation give concrete signal. Still a startup financing/customer win, not a model or platform release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

AI HOT (Curated Pool)

Thinking Machines Releases Native Multimodal Interaction Model for Real-Time Human-AI Collaboration

Thinking Machines released an interaction model that natively receives audio, video, and text input, processes foreground interaction at 200-millisecond intervals, and uses a background reasoning model for long-horizon planning and tool calls.

Why it matters: HKR-H/K/R all pass: this is more than a model notice, with a two-layer foreground/background interaction design. Pricing, access scope, and benchmarks are missing, so it sits at the lower end of 85-94.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

May 8Friday

Synced · WeChat

OpenAI launches official CLI for terminal-based model access

OpenAI released the open-source openai-cli, letting developers call Responses, cloud tools, image generation and editing, speech transcription, and TTS from a single terminal command.

Why it matters: HKR-H/K/R all pass: an official OpenAI CLI, open-source packaging, and terminal access to multimodal APIs. This is a useful developer workflow update, not a major model capability release, so it sits in low featured.

Latent Space

[AINews] GPT-Realtime-2, Translate, and Whisper: new SOTA realtime voice APIs

OpenAI released GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API, with GPT-Realtime-2 expanding context from 32K to 128K and scoring 96.6% on Artificial Analysis Big Bench Audio.

Why it matters: HKR-H/K/R all pass: an OpenAI real-time voice API refresh, a 32K→128K context jump, and a 96.6% Big Bench Audio claim. Score stays at 86 because this is a major API update, not a flagship foundation-model release.

QbitAI · WeChat

OpenAI releases three realtime voice models for reasoning, translation, and transcription

OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper as API models, covering 128K-context voice reasoning, streaming translation from more than 70 input languages into 13 output languages, and realtime transcription priced at $0.017 per minute.

Why it matters: OpenAI shipped three realtime voice APIs across reasoning, translation, and transcription, hitting HKR-H/K/R. The 128K context, 70+ languages, and $0.017/min price make this a same-day must-write item.