Skip to content

Voice & audio

AI speech and audio: speech synthesis, real-time conversation, music generation and audio understanding.

116 picksRelated topicsMultimodalAI videoProduct updates

Latest picks

61–80 of 116

May 20Wednesday

AI HOT (Curated Pool)

Google launches Antigravity 2.0 platform, builds an OS in 12 hours

Google announced Antigravity 2.0 at I/O and demonstrated an agent building a runnable operating system from scratch in 12 hours, using 93 parallel sub-agents, more than 15,000 model calls, and 2.6 billion tokens, with API costs under $1,000.

Why it matters: HKR-H/K/R all pass: a Google I/O agent-platform release with concrete demo metrics. The post lacks availability, pricing, and replication details, so it lands in the lower 85–94 band.

AI HOT (Curated Pool)

Google releases Gemini Omni for any-input-to-any-output generation and conversational video editing

Google released Gemini Omni and Omni Flash at I/O 2026, supporting text, image, audio, and video inputs and outputs, with conversational video editing; Omni Flash is available in Gemini App, Google Flow, and YouTube Shorts, while the post does not disclose the API launch date.

Why it matters: HKR-H/K/R all pass: Google announced Gemini Omni and Omni Flash as a major multimodal update at I/O. API timing, pricing, and benchmarks are not disclosed, so it stays below 90.

r/LocalLLaMA

Floor for local meeting summarization on a 6GB GPU: Qwen3.5 0.8B works in 57s, Granite 4 350M hallucinates

The author tested VoiceFlow 1.6.0 on an RTX 3060 Laptop 6GB, where Qwen3.5 0.8B summarized a 4-minute meeting in 57 seconds with 16K context, while Granite 4 350M returned summaries in 0.6-2.8 seconds but fabricated Binance and Star Trek content.

Why it matters: HKR-H/K/R all pass: the hook is concrete, the test reports hardware/context/timing, and local meeting summarization hits privacy and cost nerves. Single Reddit experiment limits authority, so 73 featured.

May 19Tuesday

r/LocalLLaMA

21 GPUs benchmarked running a small TTS model, with 5GB peak VRAM

A Reddit user rented 21 GPUs on vast.ai to benchmark OmniVoice, a small TTS model with about 5GB peak VRAM, using xRT as the audio generation speed metric and averaging 3 voice-cloning runs with reference audio.

Why it matters: HKR-H/K/R pass: a 21-GPU TTS benchmark with 5GB peak VRAM and 3-run xRT averaging is useful to local-inference builders. Scope is niche, so it sits at the low featured band.

May 15Friday

AI HOT (Curated Pool)

Connect Grok to the Hermes Agent

xAI connects Grok subscription accounts to Nous Research’s open-source Hermes Agent across all subscription tiers, letting users run Grok 4.3 text chat and reasoning, generate spoken replies with text-to-speech, create images and videos with Grok Imagine, and connect the agent to WhatsApp or Discord.

Why it matters: HKR-H/K/R all pass, but this is a mid-weight xAI product integration with an open-source agent, not a flagship model release. Featured fits; it does not clear the 85+ same-day bar.

AI HOT (Curated Pool)

Accelerating On-Device AI: Arm and Google AI Edge Optimization Practices

Arm SME2 and Google AI Edge integrate with LiteRT, XNNPACK, and KleidiAI to optimize Stability AI’s stable-audio-open-small, delivering over 2x faster audio generation and 4x lower memory use on Arm-based mobile devices and laptops.

Why it matters: HKR-H/K/R pass via concrete 2x speed and 4x memory gains, plus an edge-deployment cost hook. Scope stays narrow to one audio model on Arm devices, so it lands at the featured threshold.

May 13Wednesday

AI HOT (Curated Pool)

Google Launches New Android Smart Assistant

Google introduced Android Intelligence at Android Show 2026, with multi-step automation across Android apps, browser-use features for Gemini in Chrome, automatic form filling, Rambler voice-note transcription, and custom Gen UI widgets; the post does not disclose rollout timing, supported devices, or pricing.

Why it matters: HKR-H/K/R all pass: the hook is Android-level agent control, the new facts are concrete automation surfaces, and the resonance is the mobile AI platform fight. Thin source detail keeps it at the low end of the 85-94 band.

TechCrunch · AI

Google adds Gemini-powered dictation to Gboard, which could be bad news for dictation startups

Google adds Gemini-powered dictation to Gboard for an initial launch on Samsung Galaxy and Google Pixel phones; the post does not disclose supported languages, pricing, offline behavior, or a rollout date.

Why it matters: HKR-H/K/R all pass: Gemini dictation lands inside Gboard with Galaxy and Pixel named, and the platform-bundling angle matters to AI startups. Missing language, pricing, offline mode, and timing keep it at the low featured band.

May 12Tuesday

TechCrunch · AI

AI Voice Startup Vapi Hits $500M Valuation After Winning Amazon Ring Over 40 Rivals

Vapi reached a $500 million valuation after Amazon Ring chose its AI voice platform over 40 rivals, and the RSS snippet says its enterprise business has grown tenfold since early 2025 as companies move support and sales calls to AI agents.

Why it matters: HKR-H/K/R all pass: Amazon Ring’s 40-rival selection and Vapi’s $500M valuation give concrete signal. Still a startup financing/customer win, not a model or platform release, so it sits at the featured threshold.

Xinzhiyuan · WeChat

OpenAI releases GPT-Realtime-2, described as a GPT-5-level reasoning audio model

OpenAI released GPT-Realtime-2 alongside Realtime-Translate and Realtime-Whisper, with a 128K context window, five reasoning-effort levels, and API pricing of $32 per million input tokens and $64 per million output tokens.

Why it matters: HKR-H/K/R all pass: realtime audio reasoning is a strong hook; 128K context, five reasoning levels, and $32/$64 per 1M tokens add substance; voice-agent cost and stack choices hit practitioners. This is a same-day OpenAI product update.

Latent Space

Thinking Machines' Native Interaction Models: TML-Interaction-Small 276B-A12B Advances Realtime Voice

Thinking Machines released TML-Interaction-Small, a 276B-parameter MoE model with 12B active parameters, and the post says it advances realtime voice through 200ms time-aligned microturns, encoder-free early fusion for audio and images under 200ms, and benchmark wins over GPT-Realtime-2 and Gemini 3.1-Flash.

Why it matters: HKR-H/K/R all pass: TML-Interaction-Small gives architecture, active parameters, 200ms interaction, and named rivals. Benchmarks still need replication, but a real-time voice SOTA claim is same-day material.

AI HOT (Curated Pool)

Thinking Machines Releases Native Multimodal Interaction Model for Real-Time Human-AI Collaboration

Thinking Machines released an interaction model that natively receives audio, video, and text input, processes foreground interaction at 200-millisecond intervals, and uses a background reasoning model for long-horizon planning and tool calls.

Why it matters: HKR-H/K/R all pass: this is more than a model notice, with a two-layer foreground/background interaction design. Pricing, access scope, and benchmarks are missing, so it sits at the lower end of 85-94.

The Verge · AI

Here’s What Mira Murati’s AI Company Is Up To

Thinking Machines announced work on “interaction models” that continuously take in audio, video, and text and respond or act in real time; the post does not disclose model size, release timing, pricing, or the final product format.

Why it matters: HKR-H/K/R all pass, but the body lacks parameters, launch timing, and product form. This is a high-interest startup direction reveal, not a usable model release, so it stays at the top of the 72–77 band.

May 8Friday

Synced · WeChat

OpenAI launches official CLI for terminal-based model access

OpenAI released the open-source openai-cli, letting developers call Responses, cloud tools, image generation and editing, speech transcription, and TTS from a single terminal command.

Why it matters: HKR-H/K/R all pass: an official OpenAI CLI, open-source packaging, and terminal access to multimodal APIs. This is a useful developer workflow update, not a major model capability release, so it sits in low featured.

Latent Space

[AINews] GPT-Realtime-2, Translate, and Whisper: new SOTA realtime voice APIs

OpenAI released GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API, with GPT-Realtime-2 expanding context from 32K to 128K and scoring 96.6% on Artificial Analysis Big Bench Audio.

Why it matters: HKR-H/K/R all pass: an OpenAI real-time voice API refresh, a 32K→128K context jump, and a 96.6% Big Bench Audio claim. Score stays at 86 because this is a major API update, not a flagship foundation-model release.

QbitAI · WeChat

OpenAI releases three realtime voice models for reasoning, translation, and transcription

OpenAI launched GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper as API models, covering 128K-context voice reasoning, streaming translation from more than 70 input languages into 13 output languages, and realtime transcription priced at $0.017 per minute.

Why it matters: OpenAI shipped three realtime voice APIs across reasoning, translation, and transcription, hitting HKR-H/K/R. The 128K context, 70+ languages, and $0.017/min price make this a same-day must-write item.

AI HOT (Curated Pool)

Apple's First AI Wearable: Camera-Equipped AirPods Enter DVT Stage

Apple’s camera-equipped AirPods have entered DVT, with launch possible in September. Each earbud uses a low-res camera for visual Q&A with the upgraded Siri. The post cites Google Gemini support and a data-upload indicator light.

Why it matters: HKR-H/K/R all pass, but this is an unconfirmed hardware rumor, not an Apple launch. DVT status, camera design, and Gemini dependency keep it in the low featured band.

AI HOT (Curated Pool)

OpenAI launches official openai-cli for terminal API calls

OpenAI open-sourced openai-cli for direct API calls from the terminal. The Apache 2.0 tool installs via Homebrew or Go and covers Responses API, structured output, image editing, transcription, and key config. The key detail is Agent workflows using cloud tools like web search and code interpreter.

Why it matters: HKR-H/K/R all pass: official OpenAI terminal tooling is clickable, with concrete install/license/API details and workflow resonance. It is still a developer tooling update, not a model or major capability release, so 76 fits the featured threshold.

The Verge · AI

Apple’s AirPods with cameras for AI are reportedly close to production

Mark Gurman says Apple’s camera-equipped AirPods are in DVT, one step before PVT. Testers are using prototypes; the cameras capture low-resolution visual input, not photos or video, for Siri queries like ingredient prompts.

Why it matters: HKR-H/K/R all pass: Gurman/The Verge provides a concrete DVT-stage Apple AI hardware update. It is still pre-production, not a launch, so it stays in the 72–77 band.

May 7Thursday

OpenAI News

Advancing Voice Intelligence with New Models in the API

OpenAI introduced new realtime voice models in its API for voice intelligence. The RSS snippet says they reason, translate, and transcribe speech; the post does not disclose counts, pricing, or limits.

Why it matters: OpenAI’s official voice API update hits HKR-H/K/R, but the available body gives capability direction only. Model count, pricing, latency, and context limits are not disclosed, so it stays at the top of 78–84.