ElevenLabs' new v4 speech model makes AI voices more expressive and consistent
ElevenLabs 发布 Eleven v4 语音模型,能更准确跟随脚本中的情绪、停顿与音效标签,并在长篇制作中保持音色一致。新架构同时驱动 Turbo 版本,官方测试中约 150 毫秒开始输出语音,对比 Cartesia Sonic 3.6 的 262 毫秒和 OpenAI GPT-4o mini TTS 的 814 毫秒。
ElevenLabs 发布 Eleven v4 语音模型,能更准确跟随脚本中的情绪、停顿与音效标签,并在长篇制作中保持音色一致。新架构同时驱动 Turbo 版本,官方测试中约 150 毫秒开始输出语音,对比 Cartesia Sonic 3.6 的 262 毫秒和 OpenAI GPT-4o mini TTS 的 814 毫秒。
Google DeepMind released Gemini 3.8 Live with Live Avatar, adding near-real-time video generation to its native real-time conversation model. The result is a dynamic visual avatar with lip sync, natural expressions and smooth turn-taking.
Why it matters: The post details Live Avatar's real-time video conversation, async tool calls and 97-language support, a useful read on enterprise multimodal interaction.
Google is testing 'Call for Me,' letting Gemini make calls to businesses. It's limited to US Pixel 11 owners with a Gemini subscription, using the beta Google Phone app. Gemini can now share user-approved personal info, expanding what it can do. You can follow the call live and take over anytime. The post doesn't disclose a launch date, pricing changes, or the business-side experience.
Why it matters: Google is testing a feature that lets Gemini call businesses and share user-approved personal info to handle bookings or order lookups. The high barrier (US, Pixel 11, paid sub, beta app) keeps it a tech preview for now, so the score stays moderate. But the direction—AI making...
Meta positioned Muse as the core of a hardware-plus-agent play at Connect. Muse now does voice and real-time video, handles long background conversations, and gets its own email address you can CC. Mac computer use lets you queue jobs and walk away. It's free for now but may take a transaction cut later; retail partners include Walmart, Best Buy, and Sephora, with productivity connectors for Box, GitHub, and Notion. Hardware updates: Ray-Ban Meta Gen 3 with better battery and mics, plus Charm, a standalone handheld gadget. No new frontier model shipped—only an MSL tease.
Why it matters: Muse updates at Meta Connect are substantive: voice, real-time video, background tasks, email address, Mac desktop control, plus named retail and productivity partners. Not a vague launch — verifiable integration list. Score held back because this is a paid Latent Space newsle...
Meta teased the Muse Charm at the end of Connect, a dedicated hardware device for its Muse AI agent. It resembles a chunky strapless smartwatch with a lanyard. A fingerprint sensor on the top right activates voice input; the front has at least three mic holes and a small camera. Zuckerberg noted you don't need to unlock a phone to use it. The post doesn't disclose pricing, battery life, or a release date.
Why it matters: Meta teased a standalone Muse AI gadget at the end of Connect — a thick watch-face on a lanyard with fingerprint wake, voice, and a camera. Only looks and interaction logic are disclosed; no price, battery, or launch date, so substance is thin and the score sits right at the f...
Two weeks after launching Muse, Meta says it's working on bringing the agent to its smart glasses, including the new ones shown at Connect. You'll activate it by saying its name and can ask it to guide workouts, log meals, or help shop for products you're looking at. The glasses are also getting an FDA-cleared hearing enhancement feature for adults with mild to moderate hearing loss. The post doesn't specify a launch date or which models will get Muse.
Why it matters: Putting Muse on glasses is a key step in Meta's push to move AI assistants from phones to wearables, with three concrete use cases. But the post doesn't give a launch date or supported models, so the score sits right at the featured threshold.
Greg Brockman announced that GPT Voice can now use tools like email, calendar, and Slack, powered by GPT-6 Astra, Sol, and Luna. Voice also lands on ChatGPT Work across web and mobile, letting users create docs, presentations, websites, or sheets hands-free in the browser. Rolling out globally today in the latest app version. The post doesn't disclose latency, accuracy, or enterprise access details, so I'd discount the demo until we see real-world numbers.
Why it matters: Major update to a core OpenAI product line: GPT Voice moves from conversation toy to tool-calling work entry point, explicitly tied to the GPT-6 model family. Announced by Brockman himself with global rollout — signal strength clears featured. Held below 90 because latency and...
OpenAI added plugin access to ChatGPT Voice so it can work across email, calendar, and Slack. Voice is live on ChatGPT Work web and mobile, letting users create docs, decks, sites, and spreadsheets by speaking. OpenAI says it's powered by GPT-6 Astra, Sol, and Luna, rolling out globally in the latest app version the same day. The post doesn't cover plugin permission scopes, latency, or pricing.
Why it matters: Voice mode with office plugins is a substantive product upgrade, and naming three GPT-6 models adds density. Score held below 85 because the post doesn't disclose permission scopes or latency — two gaps that make real-world utility uncertain.
Simon Willison 用 GPT-6 Astra 开发了一个 Gemini 3.8 TTS Playground,可测试 Google 的 Gemini 3.8 文本转语音 API,支持单人或多人对话合成、试听音频并查看请求与响应细节,配置可保存为可分享的 URL。
ChatGPT Voice now works with Mail, Calendar, and Slack plugins, running on GPT-6 Astra, Sol, and Luna. ChatGPT Work on web and mobile also gets voice input—you can create docs, presentations, websites, or sheets by speaking. The post doesn't disclose rollout regions, latency, or plugin permission details.
Why it matters: Adding email, calendar, and Slack to ChatGPT Voice is a real product expansion — it moves from conversation into office automation. The three GPT-6 sub-models (Astra, Sol, Luna) are named for the first time, but with zero capability breakdown, the signal is thinner than it sho...
Google DeepMind announced Gemini 3.8 Live, combining real-time voice with Extended Thinking. The model can reason while speaking, pausing briefly for harder questions before responding. The post body only contains the title and site navigation—no parameters, latency figures, or launch dates are disclosed.
StepFun launched the StepAudio 3 series of voice models, with several topping the Artificial Analysis global rankings. The post is blocked by WeChat and does not disclose specific parameters, ranking details, or model capabilities. The title confirms the release and ranking results but does not specify which metrics or languages.
Kinney Drugs rolled back its AI phone assistant Burt after three months, following customer reports of incoherent calls, wrong dosages, and missed prescription alerts. President John Marraffa said HIPAA compliance doesn't equal a good experience. Incoming patient calls return to a touch-tone system; Burt stays only for opt-in refill texts. The article does not name the underlying model or voice vendor.
Why it matters: A pharmacy chain pulled its AI phone assistant after dosage errors and missed prescription reminders, with the CEO publicly owning the failure — a rare honest postmortem of AI in a healthcare setting. Score held back because the article doesn't name the underlying model or voi...
Anthropic released Claude Opus 5 at half the cost of Fable 5, claiming near-parity. Every's review found it argues, stops early, and fights old prompting habits. Anthropic cut over 80% of Claude Code's system prompt for Opus 5 and Fable 5 with no measurable coding-eval loss. Theo spent hours rewriting CLAUDE.md and skills files and called it worth it. ChatGPT Voice now controls the desktop app inside Work and Codex, spawning new sessions for tasks and reporting back—like a voice-driven OpenClaw. Keshav found it weaker for serious work than manually using 5.6 Sol in Codex, but decent for email, dashboards, and charts. Claude's voice mode quietly added Sonnet and Opus support plus mid-conversation tool calls to Gmail, Calendar, and Slack. Kimi K3 weights and tech report are public, with a 50% discount on Droid until Aug 10. Jensen Huang posted on X for the first time amid rumors of a US ban on Chinese open-weight models.
Why it matters: Anthropic model launch with halved pricing is a substantive update. Every and Theo's hands-on tests provide concrete signal: strong capability but awkward behavior requiring prompt rewrites. Cross-source discussion is forming, but the body is summary-only—missing full review d...
ChatGPT's desktop app now accepts voice commands that can control agents and perform multi-step tasks. It uses the ChatGPT-Live voice models launched earlier this month, works with ChatGPT Work and Codex, and can browse websites and apps. On macOS, Appshots lets it read screen content. A demo showed a developer asking ChatGPT to create a thread, make a pull request, and find a bug's root cause in one go. The smartphone version only handled conversation; the desktop update adds real execution. Anthropic also updated Claude's voice mode yesterday to operate Gmail, Slack, and other apps.
Why it matters: OpenAI brought ChatGPT-Live voice to desktop with screen reading and browser control, turning voice into a real agent driver. The dev demo is concrete and useful. Not scoring higher because it just launched — real-world stability and permission boundaries are still unknown.
Anthropic upgraded Claude's voice mode to support Opus and Sonnet for complex reasoning, plus direct access to connected tools like Gmail and Slack during voice conversations. Multilingual support is also added, though the post doesn't list which languages. I'd wait for real-world latency and tool-calling reliability data before getting too excited.
Why it matters: Anthropic shipped a substantive voice-mode upgrade, fixing both the model-capability and tool-use gaps in one go. No latency/accuracy numbers or language list disclosed, so it stays below 85. But the Claude user base has been waiting for this, and it clears the featured bar.
OpenAI rolled out voice control on ChatGPT's macOS and Windows desktop apps, letting you talk to multiple agents running inside ChatGPT Work or Codex. Powered by GPT-Live, it speaks, listens, and coordinates tasks at the same time. Available globally today for Plus, Pro, Business, Edu, and Enterprise users. The post doesn't disclose latency, concurrency limits, or which desktop actions are actually controllable—worth testing before getting excited.
Why it matters: OpenAI ships voice-controlled multi-agent orchestration to desktop, a real interaction leap. GPT-Live across all paid tiers signals production readiness. Score held back because latency and concurrency limits aren't disclosed — real-world feel is still unknown.
Anthropic expanded voice mode from Haiku to Opus and Sonnet—all three models now support it. The bigger move: voice mode can now plug into Gmail, Slack, and other apps to read your emails and messages. The post doesn't disclose latency or accuracy numbers, so I'd wait for real-world tests.
Why it matters: Anthropic rolled out voice mode to Opus and Sonnet with Gmail and Slack integration — practical and newsworthy. But no latency or accuracy data in the post, so capped below 80.
Claude voice mode now lets users pick between Opus, Sonnet, and Haiku, defaulting to the last model used in text chat. Anthropic says this handles longer, more complex tasks like coaching communication style, walking through a client pitch, or brainstorming market research. The bigger shift: voice mode can now reach into Gmail, Google Calendar, Slack, Canva, and Notion to reschedule meetings, draft emails, or create docs. OpenAI's updated voice mode still can't use external tools. The post doesn't disclose latency numbers or rollout scope.
Why it matters: Anthropic swapped voice mode's backend to user-selectable models and wired it into five productivity tools — a solid practical upgrade. Not 85+ because this is feature catch-up rather than a paradigm shift, and the post doesn't disclose latency or accuracy numbers from real us...
SpaceXAI and Cursor jointly trained Grok 4.5, landing between Opus 4.7 and 4.8 in performance but 6x cheaper than Opus and 3x cheaper than GPT-5.5 on a per-token basis. OpenAI rolled out GPT-5.6 (Sol, Terra, Luna) to all users; early testers say Sol is less smart than Fable but far more reliable. ChatGPT Voice got new GPT-Live-1 and Live-1-mini models that can talk while you speak and use GPT-5.5 in the background. Anthropic extended Fable 5 access for Claude subscribers to July 12—the post doesn't explain the repeated delays. Meta introduced Muse Image and Muse Video; image editing and text rendering look solid, but images still have an AI look, and the video model is in preview.
Why it matters: SpaceXAI + Cursor joint Grok 4.5 launch with concrete performance anchor and pricing — all three HKR axes hit. Deduction because source is a newsletter summary, not a first-party announcement, and the body is truncated with incomplete GPT-5.6 info. +3 cross-source bump to 82, ...
Some users already see Bidi 1 in ChatGPT's web and app model selector, sitting alongside Standard and Advanced Voice. Selecting it turns the voice bubble yellow. The key change is full-duplex: the model can keep listening while it speaks and respond immediately to interruptions. In a demo, a user asked it to count from 1 to 10, interrupted mid-way and told it to count backwards—it complied instantly. OpenAI hasn't announced a launch date; one outlet speculates a wider rollout this week. The post doesn't disclose pricing, regional availability, or a firm timeline.
Why it matters: A substantive upgrade to OpenAI's voice capabilities — full-duplex with interrupt response is something Advanced Voice Mode couldn't do. Only partial rollout and no official announcement yet, so not pushing past 90. But the interaction change is significant enough for featured.
Doubao real-time voice model 3.0 (Seeduplex) is a native full-duplex end-to-end voice model. It claims three strengths: precise instruction following, noise resistance, and dynamic turn-taking. It stays quiet in multi-person conversations and only joins when a specified topic comes up. It can also call custom tools during real-time interaction to schedule calendar events or send emails. False replies and false interruptions are significantly reduced. Turn-taking latency dropped by about 250ms, the interruption rate in complex scenarios fell by 40%, and user-initiated interruption latency dropped by about 300ms. Target use cases include car cockpits, smart hardware, and customer service. The post does not disclose pricing or a public launch date.
Why it matters: ByteDance's first full-duplex end-to-end voice model in public beta, with concrete latency and anti-interference numbers — not pure marketing. Deduction: invite-only, no pricing or scale disclosed yet, real-world performance unverified.
Google put Gemini into a $99.99 Home Speaker that lets you correct mid-sentence and keeps a conversation going without re-waking. Premium features like free-flowing chat and Nest camera summaries require a $10/month or $100/year Home Premium subscription. Pre-orders open now, shipping this month.
Why it matters: Google re-enters smart home with a $99 Gemini speaker, with concrete pricing and features. Not scoring higher because we only have launch info — real-world experience and Gemini Live's free-form conversation aren't verified yet.
Bloomberg's Mark Gurman tested the new Siri early and listed 7 real improvements: faster responses, on-screen awareness, cross-app actions, and more natural voice. But the core issue remains—Siri still hands off complex requests to ChatGPT and only handles simple commands itself. Gurman's take: this update pulls Siri back from 'disaster' to 'barely usable,' but it's still far from the proactive assistant Apple promised at WWDC 2024.
Why it matters: Gurman's hands-on delivers real signal with 7 testable improvements, not fluff. But the core Siri problem — complex requests still fall back to ChatGPT — caps the score at 78 rather than pushing it higher.
Google released Gemini 3.5 Live Translate in public preview through the Gemini API, offering low-latency speech-to-speech translation across 70+ languages and 2,000 language pairs.
Why it matters: HKR-H/K/R all pass: Google’s speech-to-speech translation API has a clear developer hook and concrete scale numbers. Single X-source detail and missing price, latency benchmarks, and regions keep it at 78.
Google released Gemini 3.5 Live Translate, a speech-to-speech translation model that supports more than 70 languages, starts translating before the speaker finishes, uses streaming updates, and runs through Gemini Live API, Google Meet preview, and Google Translate apps on iOS and Android.
Why it matters: HKR-H/K/R all pass: Google ties real-time speech translation to 70+ languages and streaming output before the speaker finishes. It stays at 82 because rollout scope, pricing, and benchmarks are not disclosed.
Google DeepMind released Gemini 3.5 Live Translate, an audio model for near-real-time speech-to-speech translation across more than 70 languages. It detects the language automatically and preserves the speaker's intonation, rhythm and pitch.
Why it matters: The original gives the model's language coverage, how the live translation works and the rollout pace across products, enough to judge where speech translation is usable.
Google DeepMind released Gemma 4 12B, a multimodal model with a unified encoder-free architecture, native audio input, Apache 2.0 licensing, and local laptop runtime with 16GB of VRAM or unified memory.
Why it matters: HKR-H/K/R all pass: the hook is local multimodal audio in 16GB VRAM, and the new architecture is concrete. It is a strong Google DeepMind open-model release, but not a frontier-model launch, so it stays below p1.
OpenBMB released the VoxCPM2 technical report, covering a 2B-parameter speech generation model trained on more than 2 million hours of multilingual speech data, with support for 30 languages and 9 Chinese dialects.
Why it matters: HKR-H/K/R pass via the 2B size, 2M+ training hours, and dialect coverage; the score stays at the low end of 78–84 because the post lacks benchmarks, license terms, and adoption data.
Victor M summarized 25+ open-weight model releases in one week, including NVIDIA Nemotron 3 Ultra, a 550B hybrid Mamba-MoE with 55B active parameters and a 1M-token context window.
Why it matters: HKR-H/K/R all pass: the story combines a 25+ open-weight wave with NVIDIA’s 550B, 1M-context Nemotron. Reddit/X sourcing keeps it in the 78-84 band, not p1.
RedNote released dots.tts, a 2B-parameter open-source TTS model under Apache 2.0. It uses a fully continuous architecture, supports 48 kHz synthesis and zero-shot voice cloning, and maps text directly to speech without a phoneme pipeline.
Why it matters: HKR-H/K/R pass, but the source is a Reddit summary and the SOTA claim lacks benchmark names or scores. Apache 2.0, 2B params, 48 kHz, and a no-phoneme pipeline justify low featured.
Google AI announced six updates: Nano Banana 2 is generally available, Gemma 4 12B can run fully offline on laptops, and Magenta RealTime 2 is open source.
Why it matters: HKR-H/K/R all pass: the post bundles six Google AI updates with concrete local and open-source hooks. Lacking benchmarks, licensing, and pricing keeps it below the 78+ good-quality band.
Google AI for Developers released the open-weight Magenta RealTime 2 music model, supporting MIDI, live text prompts, and gestures, with native MacBook latency under 200 ms.
Why it matters: HKR-H/K/R all pass: Google Magenta MRT2 has a concrete real-time audio hook, open weights, and sub-200ms local latency. It is strong for creative-AI builders, but narrower than a general foundation-model release.
Boson AI and LMSYS released the Higgs Audio v3 TTS service with about 4B parameters, a Qwen3-4B backbone, support for 100 languages, streaming synthesis, and text tags for controlling 20+ emotions plus style, rhythm, and sound effects.
Why it matters: HKR-H and HKR-K pass via the 4B/100-language/streaming TTS hook. HKR-R is weaker because the post lacks latency, pricing, and release-form details, so this sits at the lower featured band.
Google released Gemma 4 12B, a medium-size model that runs locally with 16GB VRAM or unified memory. It uses an encoder-free multimodal architecture, supports native audio input, ships under Apache 2.0, and includes an MTP draft model for lower latency.
Why it matters: Google’s Gemma 4 12B has clear HKR-H/K/R: 16GB local running, 12B scale, and Apache 2.0 licensing. It is a strong open-model update, not a must-write foundation-model launch.
AI dictation tools are moving into developer and office workflows, with Wispr Flow reporting over 2.5 million global downloads, 70% 12-month retention, and 100x annual user growth, while OpenAI’s gpt-4o-transcribe reached a 2.5% word error rate in a third-party evaluation cited by the article.
Why it matters: HKR-H/K/R all pass, but this is a data-backed workflow trend piece, not a model launch or platform update. It sits at the lower featured threshold.
Miso One released an 8B-parameter open-weight TTS model with one-shot voice cloning from a short sample, 110ms inference latency, GitHub self-hosting without an API, and local audio data handling; the post says API access is coming but does not disclose pricing or launch timing.
Why it matters: HKR-H/K/R all pass, but this is a single X-sourced launch with no benchmark suite, license detail, or third-party reproduction. The 8B, 110ms, self-hosted open TTS facts clear featured, not higher.
Suno raised a $400 million Series D at a $5.4 billion valuation; the RSS snippet does not disclose the lead investor, participating investors, or planned use of proceeds.
Why it matters: HKR-H/K/R pass on Suno’s $400M Series D and $5.4B valuation, a clear AI-audio funding signal. Lead investor, participants, and use of funds are undisclosed, keeping it below the must-write band.
xAI partnered with Vapi to make Grok the default engine for 12 core voices, covering more than 2.5 million voice agents, and Grok Voice ranked first in Vapi’s independent blind test.
Why it matters: HKR-H/K/R all pass: the default-engine switch has scale, numbers, and voice-agent market resonance. Single-source partnership news lacks test methodology, pricing, and migration data, so it stays in the mid product-update band.
Gemini App says Gemini Omni can add users to video creation by generating a digital avatar that resembles their appearance and voice; the post does not disclose rollout scope, pricing, or safety mechanisms.
Why it matters: HKR-H/K/R all pass: the official Gemini App post has a strong multimodal avatar hook. Scope, pricing, consent, and safety controls are not disclosed, keeping it in the mid-weight product-update band.