Skip to content
AI HOT (Curated Pool)

NVIDIA Releases NemotronLabs VoiceChat 11B: An Open Full-Duplex Speech-to-Speech Model with ~450 ms Turn-Taking and Live Tool Calling

NVIDIA 发布 NemotronLabs VoiceChat 11B:开源全双工语音模型,支持约 450 毫秒轮换与实时工具调用

NVIDIA open-sourced an 11B end-to-end speech-to-speech model that handles streaming understanding and generation in one network, skipping the usual ASR-LLM-TTS pipeline. Measured turn-taking latency is 448 ms, and it supports live tool calling during conversation. Weights and code are public.

Why it matters: NVIDIA open-sourced an 11B end-to-end speech model with 448 ms interruption latency and full-duplex turn-taking — real engineering progress, not a benchmark flex. But the MarkTechPost piece reads like a product announcement with no third-party benchmarks or head-to-head compar...

Read the original ↗Export Markdown