Skip to content
AI HOT (Curated Pool)

Google launches Gemini 3.5 Live Translate with near-real-time speech-to-speech translation for 70+ languages

Google AI 发布 Gemini 3.5 Live Translate,支持 70+ 语言近实时语音到语音翻译

Gemini 3.5 Live Translate processes raw audio streams directly and preserves the speaker's tone, rhythm, and pitch. Southeast Asian super-app Grab is exploring it for cross-language driver-passenger calls—Grab users make over 10 million voice calls per month. Developers can integrate via the Gemini Live API with LiveKit, Fishjam, Pipecat, or Vision Agents. LiveKit already demonstrated real-time multilingual understanding in virtual meeting rooms; Software Mansion used the MoQ protocol to break through streaming bottlenecks; VisionAgents AI showed dynamic language switching. Developers can try it now in Google AI Studio and grab Cookbook sample code.

Why it matters: Google ships end-to-end speech translation with a Grab pilot at 10M+ monthly calls—strong tech signal and real-world validation. Held below 85 because the post doesn't disclose latency in ms or translation quality metrics, so the claim is directional for now.

Read the original ↗Export Markdown