ByteDance Seed launches SeedRealtime, a native audio-video full-duplex model, now live in Doubao
字节 Seed 发布 SeedRealtime 音视频全双工大模型,走向全模态自然交互
SeedRealtime fuses audio, video, and text into a single end-to-end model, ditching the cascaded ASR-VLM-TTS pipeline. It watches, listens, and speaks in a continuous stream, deciding in real time when to jump in and whom to track. Human evals show half the turn-taking issues vs. cascaded systems—fewer cut-offs, late replies, or false triggers from background chatter. It also acts proactively: it can alert you when a target exhibit appears in a museum or correct a coffee-making mistake on the spot. The model is now fully rolled out in Doubao's video call feature.
Why it matters: ByteDance Seed released SeedRealtime, a native audio-video full-duplex model with a unified architecture. Human eval shows interaction-rhythm issues halved vs. cascade systems, and it's already live on Doubao App. This is the first scaled full-duplex multimodal launch from a m...