Skip to content
AI HOT (Curated Pool)

ByteDance launches Seed Audio 1.0, a single model that generates dialogue, sound effects, and ambience end-to-end

字节跳动发布 Seed Audio 1.0 音频创作模型,统一建模人声与音效实现端到端影视级音频生成

Seed Audio 1.0 uses a universal acoustic encoder to model speech, sound effects, and ambience as one scene instead of stitching separate outputs. A single prompt controls character lines, emotion, and sound cue timing at 100ms precision. It does zero-shot voice cloning from a reference clip, generates roughly 2 minutes per pass, and can extend audio while keeping the voice consistent. It covers 20+ languages and adapts rhythm and pronunciation per language. Human evals show >90% usability in film, podcast, and short-drama scenarios, with MOS above 4 for most languages. Available now on Volcano Engine's experience center.

Why it matters: ByteDance drops a unified audio creation model that generates speech, sound effects, and ambience end-to-end with 100ms timeline control — a flagship release from a major Chinese lab. Held below 85 because we only have the official blog post so far; no third-party tests or hea...

Read the original ↗Export Markdown