Skip to content
AI HOT (Curated Pool)

ByteDance launches Doubao-Seed-Audio 1.0: one prompt generates multi-character dialogue, music, and ambience

豆包音频生成模型1.0发布,重新定义AI音频创作

ByteDance's Volcano Engine released Doubao-Seed-Audio 1.0, an end-to-end audio generation model driven by text or audio references. A single prompt can arrange multi-character dialogue, emotional tone, background music, and ambience, keeping voice timbre consistent for up to 2 minutes and across extensions. It works zero-shot with no extra training and supports one voice playing multiple roles. The model is in invite-only testing on Volcano Ark API with 30 free minutes per user, and will roll out to Jianying, Jimeng, and Tomato.

Why it matters: ByteDance/Volcano Engine released Doubao Audio Generation Model 1.0, a domestic flagship model launch that policy treats on par with US lab releases. End-to-end multi-character audio generation is a differentiated capability, hitting all three HKR axes. Score stays at 78 rathe...

Read the original ↗Export Markdown