Tongyi Lab releases Qwen-Audio-3.0-TTS; Plus variant tops Artificial Analysis leaderboard
通义实验室发布 Qwen-Audio-3.0-TTS 实时语音合成模型
Tongyi Lab released Qwen-Audio-3.0-TTS, a real-time speech synthesis model with Flash and Plus variants. Flash hits ~300ms first-packet latency and 3.87 average WER/CER. Plus tops the Artificial Analysis leaderboard, reaches 82.75 speaker similarity, and supports 16 languages plus 20 Chinese dialects. The post does not disclose parameter count, inference hardware, or API pricing.
Why it matters: Qwen Lab drops a real-time TTS model with two clear tiers: Flash for low latency (~300ms first packet) and Plus topping the Artificial Analysis leaderboard. Missing param count, hardware requirements, and pricing keeps this from scoring higher — those are the gaps that matter ...