Zyphra open-sources ZONOS2: an 8B-param, 900M-active real-time TTS with high-fidelity voice cloning
ZONOS2: real-time TTS with 8B params, 900M active, and high-fidelity voice cloning
Zyphra released ZONOS2, an open-source real-time TTS model with 8B total params and only 900M active at inference. It uses a sparse MoE design to balance speed and expressiveness, with zero-shot voice cloning as the headline feature. The model reads raw UTF-8 bytes instead of using a phonemizer, which helps with Chinese, Korean, Japanese, and mid-sentence code-switching. Audio runs through a 44.1kHz Descript Audio Codec for studio-quality output, and training data scaled from 200K to over 6M hours. On the TTSDS prosody benchmark it scores 88.7, ahead of Qwen 3 TTS, Cartesia Sonic 3.5, and ElevenLabs V3. Weights are Apache 2.0, and inference is also available on Zyphra Cloud running AMD hardware.
Why it matters: ZONOS2 combines real-time inference, high-fidelity cloning, and Apache 2.0 licensing in local TTS — the engineering choices are worth a look. Score stays at 74 because we only have the Reddit post; no independent benchmarks or side-by-side comparisons yet, so real-world clonin...