Skip to content
Hacker News front page

Nari Labs pushes Qwen3-TTS to sub-50 ms time-to-first-audio at 10 RPS on a single H100

How We Made a Text-to-Speech Model Respond in Sub-50 ms

Nari Labs open-sourced a Qwen3-TTS 1.7B CustomVoice serving implementation that hits sub-50 ms p95 time-to-first-audio at 10 RPS on a single H100 SXM with zero underruns. They benchmarked against vLLM-Omni, SGLang-Omni, VoxServe, and M*—default p95 latencies ranged from 277 to 1,160 ms at 1 RPS. At full utilization the system costs roughly $2 per 1M characters, compared to $100 for ElevenLabs V3 and $49 for Cartesia Sonic 3.5. Key optimizations include dynamic leading-silence trimming (~80 ms saved) and tuned codec-frame accumulation. Code and benchmarks are public; the post does not disclose underrun details at higher concurrency or long-form performance.

Why it matters: Nari Labs open-sourced a deployment recipe for Qwen3-TTS 1.7B that hits sub-50 ms p95 time-to-first-audio at 10 concurrent requests on a single H100—an order-of-magnitude improvement over vLLM-Omni and others. The post includes concrete benchmarks and reproducible optimization...

Read the original ↗Export Markdown