Nari Labs tops Coval voice AI benchmarks on latency and accuracy for both STT and TTS
Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost
Nari Labs placed both its Qwen3-ASR and Qwen3-TTS 1.7B models on the quality-latency Pareto frontier in Coval's voice AI benchmarks. STT hits 44 ms median time-to-final-segment with 3.6% WER, second only to AssemblyAI's 3.5% but at less than a quarter of the cost. TTS achieves 63 ms median time-to-first-audio and 3.8% WER, ranking first, while tying for the cheapest public price at $10 per 1M characters. The official Qwen3 TTS Flash Realtime endpoint scores 8.8% WER and 692 ms latency on the same benchmark, so Nari's serving stack makes a big difference. The post doesn't disclose the audio dataset makeup or p95/p99 tail latencies.