Hacker News·3 min read·hard

Show HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost

T
toebee
AI Summary

Nari Labs has announced its Qwen3-TTS and Qwen3-ASR models, which currently lead industry benchmarks for voice AI latency and cost-efficiency. The models are designed to improve the responsiveness of voice agents by minimizing time-to-first-audio and word error rates.

Why it matters

Advancements in low-latency voice AI are critical for the development of more natural and efficient real-time conversational AI agents.

Dive DeeperCreate a free account to unlock

Nari Labs leads Coval’s voice AI benchmark by sitting on the quality-latency Pareto Frontier for both Text-to-Speech and Speech-to-Text. We also lead the latency-cost and quality-cost Pareto Frontier out of all publicly available models on the benchmark.

Coval is a leading provider of voice AI evaluation and benchmarks. They help speech AI agents perform better in production and publish one of the most widely cited benchmarks in the industry.

The Text-to-Speech (TTS) benchmark evaluates latency from text input to first audible chunk of audio (time-to-first-audio or TTFA) and Word Error Rate (WER). The Speech-to-Text (STT) benchmark evaluates latency from user’s finalize request to the final text output (time-to-final-segment or TTFS) and Word Error Rate (WER).

TTFA and TTFS are critical for voice agents, where latency can make a voice AI agent feel unresponsive. Low WER is an obvious key factor for model performance as well.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyaibusiness

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in