LLM Tools|Index 05
Nari Labs Optimizes Open-Source Speech Models for Speed and Cost
Nari Labs has released an optimized inference engine for Qwen3-TTS and ASR, claiming sub-50ms latency and superior accuracy over established proprietary services at a fraction of the cost.
- Via
- AITECH TOKYO Editors
- Dateline
- September 14, 2026
- Date
- September 14, 2026
- Time
- 6 min read
Source
Hacker News TopTagline
Optimized open-source speech models, faster and cheaper.
Who & Why
For a Tokyo-based developer or product manager building voice-enabled applications who needs high-performance, cost-effective, and accurate speech-to-text or text-to-speech without vendor lock-in.
vs. Existing
This directly competes with proprietary speech APIs like 11Labs and Cartesia, offering superior accuracy and latency at a lower cost, while also outperforming Alibaba's official Qwen endpoints by optimizing open-source models.
Tokyo Take
While promising for global developers, the immediate impact for Tokyo professionals depends on robust Japanese language fine-tuning and local adoption, as the current benchmarks focus on English. The cost advantage could be significant if JPY pricing is competitive.
Nari Labs, a startup focused on efficient AI inference, has open-sourced a specialized engine for Qwen3-TTS (Text-to-Speech) and Qwen3-ASR (Automatic Speech Recognition) models. This initiative aims to make high-quality, open-source speech technology accessible and affordable for developers.
The company built its inference engine to address the perceived shortcomings of existing multimodal inference systems, which it argues are not well-suited for speech models. Their solution, available on GitHub, demonstrates that open models can achieve performance benchmarks typically associated with closed-source offerings.
Nari Labs reports significant performance gains. Their Qwen3-TTS endpoint achieves sub-50ms latency at 10 requests per second (RPS), positioning it as a top contender in speed. For Qwen3-ASR, they claim the lowest latency among benchmarks.
Beyond speed, Nari Labs asserts superior accuracy and cost-effectiveness. Measured against the Coval voice AI benchmarks, their Qwen3-TTS endpoint ranks second in latency and first in accuracy (WER) among competitors like 11Labs and Cartesia, while being the cheapest option available.
Similarly, their Qwen3-ASR endpoint is second in accuracy, just 0.1% behind the leader, and is the second cheapest model on the list. The company notes that even Alibaba's official endpoints for Qwen models perform worse in accuracy and latency compared to their optimized serving.
"We want to continue to push prices down to make speech technology a commodity."
This development could significantly impact developers and businesses looking to integrate sophisticated speech capabilities into their applications without incurring the high costs or vendor lock-in of proprietary services. For a Tokyo professional, this means potentially cheaper and more performant speech-enabled applications, especially if these optimizations extend to Japanese language models or are adopted by local service providers.
Nari Labs also indicates future work in other audio domains like diarization, as well as video and world model inference, suggesting a broader ambition to optimize multimodal AI.
Adjacent Tools
LLM Tools
Nvidia CEO Vows No AI Slowdown
Jensen Huang's statement to Trump sets the expectation for continued rapid advancement in AI, with significant implications for global tech strategy.
LLM Tools
iOS 27's Siri: A Conversational Assistant Finally Comes of Age
Apple's voice assistant, long a source of frustration, has been fundamentally re-engineered in iOS 27, offering contextual understanding and deep application integration that fundamentally redefines its utility for everyday tasks.
LLM Tools
Jaron Lanier: AI is Just People, Not a Mind
A leading critic challenges the notion of AI as an independent intelligence, reframing it as a mirror of human data and labor.