September 14, 2026

LLM Tools|Index 05

Nari Labs Optimizes Open-Source Speech Models for Speed and Cost

Nari Labs has released an optimized inference engine for Qwen3-TTS and ASR, claiming sub-50ms latency and superior accuracy over established proprietary services at a fraction of the cost.

Via
AITECH TOKYO Editors
Dateline
September 14, 2026
Date
September 14, 2026
Time
6 min read
Nari Labs Optimizes Open-Source Speech Models for Speed and Cost

Tagline

Optimized open-source speech models, faster and cheaper.

Who & Why

For a Tokyo-based developer or product manager building voice-enabled applications who needs high-performance, cost-effective, and accurate speech-to-text or text-to-speech without vendor lock-in.

vs. Existing

This directly competes with proprietary speech APIs like 11Labs and Cartesia, offering superior accuracy and latency at a lower cost, while also outperforming Alibaba's official Qwen endpoints by optimizing open-source models.

Tokyo Take

While promising for global developers, the immediate impact for Tokyo professionals depends on robust Japanese language fine-tuning and local adoption, as the current benchmarks focus on English. The cost advantage could be significant if JPY pricing is competitive.

Nari Labs, a startup focused on efficient AI inference, has open-sourced a specialized engine for Qwen3-TTS (Text-to-Speech) and Qwen3-ASR (Automatic Speech Recognition) models. This initiative aims to make high-quality, open-source speech technology accessible and affordable for developers.

The company built its inference engine to address the perceived shortcomings of existing multimodal inference systems, which it argues are not well-suited for speech models. Their solution, available on GitHub, demonstrates that open models can achieve performance benchmarks typically associated with closed-source offerings.

Nari Labs reports significant performance gains. Their Qwen3-TTS endpoint achieves sub-50ms latency at 10 requests per second (RPS), positioning it as a top contender in speed. For Qwen3-ASR, they claim the lowest latency among benchmarks.

Beyond speed, Nari Labs asserts superior accuracy and cost-effectiveness. Measured against the Coval voice AI benchmarks, their Qwen3-TTS endpoint ranks second in latency and first in accuracy (WER) among competitors like 11Labs and Cartesia, while being the cheapest option available.

Similarly, their Qwen3-ASR endpoint is second in accuracy, just 0.1% behind the leader, and is the second cheapest model on the list. The company notes that even Alibaba's official endpoints for Qwen models perform worse in accuracy and latency compared to their optimized serving.

"We want to continue to push prices down to make speech technology a commodity."

This development could significantly impact developers and businesses looking to integrate sophisticated speech capabilities into their applications without incurring the high costs or vendor lock-in of proprietary services. For a Tokyo professional, this means potentially cheaper and more performant speech-enabled applications, especially if these optimizations extend to Japanese language models or are adopted by local service providers.

Nari Labs also indicates future work in other audio domains like diarization, as well as video and world model inference, suggesting a broader ambition to optimize multimodal AI.

The Briefing

World AI tech, read from Tokyo. Once a week, in Japanese.

Each Friday: the five global AI tech stories Japanese business professionals should know about this week, translated and read through a Tokyo lens — what it means for Japan, what to act on, what to keep watching.

We respect your inbox. Unsubscribe anytime.