Models

Alibaba's Qwen launches Audio 3.1 with five models, cuts prices up to 95%

Alibaba's Qwen team has released Qwen-Audio-3.1, a five-model lineup for speech recognition, text-to-speech, and real-time interaction, with AI audio prices cut by up to 95 percent.

By Rebecca Stone3 min read

Updated

Why it matters

  • Alibaba's Qwen team released Qwen-Audio-3.1, a lineup of five models for ASR, TTS, and real-time interaction.
  • Alibaba cut AI audio prices by up to 95 percent alongside the release.
  • ASR-Next adds multi-speaker identification with timestamps plus detection of emotions, ambient sounds, and machine noise.
  • The base ASR model improves multilingual and dialect recognition and automatically removes filler words and repetitions.
  • The TTS model handles multilingual speech synthesis.

Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models covering speech recognition (ASR), text-to-speech (TTS), and real-time interaction — and cut prices for AI audio services by up to 95 percent.

The announcement signals Alibaba's intent to compete on cost in an AI audio market where voice interfaces are becoming a primary way users interact with models. Aggressive price reductions of this scale pressure rivals to follow, especially in markets where Chinese cloud providers already undercut Western competitors on inference costs.

What does the Qwen-Audio-3.1 lineup include?

The release comprises five models spanning three core audio capabilities:

  • ASR (automatic speech recognition) — the base speech-to-text model
  • ASR-Next — an advanced recognition model with additional perception features
  • TTS (text-to-speech) — multilingual speech synthesis
  • Real-time interaction models for live voice applications

The base ASR model improves multilingual and dialect recognition, according to the company. It also automatically cleans up filler words and repetitions from transcripts — a feature aimed at producing usable text from raw, conversational speech without post-processing.

What can ASR-Next do?

ASR-Next, the more capable recognition model in the lineup, adds multi-speaker identification with timestamps. That means the model can distinguish who said what in a conversation and anchor each utterance to a point in time — a requirement for meeting transcription, call-center analytics, and subtitle generation.

Beyond speaker separation, ASR-Next detects:

  • Emotions in speech
  • Ambient sounds
  • Machine noise

Perception of non-speech audio broadens the model's use cases beyond dictation. Systems that recognize background machinery, crowd noise, or a speaker's emotional state can support industrial monitoring, accessibility tools, and customer-service applications that respond to more than words.

What about text-to-speech?

The TTS model handles multilingual synthesis, allowing a single system to generate spoken audio across languages. Combined with the real-time interaction models, the lineup covers the full pipeline from hearing speech, to understanding it, to responding in natural voice — the architecture behind modern voice assistants and AI agents.

Why do the price cuts matter?

Alibaba is slashing AI audio prices by up to 95 percent alongside the release. The move mirrors the broader price war across Chinese AI providers, which have repeatedly cut inference costs for large language models to win developers and enterprise customers.

For audio specifically, cost has been a gating factor. Applications that transcribe thousands of call-center hours, stream real-time voice interaction, or synthesize speech at scale rack up substantial inference bills. A 95 percent reduction changes the economics of those workloads and could shift voice-enabled products from premium features to default components.

It also positions Qwen against dedicated speech AI providers — companies such as those specializing in transcription and voice APIs — whose pricing Alibaba now undercuts directly. Developers building on Alibaba's cloud gain a full-stack audio option at a fraction of previous cost.

What's the context for the release?

Qwen is Alibaba's flagship AI research and model family, and the group has shipped successive model updates across text, vision, and audio modalities. Audio has grown strategically important as voice becomes a primary interface for AI agents, customer-service bots, and multimodal assistants.

The Qwen-Audio-3.1 release consolidates recognition, synthesis, and real-time interaction into one coordinated lineup. That bundling matters for developers: instead of stitching together ASR, TTS, and dialogue components from separate vendors, they can build end-to-end voice systems within a single ecosystem — now at dramatically reduced cost.

With five models, improved multilingual and dialect coverage, and perception features like emotion and noise detection baked into ASR-Next, Alibaba is betting that low prices plus breadth of capability will pull developers onto its platform as voice-driven AI applications scale.

Original: fun-resource-shanghai.oss-cn-shanghai.aliyuncs.com

Share this article:

More from Rebecca Stone

Rebecca Stone

Show full bio

Correspondent covering consumer brands and retail at AI In Context.

214 articles

Related articles

  1. OpenAI ships three realtime audio models, led by GPT-Realtime-2
  2. Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
  3. ElevenLabs launches Eleven v4 with tighter voice control
  4. Google Ships Gemini 3.8 Live and Extended Thinking Models
  5. Google Launches Gemini 3.1 Flash TTS With Audio Tags and 70+ Languages

« Previous articleNext article »