All articles
LLM API Comparison & Cost/

Qwen-Audio-3.1-Realtime-Plus Pricing and Specs

A technical breakdown of API costs, context limits, and benchmark upgrades over 3.0.

4 minutes read

Qwen-Audio-3.1-Realtime-Plus Pricing and Specs. A macro close-up of a high-density liquid-cooled server blade rack, where vibrant blue coolant tubing weaves through tightly packed GPU modules, illuminated by a sharp, rhythmic strobe of — GPUniq
Data as of 4 sources

Drafted with AI tools and edited by a human. The figures and conclusions were checked by a GPUniq editor.

TL;DR

Qwen-Audio-3.1-Realtime-Plus costs $6.40 per 1M audio input tokens and $24.00 per 1M audio output tokens in Singapore on Alibaba Cloud Model Studio. Context capacity reaches 262,144 tokens while false responses to background speech drop from 73.0% down to 13.0%. A full-duplex session constraint prevents Web Search and Function Calling from running simultaneously.

How Much Does the API Cost?

Alibaba Cloud Model Studio bills qwen-audio-3.1-realtime-plus by token modality and region. Audio processing uses separate rates for input ingestion and output synthesis.

RegionText In ($/1M)Audio In ($/1M)Text Out ($/1M)Audio Out ($/1M)
Singapore (SGP)0.806.406.4024.00
China (Beijing)0.6885.5015.50120.628

In Singapore, text input costs $0.80 per 1M tokens. Text output and audio input both cost $6.40 per 1M tokens. Synthesized audio output costs $24.00 per 1M tokens.

Deploying in Beijing drops rates to local pricing. Text input costs $0.688 per 1M tokens. Text output and audio input cost $5.501 per 1M tokens. Audio output costs $20.628 per 1M tokens.

Audio output drives total billing quickly. Live streams consume far more tokens than plain text responses.

How Does 3.1 Compare to 3.0?

Benchmark data shows performance gains over Qwen-Audio-3.0-Realtime across speech tasks.

Metric / BenchmarkQwen-Audio-3.0-RealtimeQwen-Audio-3.1-Realtime-PlusDelta
Audio MultiChallengeBaselineBaseline + 5.09%+5.09%
14-language QA (Big BENCH Audio)BaselineBaseline + 6.40%+6.40%
τ²-Bench Audio Task SuccessBaselineBaseline + 3.59%+3.59%
τ-Voice Adaptation (Speech-to-Text)78.4%82.0%+3.60%
Background Speech False Response73.0%13.0%-60.00%

Accuracy on the half-duplex speech-to-text adaptation of τ-Voice rose from 78.4% to 82.0%. Audio MultiChallenge improved by +5.09%, while the 14-language QA evaluation on Big BENCH Audio gained +6.40%.

The major operational fix is background noise handling. On Full-Duplex-Bench v1.5, false responses to background chatter dropped from 73.0% down to 13.0%.

What Is the Maximum Context Window?

qwen-audio-3.1-realtime-plus supports 262,144 tokens of total context.

The context splits unevenly between input and generation:

  • Maximum Input Allocation: Up to 245,760 tokens for audio streams, text prompts, and instructions.
  • Maximum Output Generation: Capped at 16,384 tokens per response.

The API automatically truncates history after approximately 50 turns or 300 seconds of continuous audio. Track frame duration client-side instead of assuming endless buffer retention.

Does It Support Search and Tools?

The model supports Function Calling and real-time Web Search. However, Web Search and Function Calling cannot both be enabled at the same time.

{
  "model": "qwen-audio-3.1-realtime-plus",
  "modalities": ["audio", "text"],
  "enable_web_search": true,
  "tools": [] 
}

If your stack requires external lookups alongside search, route them sequentially. Execute database lookups first using function calls, then pass that context into a search-enabled session.

The API offers system voices like longanqian_v3.1 and longanhuan_v3.1 alongside voice cloning. Note that setting an output voice language operates on a best-effort basis. The engine does not guarantee language overrides, and text output language remains unaffected.

What Are the Audio Stream Requirements?

You cannot download model weights. qwen-audio-3.1-realtime-plus is API-only, hosted on Alibaba Cloud Model Studio via WebSocket, AOQ, or WebRTC protocols.

Audio chunks must match exact specs to prevent errors:

  • Audio Input Format: Single-channel mono PCM, 16 kHz sample rate, 16-bit depth.
  • Audio Output Format: Single-channel mono PCM, 24 kHz sample rate, 16-bit depth.
DirectionFormatSample RateBit DepthChannels
Client InputPCM16 kHz16-bitMono
Stream OutputPCM24 kHz16-bitMono

Turn control uses push-to-talk, client VAD, or the native smart_turn auto-detection. Audio input quality may degrade over weak networks if format and sample rate requirements are not maintained.

Should You Upgrade to Version 3.1?

Upgrade if you deploy voice agents in noisy spaces. Reducing background false triggers from 73.0% to 13.0% fixes random unprompted interruptions.

Upgrade immediately for:

  • Noisy Environments: Warehouses, kiosks, or mobile apps with ambient noise.
  • Multilingual Coverage: Voice apps using English, Chinese, Japanese, Korean, German, French, Spanish, Portuguese, Russian, Indonesian, Italian, or Chinese dialects.

Hold off if:

  • Your backend executes function calls inside active web search sessions. You must refactor that workflow into separate calls before upgrading.

Methodology

Facts are compiled from the vendor's own announcement and documentation, linked below, at the time of writing; where the vendor hasn't published a number it is marked as such. GPUniq prices, when the model is already available here, come from our live catalog.

Sources

  1. 1.Alibaba Cloud Model Studio: qwen-audio-3.1-realtime-plus — Alibaba Cloud Documentation Center (accessed )
  2. 2.Speech-to-speech models overview — Alibaba Cloud Model Studio (accessed )
  3. 3.Qwen-Audio-3.1-Realtime: Towards Reliable Agentic Voice Interaction (arXiv) — arXiv (accessed )
  4. 4.Qwen-Audio-3.1-Realtime — Beyond conversation (for capabilities & demos) — QwenAudio official site (accessed )

Want cheap GPUs for your next project?

Browse live GPU prices and rent the right card in seconds — H100, A100, RTX 4090, and 50+ more models.

More from the GPUniq blog