Alibaba's Qwen Releases Qwen-Audio-3.1-Realtime, a Full-Duplex Voice Model, and Cuts Audio API Prices by Up to 95%
Alibaba's Qwen launched Qwen-Audio-3.1-Realtime, a full-duplex voice agent model, cutting audio API prices by up to 95%.
The model is offered only as a managed API. Qwen-Audio-3.1-Realtime-Plus is live on QwenCloud over WebSocket, and no open weights were announced. Its model page lists text and audio as both input and output, a 262K-token context with a 245K maximum input and 16K maximum output, and default limits of 60 requests and 100,000 tokens per minute. Pricing is $6.4 per million audio input tokens and $0.8 per million text input tokens, with output billed at $24 per million tokens; output text is not charged. Function calling, web search, structured outputs, context cache and fine-tuning are listed as supported.
A companion model, Qwen-Audio-3.1-ASR-Flash-Filetrans, targets offline long-audio transcription and supports hot words, speaker separation, punctuation and multilingual recognition including Chinese dialects. It costs $0.15 per million input tokens and $0.47 per million output tokens, according to the same report.
The system runs two models that share the same audio encoder and LLM design. A full-duplex decision model predicts whether to keep listening, speak, stop or resume, while a speech-to-text model writes the response content as text. A context-aware voice renderer then converts that text into streaming speech, conditioned on conversation history, voice cues and acoustic context.
Training is organized into three layers: Think, Act, and Speak and Coordinate. In the Think layer, M²-OPD Core-Cocktail supervised fine-tuning re-anchors the audio model to its source text LLM using million-hour-scale paired data, followed by multimodality OPD, in which a Text Teacher and a frozen Audio Reference score each token of the student's own trajectory in an on-policy distillation setup rather than imitating pre-written answers. Domain experts for empathy, pragmatic intent and acoustic scenes are then trained with GRPO and merged by Multi-Teacher OPD into one deployable model.
The Act layer bundles each training domain with a tool pool, a stateful JSON database and a natural-language business policy, seeded from open-source tool and MCP server definitions. Every task defines one of three outcomes: a write, a justified refusal, or an unsupported request. Scoring checks terminal state first, then permitted writes, then behavioral assertions, and a fluent reply cannot rescue a failed state check. GRPO rewards arrive at dialogue, milestone and turn level. Search training penalizes redundant queries; mean queries per search call fell from 4.37 to 1.05, while trigger F1 slipped from 60.87% to 58.61%.
The Speak and Coordinate layer decides whether, when and how to speak. On Full-Duplex-Bench v1.5, replies to people talking to someone else fell from 0.13 to 0.03, and on v3.0 the filler rate dropped from 0.7590 to 0.2960. The report notes trade-offs: after interruptions, the unwanted resume rate rose from 0.035 to 0.130, and interruption stop latency is 1.116 seconds, against 0.383 seconds for GPT-Realtime-2.
Against Qwen-Audio 3.0, Audio MultiChallenge rises from 47.12 to 52.21, the 14-language BBA average climbs from 81.7% to 88.1%, and FLEURS word error rate falls from 9.01 to 3.98. Task success on a τ-Voice adaptation rises from 78.4% to 82.0%, and replies to background speech on Full-Duplex-Bench v1.5 drop from 73.0% to 13.0%. The report states that the τ-Voice figures use a half-duplex speech-to-text adaptation and are not comparable to official full-duplex results.
In a 50-session human red-team study cited in the report, GPT-Realtime-2 leads with 96.00% against Qwen's 92.00%. Multi-turn attack success falls to 26.0% in Chinese and 23.5% in English.
Compared with competing APIs, the report lists Qwen-Audio-3.1-Realtime-Plus at a 262K context, 16K maximum output, $6.4 per million audio input tokens and $24 per million audio output tokens; OpenAI's GPT-Realtime-2 at 128K context, 32K maximum output, $32 and $64; and Google Gemini 3.8 Live at 131,072 context, 65,536 maximum output, $3.00 and $12.00. None of the three offers open weights. The report cautions that the prices are list rates checked on September 28, 2026, and that token rates are not directly comparable because each provider tokenizes audio differently.