AI News Feed
Market watch
Large Language Models

Alibaba's Qwen-Audio-3.1-TTS Tops Global Speech Leaderboard With 1177 Elo

Alibaba's Qwen-Audio-3.1-TTS ranked first worldwide on Artificial Analysis's Controlled Voice Arena with an Elo score of 1177, according to Leiphone. The multilingual speech synthesis model is part of the Qwen-Audio-3.1 series unveiled at the 2026 Yunqi Conference.

The model is Alibaba's latest-generation text-to-speech system and supports synthesis across multiple languages and dialects. According to Leiphone, it lets a single voice timbre migrate naturally between languages, so one voice can speak several languages without obvious seams. Where conventional speech synthesis concentrates on whether individual characters are pronounced correctly, Qwen-Audio-3.1-TTS places more emphasis on whether a voice fits the specific content and the setting in which it is used. Users can steer emotion, speaking rate and delivery style through instructions, so that generated speech does not merely sound correct but matches the intended tone.

The TTS model belongs to the Qwen-Audio-3.1 series that Alibaba released at its 2026 Yunqi Conference. The series also covers automatic speech recognition and realtime voice interaction models, and API services for all three have been listed on the Qwen AI platform, Leiphone reported.

Beyond the hosted services, Alibaba's speech team has open-sourced several models, including Qwen-Audio-Agent and CosyVoice, which together have drawn tens of thousands of stars on GitHub and attention from developers at home and abroad, according to the same report.

Leiphone also reported that prices for the full Qwen-Audio line of speech models have been cut, with the largest reduction reaching 95 percent, lowering the cost of speech applications for developers and enterprises.