StepFun Launches StepAudio 3 Voice Models, Several Top Artificial Analysis Rankings
Leiphone reported that StepFun launched five StepAudio 3 voice models for real-time interaction, speech recognition, TTS, audio and music generation; several ranked first on Artificial Analysis.
The report said StepAudio 3 moves beyond competing on single capabilities such as speech recognition or speech generation. The series extends to full-stack voice understanding and generation, real-time reasoning, task execution and audio creation, allowing AI to listen and speak, understand the semantics and emotion behind sound, and take part in fuller interaction and content production.
StepAudio 3 Realtime, aimed at real-time voice interaction, supports native full-duplex conversation. The report said it can decide when to respond and when to wait during continuous exchanges, and handle interruptions and continuous feedback. It scored 98.9% on the Artificial Analysis Conversational Dynamics leaderboard, ranking first globally.
The model also adds speech reasoning and task execution to the low latency and natural dialogue expected of real-time voice systems. It can understand semantics, tone, emotion and environmental sounds at the same time, and supports reasoning and speech generation in parallel. Tool calls and long tasks can run asynchronously without interrupting the current conversation, according to the report. It also ranked first on the Artificial Analysis Speech Reasoning leaderboard.
For speech recognition, StepAudio 3 ASR combines high-precision recognition with the context understanding and knowledge capabilities of a large language model. The report said it aims not only to hear clearly but also to improve understanding of professional terms, names, place names, dialects and complex context. It supports Chinese, English, dialects, mixed Chinese-English speech, long audio and professional domains, and is optimized for whispering, fast speech, unclear connected speech and background music. On the Artificial Analysis non-streaming speech recognition accuracy leaderboard, it recorded a word error rate of 1.7%, tied for first globally. A lower WER means fewer transcription errors.
StepAudio 3 TTS targets human-level voice generation. Beyond timbre, intonation, rhythm, pauses and emotional expression, it can generate paralinguistic cues common in real conversation, such as laughter, hesitation, repetition and self-correction. It uses a streaming generation architecture that allows generation and playback to happen at the same time, meeting the needs of real-time voice applications.
StepAudio 3 Gen combines voice, sound effects, ambient sounds and background music, which had previously been separate generation steps. Users can describe characters, scenes and plots in natural language and generate complete audio containing multiple sound elements in one pass, while controlling character timbre, emotion and the timing and order of different sounds. The report said it is suited to film, animation, games, radio drama, audiobooks and short video production.
StepAudio 3 Music goes beyond generating a song from a single sentence. It supports songwriting, a cappella scoring, cover songs and multi-turn interactive creation based on ABC notation. Users can work with the model on arrangement and iteration through lyrics, a cappella singing or reference songs, according to the report.
The report said voice large models have moved over the past year from single speech recognition and generation capabilities toward competition that combines real-time interaction, sound understanding and content creation. With the StepAudio 3 series, StepFun's audio model lineup covers listening, speaking, understanding, reasoning and creation. All five models are now available on the StepFun Open Platform.