AI News Feed
Market watch
Large Language Models

StepFun Releases Step 5 Preview, Claims Global Open-Source Top Three on AA Index

StepFun released Step 5 Preview, a flagship base model for coding, software engineering, professional work and finance, saying it ranks in the global open-source top three on the Artificial Analysis Intelligence Index and costs one-eighth as much per task as Claude Opus 5. It will be open-sourced on Oct. 15.

Step 5 Preview uses a sparse MoE architecture with 600 billion total parameters and 27 billion activated parameters, supports a 1 million token context window and natively handles text and visual inputs, according to Leiphone. StepFun positions it as a main work model for real-world agentic tasks that can handle complex jobs over sustained interactions.

APPSO, a media outlet under Ifanr, reported after testing the model through WorkBuddy that Step 5 Preview scored 44 on the Artificial Analysis Intelligence Index, placing it in the top three open-source models and on the Pareto frontier balancing intelligence and per-task cost. Leiphone reported the per-task cost is one-eighth that of Claude Opus 5. StepFun said the model improves the boundary between intelligence and cost.

In public and internal evaluations, StepFun said Step 5 Preview performed well on complex software engineering tasks, including long-horizon planning, as well as professional knowledge work and finance. The company built three internal evaluations around company research, a difficult financial workflow, and introduced FrontierFinance, which contains 220 expert questions and 11,543 evaluation criteria. Official results showed Step 5 Preview was close to Claude Opus 5 on four finance benchmarks, including information retrieval, corporate valuation and financial research, according to APPSO.

APPSO's hands-on tests covered programming, design, 3D generation, deep research and data analysis. In Three.js tasks, the model generated a five-in-a-row game with human-versus-machine play, a Venice speedboat scene, a Cappadocia hot-air balloon scene and a Niagara Falls recreation, maintaining scene structure, interaction logic and runtime efficiency. It also produced a web-based operating system page that APPSO said was more complete than earlier tests with DeepSeek V4 Pro and DeepSeek V4.1 Flash. For design and 3D assets, it handled space planet and nighttime sports car themes and converted a casually taken photo into a model with basic structure and spatial detail within minutes, APPSO reported.

On research and financial tasks, APPSO said Step 5 Preview was asked to produce a folding-phone project report by linking user complaints, cover-screen ecosystem and after-sales costs; plan a seven-day Guangzhou-Foshan trip under budget, stamina and strict vegetarian constraints; and analyze a messy sales table by cleaning data, tracking negative-margin orders by region and category, and generating monthly sales trend charts. The model followed constraints and produced reviewable output, according to APPSO.

StepFun has also built a multimodal model matrix. In the past quarter it released the Step Edge on-device model and StepAudio 3 speech models. Step Edge combines text and vision base models with speech, GUI and generation capabilities, supports simple tasks on device and complex tasks in the cloud, and ranked first in 29 evaluation items, according to APPSO. StepAudio 3 Realtime scored 98.9% on Artificial Analysis conversational dynamics and 99.7% on speech reasoning tests, while StepAudio 3 ASR had a 1.7% word error rate, tied for first globally in non-streaming speech recognition, APPSO reported.

On deployment, StepFun's official data show its models are installed on more than 42 million phones and serve nearly 20 million people daily, with partnerships covering more than 50% of leading Chinese phone brands. Use cases include screen-aware Q&A, intelligent search, content generation and cross-app task execution. In automobiles, StepFun, Geely and Qianli Technology jointly developed the vehicle intelligent agent Super Eva, first used in the Zeekr 8X. Super Eva integrates the Step 3.5 Flash base model and StepFun's speech and visual understanding models. Geely Galaxy M9 also adopted StepFun's end-to-end speech model, which the companies said was the industry's first mass production deployment of such a model, according to APPSO.

In July, StepFun launched STEPX, a large-model-native AI terminal brand, and said it will soon release the large-model-native agent phone STEPX Neo, according to APPSO. APPSO reported that this gives StepFun a model-software-hardware loop spanning models, systems and terminals.