StepFun Launches Step 5 Preview, an Open-Source Flagship With 27B Active Parameters
StepFun has released Step 5 Preview, a 600B-parameter MoE flagship with 27B active parameters and 1M context, ranking among the top two open-source models on Artificial Analysis while undercutting rivals on cost. Tests show strong agentic generation in 3D games, Blender and websites, though human review is still needed.
QbitAI said it tested Step 5 Preview through an API, first asking it to operate Blender and build a Backrooms scene. After about an hour, the model produced a rendered MP4 that included VHS videotape texture, lens noise, dim yellow lighting and footstep sounds during character movement, the report said. In a second task, it created a LEGO Racers-inspired racing game with a track, cars, human and AI drivers, real-time rankings and engine sound effects. The test also generated a neon parkour game, a writing website and a 3D model of a flower house.
The tests also found remaining flaws. The first version often missed small details and required two or three rounds of corrections to become usable, according to QbitAI. The writing website had module overflow bugs and still needed review and editing. QbitAI said the output was not yet at the level of some computer-use demos but was usable for work.
StepFun attributed the agent improvement to a network design it calls Narrow but Deep. The model has 92 Transformer layers, which the company says gives information longer propagation and reasoning paths for multi-hop agent tasks. StepFun also used sparse techniques including Sparse MoE, Hybrid Sparse and Sparse GQA, along with low-level operator and data-access optimizations, to control the cost of long contexts. The model was trained around agent loops, using long-horizon reinforcement learning and context compaction so it can keep track of goals after dozens of tool calls.
The release marks a return to the frontier for StepFun after two previous Flash-tier models, Step 3.5 Flash and Step 3.7 Flash, QbitAI reported. The report placed the launch in a crowded field where frontier models are closer than before: as of March 2026, the top models from Anthropic, xAI, Google and OpenAI were separated by less than 25 Elo points on Arena, according to the Stanford 2026 AI Index cited by QbitAI. Companies are increasingly choosing distinct strengths such as front-end development, low cost or computer use, rather than chasing a single benchmark.
The same shift is also being felt by users. In an essay published by sspai, the author argues that AI agents are making execution skills cheap. The essay compares the situation to the middle-income trap: after early gains from AI, users face rising competition as equally capable models become cheaper and information gaps disappear. Prompt techniques and tool-specific know-how can deliver short-term advantages, but they depend on immature tools and unequal access, and tend to depreciate as models, harnesses and community practices improve, the author writes.
The sspai essay distinguishes external capital, such as users, social accounts, distribution and reusable workflows, from internal capital such as judgment and taste. It argues that execution ability is increasingly rented from AI, while judgment and taste are less likely to lose value as models improve. The author recommends making one's own prediction before asking an agent, seeking evidence that could disprove that judgment, and treating mistakes as material for transferable principles rather than one-off corrections. Taste, the essay says, develops through comparison across many good products and cross-disciplinary works, as well as through actual creation.