AI News Feed
Market watch
Large Language Models

Chinese AI Teams Push Real-Time World Models, Open-Source Foundation Models and 3D Asset Pipelines

Leiphone's Sept. 30 reports detail tests of Step 5 Preview, PixVerse R2, HiDream-O1-Video, Aholo Lux3D and AgnesCode, showing Chinese AI teams pushing real-time world models, open-source foundation models, video causality and 3D asset pipelines into production workflows.

Leiphone reported that StepFun's Step 5 Preview scored 44 on the Artificial Analysis Intelligence Index, placing it second among global open-source models and inside the range of top flagship models. According to the AA model page cited in the report, Step 5 Preview supports text and image input and text output, has a 1M-token context window and 600B parameters, and runs at 99.8 tokens per second. Its cost per Intelligence Index task was $0.71, with input at $1.00 per million tokens, output at $2.70 per million tokens and a 95% cache discount. In Leiphone's test, Claude Code CLI called Blender and Three.js/Vite to build a playable 3D web game, Star Leap Island: Energy Delivery, with an original mushroom-headed character, six floating platforms, five energy orbs and a finish beacon. The report said about 70% of internal and external reviewers believed Step 5 Preview could autonomously complete medium-to-high complexity coding tasks. On ALE-CLI, FrontierFinance and DRACO, it ranked behind only GPT-6 Astra or Claude Opus 5 and ahead of other open-source models. StepFun's device reach was also cited: by the end of 2025, its phone-side models had more than 42 million installations, served nearly 20 million people daily and covered more than 50% of top domestic phone brands. StepAudio 3 Realtime ranked first on AA Conversational Dynamics at 98.9%, with speech reasoning accuracy of 99.7%, while StepAudio 3 ASR had a word error rate of 1.7%, tied for first.

AiShi Technology's PixVerse R2 has fully launched, allowing users to enter a web experience with action commands and interactive film-game interaction. Leiphone reported that AiShi proposed the product direction of a general real-time world model in January 2026 and launched R1. For R2, an analysis of HAR network records for the Winter Palace scene found that the client observed only one session_init and one generation_started. Eight recommended prompts shared one content stream, and prompt history accumulated continuously from the first to the eighth round. The generation window lasted about 119.5 seconds. The records also showed 3,355 structured action_frames, with a median sending interval of about 29.944 milliseconds, or about 33.4 Hz. Each prompt took about 3.9 to 4.5 seconds from sending to prompt_generated; in rounds three to eight, actions continued to enter the control channel about 8 to 26 milliseconds after a prompt was sent, with about 129 to 152 action frames in each window. The server-side prompt_updated median was about 1.33 seconds, while the full prompt_generated event median was about 4.27 seconds. Winter Palace started from a first frame and a 345-character world setting and used pipeline=r2. The interactive film-game Zero Mark used a director pipeline, maintaining an about 20-minute session with 53 valid state responses pointing to the same session, staying ACTIVE, and eight prepared 1080p, 24fps key cinematic assets. R2's scaling is described as occurring across model capacity, data, tasks, control signals and time scales, with Omni Causal AR raising the capability ceiling and Real-Time Acceleration bringing it back to low latency. PixVerse says it has 150 million global users and opened world model experience to ordinary users in January 2026.

HiDream.ai released HiDream-O1-Video-1.0, which it describes as a video generation model that better understands the real world. The model is placed in a route of one UiT and four models evolving from the same source, with front-end multimodal intent understanding, 5- to 20-second dynamic duration, 1080P output, physical change generation and native audio-video. Leiphone ran three original image-to-video tasks and compared the model with Seedance 2.5 and MiniMax H3 under the same reference images and original Chinese prompts, with a target duration of 10 seconds. In a wok-cooking task, HiDream's ladle re-entered the wok at 2.4 seconds, oil-popping sound peaked at 2.5 seconds and flame area began its fastest growth at 2.6 seconds, with action, sound and result about 100 milliseconds apart. Seedance showed wine pour at 0.25 seconds, popping at 0.54 seconds and fire at 0.58 seconds, while MiniMax showed fire at 1.5 seconds and the corresponding high-frequency popping at 2.26 seconds, putting flame about 0.5 seconds before sound and away from the wok center. Flame peaks were 54.7% of the frame for more than four seconds for Seedance, 36.2% for HiDream and 30.6% for about 0.6 seconds for MiniMax. In plating, HiDream's transfer lasted about 0.9 seconds with continuous flowing noodles, while Seedance took about 0.5 seconds with a clumped mass and MiniMax took about 0.42 seconds with a nearly rigid translation. In a drummer task, HiDream was the only one without a hard cut or sudden frame change, with maximum frame-to-frame change no more than five times the median. The snare, tom, kick and cymbal order held and strike positions, timing and timbre corresponded, although a cymbal sound appeared before the drummer moved. Its final hit around 4 seconds had a wide spectrum; stage lights warmed from 4.5 seconds and crossed the cool-warm boundary at 4.7 seconds over about 270 milliseconds. Cheering ran from 5.5 to 7.9 seconds, with more than 96% of energy in the 400-2000 Hz vocal band for three and a half seconds. Mid-frequency made up 47.8% and low frequency 9.3% of the soundtrack, the closest among the three to a real drum kit. Seedance's opening was well synchronized, but its final hit at 4.66 seconds missed the cymbals and was followed 50 milliseconds later by a hard cut with frame change 11.3 times the median. MiniMax offered the highest single-frame specification but had the most event-level problems, according to the report.

Manycore's Aholo Lux3D was tested with GPT-6 Astra in Codex and Blender to turn three photos of a bronze rat head from the Old Summer Palace into a 3D asset. GPT-6 Astra understood the task, split the roles of the reference images and scheduled Aholo Lux3D. Aholo Lux3D first used G1-Turbo to generate a fast GLB for contour validation, then G1 for a high-quality version. G1-Turbo took about 92.9 seconds and G1 about 559.7 seconds. The two calls cost 21 credits, about 1.47 yuan. The final rat-head model had about 190,500 vertices and more than 285,600 polygons, with height normalized to about 40 centimeters and a base about 3.5 centimeters high. The output included GLB files, a Blender project and three 1400 by 1400 render images. According to official comparisons cited in the report, Aholo Lux3D's standard version ranked first in PSNR, fourth in LPIPS and third in CMMD among several existing models. The fast version averages about 20 seconds per 3D model and has been used for quick game scene construction, embodied intelligence training and large-scale SKU screening. The report describes this as a 3D AI Harness, in which an agent schedules, a 3D model generates and professional software checks and post-processes assets.

Leiphone also tested AgnesCode with Agnes 3.0 Flash against Codex with GPT-5.6 Sol in the FlagOS Open Compute Global Competition Season 2 Track 1, which covers SGLang framework operator optimization across multiple chips. The task was Task66 dsv3_fused_a_gemm, a fused QKV-A down-projection operator used in DeepSeek-V3 decoding, tested on eight chips. Codex completed the task in 8 minutes 57 seconds, while AgnesCode took 20 minutes 36 seconds after a tester prompted it once. Official automated evaluation showed Codex passed seven chips and AgnesCode passed six. On the five chips both passed, AgnesCode had a higher speedup on four. On international general chip B, AgnesCode reached 4.81x versus Codex's 3.71x, about 29.6% higher. On Tianshu Zhixin, AgnesCode reached 2.16x versus 1.72x, and on Kunlunxin only AgnesCode passed, with a 2.56x speedup. Agnes 3.0 Flash scored 36 on the AA Intelligence Index, close to GPT-5.6 Luna(max)'s 38, and was priced at about $0.03 per million tokens, placing it on the Pareto frontier. The report said the free model still showed instability in delivery details such as entry naming, submission rules and coverage.