Shengshu Technology Opens Vidu Q4 Preview, Pushing AI Video Generation to 0.09 Yuan per Second
Shengshu Technology has released a preview of its flagship Vidu Q4 video model, which QbitAI tested at as little as 0.09 yuan per second in 720p. The build adds 4K output, up to 15 reference images and three reference audio clips for character consistency.
The preview is the first Q4 build released to creators and concentrates on three areas that Shengshu says are where AI video most often gives itself away: more convincing character performance, more dynamic camera language and more forceful visual effects.
According to QbitAI, pricing depends on the specification chosen rather than being a flat rate. On the MaaS side, image-to-video costs about 0.6 yuan per second at 720p and 0.75 yuan per second at 1080p under a tiered system. Through the Vidu SaaS product, a two-month limited-time promotion during the preview period brings parts of the usage cost down to several cents per second, the report said.
The model outputs up to 4K and can generate 2K and 4K footage in volume, whether through the API or directly inside the Vidu product. A single generation can take up to 15 reference images, so a character, a costume, a prop and a scene can each be supplied separately and their relationships described in text, along with as many as three reference audio clips that fix a character's voice from one scene and emotion to the next. Clips run to a maximum of 16 seconds, which QbitAI said leaves room for continuous action, camera moves and first-person FPV shots of the kind flown by racing drones.
In its tests, QbitAI uploaded images and audio to the reference-to-video page and slotted them into a prompt. The outlet generated a shot-reverse-shot English-language dialogue scene, a comic-style fight sequence submitted by a Vidu user identified as Xinghuo 279, a science-fiction scene of a megastructure suspended above a sea of clouds, and a coffee commercial. Six clips totalling about 110 seconds came to less than 10 yuan at 720p pricing, it said.
QbitAI singled out the lighting in the science-fiction clip. When the structure lit up, an astronaut's white suit, the surrounding railings and the clouds below took on a warm orange cast, and the image returned to a cooler tone as the light dimmed. The effect, the report said, changes the surrounding scene rather than sitting on top of it as a filter.
On performance, the preview strengthens the detail of facial expressions, the transitions between emotions and body movement. Voice is treated as part of the acting: with reference audio, a character keeps the same timbre across scenes while still speaking with emotion.
Camera work is the second focus. QbitAI said the model holds the relationship between a person, a subject and the space around them more steadily in tracking, orbiting, push-pull and crane shots, and can cut between wide, medium and close views within one generation.
The third area is effects and texture. Generating an explosion is not difficult, the report noted; the harder part is what follows it, including whether firelight illuminates its surroundings, whether smoke and particles disperse along the direction of motion, and whether an actor's movements keep pace with the effect. QbitAI said Vidu Q4 was built to let effects change together with people, environment and camera.
Shengshu Technology said Vidu's research team includes a number of young engineers who track games, anime, film and visual effects closely, and QbitAI attributed the model's distinctive look to that technical and creative taste being trained into it. The preview already handles both high-energy action and quieter dramatic beats, the outlet wrote, and a formal release is still to come.