Shengshu Technology Releases Vidu S2 for Real-Time AI Video Editing and Digital Human Interaction
Shengshu Technology has released Vidu S2, a real-time video model that combines interactive digital humans with live editing for streams, according to an APPSO review. It supports 720p output, dynamic reference images, and real-time clothing, background, style, and character changes.
S2-Avatar supports 720p output and improves action instruction following and performance. It can take a reference image during a live stream and follow commands to interact with objects, change clothing, or change the background. Voice or text can drive actions such as arm swings, twists, steps, and leg raises. Users can upload photos of real people, anime characters, or pets to create a digital character, then enable microphone and camera permissions to talk with it.
APPSO tested the avatar by creating a Steve Jobs digital human from a full-body reference image and a character description covering Jobs's personality, expression, and views on products, design, and user experience. When asked whether the iPhone Duo met his expectations, the avatar replied, "The iPhone Duo is smoother than I expected, but the price makes people hesitate." APPSO also asked the avatar to dance, testing full-body motion control. The review said the result was more than animating a photo: the system tries to turn a fixed character image into a digital role that can hold objects, evaluate products, and perform complex actions.
APPSO also held video calls with an avatar named Himari and a 30-something American delivery driver named Tom Miller. Himari said she was wearing a school uniform; Tom Miller said he was busy with deliveries. To address collapse and drift in long streaming generation, Shengshu developed a training technique called SRF, or Self-Replay Forcing. The review described it as giving the model self-correcting memory so it can review its own generation history and fix errors.
Vidu S2-Editing's main upgrade is that a single reference image can edit a playing video stream in real time. The functions include virtual try-on, background replacement, style change, and character replacement. APPSO uploaded clothing reference images with clear garment edges and got try-on results for a blue polo shirt, a beige coat, a black leather jacket, and a baseball jacket. The model showed fluid motion, stable facial features, and realistic fabric deformation. The review said live commerce could use the feature: operators upload white-background product images, and a digital host changes outfits or presents products while speaking, turning live streaming from one-way broadcasting into responsive shopping guidance.
S2-Editing can also be used for real-time special effects and video packaging. The review said traditional stylized rendering and background keying depend heavily on post-production work, with timelines measured in days. With Vidu S2-Editing, a camera feed and different style choices can give a live stream different visual qualities. Character replacement can switch a person to another image in real time; APPSO recommends a front-facing, clearly visible half-body reference image. In tests, the review said the system achieved precise perspective alignment and natural occlusion, with few collage artifacts. Backgrounds shifted from a beach to a cafe with cluttered backgrounds removed in real time.
The system is aimed at interactive entertainment. Vidu S1 already moved in this direction, and S2 adds dynamic reference images, complex motion, and real-time editing while raising digital human output to 720p. The review noted that continuous video generation is harder than adding a chat box to video: a raised hand cannot pause while the model computes, and a real-time system cannot know the next sentence in advance. Shengshu's full-chain development for this system is led by Zhang Jintao, described as a post-2000 PhD student of Zhu Jun and the company's head of streaming video generation and inference. Vidu S2 uses a two-stage Backbone-Refiner architecture. The backbone network handles skeleton motion, physical rules, and camera language at very low latency; the refiner network works asynchronously to reconstruct 720p details, textures, and lighting in milliseconds. Temporarily added reference images are coordinated by a VLM Agent, which parses the motion phase and spatial state of the current frame and aligns features at millisecond timestamps so new elements follow inertia and lighting physics. The review said Vidu S2 also uses reinforcement learning for causal streaming generation at the deployment end, reducing instruction drift.
The review frames the shift as a change in the basic unit of video. For more than a century, video has been a completed span of time that is shot, edited, compressed, distributed, or generated by AI, with the audience arriving after the content is finished. Real-time generation lets a viewer's intent enter while the image continues. An analysis cited in the review said that by the end of 2026, creators may no longer wait for render queues, AI systems may support real-time interaction with the image itself, and advertising could move from one video for a million viewers to a million videos each made for one person. Bloomberg reported that ByteDance is also preparing an AI model that would let users create real-time spatial video, with Zhang Yiming personally overseeing development and the model possibly launching as soon as next month. Shengshu's Vidu S series focuses on generation, extending video from one-time output into an interactive process that can continually receive input, understand instructions, and respond immediately. Vidu S2 extends this to real-time spatial video generation and editing: generated or edited frames can be converted into synchronized left-eye and right-eye views for compatible headsets to present stereoscopic depth, and existing stereoscopic video can also receive real-time edits.
Zhang Jintao cited two Vidu S1 community cases in a podcast. One user uploaded a photo of their dog and chatted with a digital dog on screen. Another player used game footage as input so an anime character could watch the game and talk with them, reacting to changes in the image. Zhang said that in the future a large amount of visual content may be generated in real time by AI, and each person may have an independent interactive session. The review argued that real-time interaction is a basic human visual entertainment need, noting that before the digital revolution most visual entertainment was real-time and online. Offline pre-made video has a playback ceiling, but real-time interactive video gives each person a unique session. Adobe's MotionStream lets creators drag the camera and change object motion trajectories with a mouse and sliders while video is being generated. HeyGen's LiveAvatar focuses on real-time interactive AI digital humans that can be customized with a user's own data and has been used by teams for customer service, live streaming, and enterprise communication.