AI News Feed
Market watch
Large Language Models

Open-weight MiniMax H3 vs Seedance 2.0 Fast: Test Shows Promise and Gaps in Physics Simulation

Leiphone tested MiniMax H3 against Seedance 2.0 Fast. The open-weight H3 reconstructed a pottery-breaking event but lagged in speed, realism, and physics.

The test was inspired by a "Detective Conan" episode. The editors gave both models a screenshot of a shattered pottery jar and a storyboard prompt describing how a cat entered a studio, knocked the jar off a pedestal, and caused it to fall and break. H3 was run locally on an NVIDIA RTX A6000 GPU through ComfyUI, while Seedance 2.0 Fast was accessed via an API. Both models were asked to produce a 10-second video showing the cause-and-effect chain from the cat's entry to the jar's shattering.

H3's output showed a complete narrative: the studio, the intact jar, the cat entering through a high window, the jar wobbling and falling, and finally the fragments on the ground. The result preserved the reference image's key visual features, such as the brown clay material, light-colored fracture edges, and low-saturation animation style. Leiphone noted that H3 is not doing real physics inference but generating a visually and acoustically plausible path under multi-modal constraints. Its strengths lie in unified multimodal modeling, reference-image conditioning, and synchronized audio-visual generation.

However, in a direct comparison with Seedance 2.0 Fast, the closed-source model had clear advantages. Seedance completed the 10-second video in 3 minutes 23 seconds through the API, while H3 took 54 minutes 4 seconds on the local GPU. Leiphone cautioned that the time gap is not simply a model capability difference, as the local GPU was older and the inference environment lacked the engineering optimization of a commercial API. In terms of cost, Seedance charges 0.60 yuan per second, making a 10-second video cost 6 yuan, whereas the open-weight H3 has no per-second fee but incurs electricity, hardware depreciation, and waiting time.

More importantly, Seedance's output had higher completion and visual continuity. The closed-source model showed the key contact point clearly: the cat actually hitting the jar, the jar tipping, falling, and fragments sliding to a stop. The physics of the impact, trajectory, and fragment motion were more convincing and closer to a real-life accident scene. H3's output still appeared somewhat animated and generated; the moment of contact between the cat and jar was not clearly visible, and the audio at the beginning did not match the action. The open-weight model also failed to show the cat escaping through the window as prompted.

Leiphone concluded that the two models represent two different paths. The open-weight H3 provides a researchable starting point for localized deployment, multimodal audiovisual generation, and reverse reasoning from results to processes. The closed-source Seedance 2.0 Fast benefits from platform-level engineering optimization and shows maturity in realistic rendering and physical stability. The report noted that H3 is not fully open-sourced, as some modules are still only available through the official API due to system complexity. Therefore, the local open-weight version may not fully match the official demonstration quality. Ultimately, current video models can already expand a "result image" into a "process video," but reliable physical accident reconstruction still requires further improvements in action causality, force logic, fragment motion, and audio-visual synchronization.