AI News Feed
Market watch
Large Language Models

Doubao 2.1 Pro Updated to 0915 Build and Added to Doubao Work

Volcano Engine said Doubao 2.1 Pro has been updated to its 0915 build, with the new API live on Volcano Ark and the model available inside Doubao Work, adding multimodal reasoning, coding and general agent capabilities.

According to Volcano Engine, the new build improves multimodal understanding, including video reasoning and 3D object image recognition. Coding ability also improved substantially, letting the model handle cross-file, long-horizon problems inside complex code repositories. In multimodal coding scenarios where the model writes code from images, the latest Doubao 2.1 Pro has reached a leading level in China, the company said, serving delivery work such as front-end web design, game development and 3D modeling.

In one test case, the updated model coordinated multiple sub-agents working in parallel. On 1,000 real historical issues from the open-source game repository Luanti, which contains about 387,000 lines of code, it fixed 83% of the problems to a mergeable standard within nearly 36 hours. On a 3D modeling task, it read a courtyard design drawing covering the four seasons and generated, from scratch, a code-based 3D dynamic scene that could be demonstrated with controllable camera movement.

General agent abilities also improved, particularly in long-horizon scenarios such as financial investment research and office automation that require repeated tool calls and online research. Volcano Engine said the model strengthened evidence tracing, retrieval of authoritative sources, timeliness judgment and data verification, reducing hallucination and making complex long-horizon tasks more stable. In a task verifying financial report information, the new model retrieved content across languages on a broad scale, cross-checked different sources and produced a traceable due diligence report, with the research method and quality of the delivered report moving closer to expert level, according to the company.

Token efficiency was optimized as well, with fewer reasoning rounds and tool calls. Token consumption for image and video reasoning fell by more than 30% compared with the previous generation, which Volcano Engine said substantially lowers costs for enterprises and developers.