Alibaba Lays Out Next-Generation Model Roadmap at Yunqi Conference
Alibaba said at its annual Yunqi Conference that future Qwen models will scale toward 5 trillion to 10 trillion parameters, while new multimodal and AI-for-chip-design efforts aim to speed development and lower costs. Leiphone reported.
Qwen large language model head Liu Dayiheng used his keynote to answer the most immediate question: how large the next models will be. After Qwen3.8 was released in August 2026 with 2.4 trillion parameters, he presented a roadmap in which Qwen4 will be followed by Qwen4.5 and Qwen5, with plans to reach 5 trillion to 10 trillion parameters. Two years earlier, Qwen2.5's flagship model had 72 billion parameters; by Qwen3.8 the figure had grown 33 times. Liu reiterated Alibaba's view that the intelligence ceiling of foundation models still depends on continued investment in compute, data and model scale, and said the company will keep expanding pretraining compute along the scaling law.
Alibaba is pairing scale with architecture changes to control costs. Qwen3.8-Flash, the company said, uses sparse attention, linear attention and external memory to activate only 6 billion parameters per inference. Compared with the previous generation, training cost fell to one-ninth, the input price per million tokens fell to one-quarter, and prefill throughput for very long context rose 8.6 times, while model capability continued to improve. The approach reflects Alibaba's plan to support larger models with lower training and inference bills rather than simply making models bigger.
The second question is how soon Qwen4 and Qwen5 will arrive. Model competition has become fast enough that a leading version can be matched within weeks or months, making iteration speed a more durable advantage. Alibaba and other frontier labs are exploring recursive self-improvement, in which models help design experiments, analyze results and improve training methods. Leiphone cited Anthropic's experiment in which Claude was given code for training a small model and asked to speed it up: it reached 3 times faster last year and 52 times faster this year, while human researchers typically achieved about 4 times in half a day. OpenAI trained a GPT-Red model to attack other models and used the samples to train next-generation GPT; the result, according to the report, was that GPT-5.6 failed about six times less often on the hardest prompt-injection attacks than a model from four months earlier.
Alibaba's experiments extend that idea into chips and infrastructure. Leiphone reported that Qwen3.8-Max was given a real network-on-chip module specification with no reference code or ready-made test cases. Over more than 60 hours, the model called EDA tools on its own and completed circuit description, verification and back-end physical implementation without round-by-round human intervention. The final chip physical area fell 42 percent and estimated power consumption dropped 59.5 percent. For a new T-Head GPU that had never been adapted before, Qwen3.8-Max modified and optimized the SGLang inference framework and completed deployment in 56 hours, raising single-instance throughput in daily conversation scenarios by 96 percent. Alibaba also disclosed that Qwen3.8-Max ran autonomously for more than a month and completed 33 valid experiments, building processes, constructing data, designing experiments, locating defects and pushing model updates.
The report placed those efforts in the context of China's domestic chip ecosystem. CUDA has long bound models, operators, frameworks and toolchains to Nvidia chips, and moving a domestic chip from "able to run" to "runs well" often requires lengthy adaptation. Liang Wenfeng has said domestic hardware must first solve ecosystem problems and pointed to higher-level programming approaches such as TileLang, which let developers reimplement low-level operators with less code instead of rewriting complex CUDA line by line. DeepSeek has begun exploring the use of AI to write TileLang, according to Leiphone.
Multimodal models are the third focus. Zheng Bo, the ATH technology vice president responsible for audiovisual and other multimodal models, said the next-generation video generation model will be unveiled in November. He described the next stage as "Agent Video," moving beyond generating realistic imagery or using reference images, video and audio for control. In his description, a user gives an idea, and the model breaks down the plot, designs storyboards and arranges characters and scenes to produce a complete video. The new version will pursue longer, more controllable, more complete and more intelligent generation.
Alibaba also introduced upgrades across image, speech, music and world models. Qwen-Image-3.1 is optimized for real commercial scenarios such as e-commerce, creative work and design, compressing visual design that once took hours into seconds and supporting precise image editing, the company said. Qwen-Audio-3.1 series upgraded automatic speech recognition, text-to-speech and realtime voice interaction; the realtime model can listen, think and speak simultaneously, with stronger multilingual interaction, role-play and empathy. The models have been used in Qwen Office and Qoder and integrated into hardware including the QwenNote A2 companion assistant, QwenNote Eva desktop robot and Qwen AI glasses. Music model HappyShrimp 1.1 was released with improvements in musical expression, vocal quality, multilingual singing and instruction understanding, supporting more than 100 musical styles and more than 10 mainstream languages for both vocal songs and instrumental music. World model HappyOyster 2.0 Preview more accurately simulates object motion and collisions, uses 3D spatiotemporal memory to keep scenes stable and reduces interaction latency. Alibaba is working with universities and industry partners on a world model benchmark and arena.
Zheng Bo said that within three years the industry could see a native omni-modal unified generation model that can generate images, video and music at the same time and understand space and the world. Across the conference, Alibaba framed its model strategy around three questions: how large the next Qwen will be, how fast model iteration can become, and whether multimodal models can compete. The company described two paths—one chasing the frontier of AGI and one bringing AI into everyday productivity—that converge on the same goal of moving higher intelligence out of the lab and into the real world.