AI News Feed
Market watch
AI Chips & Compute

Alibaba Unveils Zhenwu V900 AI Chip and 20GW Cloud Plan as Qwen4 Enters Training

Alibaba introduced the Zhenwu V900 AI chip, a plan for more than 20GW of global data center capacity by 2032, and said Qwen4 is training, with future Qwen models targeting 5-10 trillion parameters.

The Zhenwu V900 delivers three times the performance of the Zhenwu M890 released in May, according to Alibaba. The new chip is scheduled for mass production and commercial release in the first quarter of 2027. Existing Zhenwu chips are already used by more than 650 customers in automotive, finance, energy and manufacturing, the company said. Based on the V900, a single cluster can scale to 500,000 cards, a capacity intended to support training and inference of frontier models. Alibaba also said its self-developed M890 AI supernode has supported efficient inference for models with more than 2 trillion parameters and began scaling the supernode on Alibaba Cloud this quarter. The announcements followed Huawei's unveiling last week of new AI infrastructure that it said can scale to as many as one million processors.

Alibaba also detailed progress across its Qwen model family. Qwen4 is already in training on a new architecture, while Qwen4.5 and Qwen5 are planned to expand to 5 trillion to 10 trillion parameters, according to QbitAI. Alibaba said Qwen3.8-Max scored 45 on Artificial Analysis, up from 40, placing it alongside models from Claude and GPT at the frontier. The company described progress in recursive self-improvement, or RSI, in which models design experiments, build training data and locate defects with no human involvement over more than a month, completing 33 effective iterations. In one example, Qwen3.8 adapted and optimized the SGLang inference framework for a next-generation T-Head GPU it had not seen before, raising single-instance throughput by 96%. In another, it ran for more than 60 hours and called EDA tools more than 10,000 times to complete a physical implementation optimization that reduced area by 42%, cut standard cell count by 29% and lowered power consumption by 59.5%, according to the company. Alibaba said it has open-sourced more than 460 Qwen models, with total downloads exceeding 3 billion and more than 300,000 derivative models. Qwen3.8-27B became the most-liked model in Hugging Face history, surpassing DeepSeek-R1 and Meta-Llama3.1, the company said. For devices, Alibaba released Qwen Intelligence for mobile partners and said its open-source Qwen27B model can run on desktops for local deployment.

On multimodal and generative models, Alibaba introduced Qwen-Image-3.1 for e-commerce, design and other commercial scenarios, saying it can compress visual design work from hours to seconds. Qwen-Audio-3.1 covers speech recognition, speech synthesis and real-time interaction, while new next-generation audio models include Qwen-Audio-3.1-TTS-Next and Qwen-Audio-3.1-ASR-Next. Wan3.0 ranked first on Artificial Analysis for text-to-video and video editing, and Alibaba said a next-generation video model will be released in November with stronger consistency, control and narrative understanding. HappyOyster 2.0 Preview was also shown, along with Qwen3.8-LiveTranslate, a simultaneous interpretation model that cuts average latency to about 2.3 seconds, compared with an average of about 4 seconds for human interpreters, according to the company.

At the same event, Alibaba introduced a new series of Qwen AI hardware: the Qwen AI glasses N1 and N1 Pro, and Qwen AI ear-clip earphones. The products opened online reservations on Sept. 22 and will go on sale on Oct. 13, according to Leiphone. The N1 series uses a 50-megapixel sensor and supports first-person shooting. The N1 Pro adds eye tracking and iris payment, allowing the AI to use gaze information to identify what a user is asking about. The ear-clip earphones combine Bose tuning with AI assistant functions, including face-to-face translation, AI note-taking and voice-driven tasks. Song Gang, general manager of Qwen AI hardware, said AI wearables are an important carrier for Qwen and a core part of its Personal Agent strategy, moving AI from waiting for commands to proactive service.

Alibaba CEO Eddie Wu told the conference that machine thinking is becoming the main source of thought and that intelligence is becoming a scalable commodity. He said total machine thinking would eventually exceed human thinking by more than 1,000 times, compared with less than 3% today, and that machines could handle 99.9% of thinking work in the future. Wu compared today's AI coding to the light bulb in 1882, saying it replaces existing work but has not yet created a new era. He identified AI models, AI chips and AI cloud as the three foundations of the machine intelligence era and said Alibaba would keep investing in them as a long-term strategic choice. The company said medium- and long-term AI demand far exceeds supply, and that a global shortage in AI data center supply chains is limiting how fast it can add computing capacity.

At a separate event the same day, Inspur Information used the AI Computing Conference 2026 to argue that AI compute gaps cannot be solved by adding cards alone. According to QbitAI, IDC's 2026 China AI computing power development assessment report said global AI compute demand satisfaction would fall from 79% in 2024 to 71% in 2027 and recover only to about 77% by 2030, leaving an absolute compute gap of $380.9 billion by 2030, nearly ten times the 2024 level. IDC also projected that global annual AI inference tasks would grow by 4,000 trillion by 2030, with token consumption growing at a compound annual rate of 4,822.6% from this year to 2030. Inspur launched the Yuanbrain SD200 Ultra supernode AI server, which packs 128 chips into one system with 8TB of memory and 64TB of extended storage, supports 10-trillion-parameter models on a single machine, and cuts All-to-All communication latency to 0.69 microseconds, according to QbitAI. It also introduced the Yuanbrain HC2000 multi-compute rack, featuring a liquid-cooled design with more than 300 kilowatts of power per rack, support for up to 256 AI accelerator cards and 200Tbps aggregate bandwidth, and said the system achieved a tenfold increase in token capacity under equal investment in a typical Agent task scenario.