Alibaba Launches Enterprise Agent Infrastructure and Multimodal Models; Intel Says CPU Value Is Rising
On Sept. 22, 2026, Alibaba's Qwen Office released Enterprise Context and QwenNote A2 for enterprise agents, while Alibaba Cloud unveiled a full multimodal model lineup. Intel separately said AI agents are shifting data-center compute ratios toward one CPU per GPU.
Qwen Office's Enterprise Context aims to compress and structure enterprise information scattered across group chats, documents, knowledge bases, and business systems into knowledge agents can call as needed, according to Leiphone. The system updates continuously with business changes and sends only the most relevant slice for a task, using a built-in dedicated model to reduce token costs. Qwen Office CEO Chen Yusen said a mid-to-large enterprise may have about 150PB of internally registered data, while the strongest context windows are at the million-token level.
QwenNote A2 is an upgrade of DingTalk's 2025 recording card A1 and is positioned as an agent entry point rather than only a recorder. It has six microphones, picks up sound up to eight meters away, claims 98 percent transcription accuracy, and recognizes 120 languages and 41 dialects. It includes a 4G IoT card, can work without a phone, supports task assignment via a long-press AI key, and offers real-time voice chat. It can transcribe for 40 hours continuously, standby for up to 18 days, and retains only text and summaries, not raw audio.
Chen described a three-step approach: connect, understand, reuse. Enterprise knowledge is first connected across systems, including images and video through multimodal models; then structured into callable entities such as people, projects, and process rules; then retrieved on demand for specific tasks. Permissions inherit the enterprise's IM organization, identity, and access-control system. Guming Tea's operations knowledge space for nearly 10,000 frontline employees was cited as an example: recipes, equipment checks, and opening and closing procedures are split into entities and relations, with different permissions for store managers, trainers, and area managers.
Alibaba Cloud's Yunqi updates included Qwen3.8-Max, which is described as starting to train itself; Qwen3.8-Flash, which previews the next-generation architecture; Qwen3.8-Omni, which adds a new thinking partition; and Qwen4, already in training, with Qwen4.5 and Qwen5 on the road map. The company is targeting parameter scales of 5 trillion to 10 trillion. It also showed Qwen-Image-3.1 for image generation, Happy Shrimp 1.1 for music, HappyOyster 2.0 Preview for world models, Qwen-Audio-3.1 and Qwen3.8-LiveTranslate for speech and simultaneous translation. A next-generation video model is planned for November.
QbitAI reported that Wan3.0 supports single-generation clips of up to 30 seconds and accepts structured inputs such as documents, tables, and web pages; Alibaba said it ranked first in Artificial Analysis's text-to-video and video-editing lists. Director Lu Chuan used AI to rebuild a 400-year-old Ming dynasty disaster scene in five days, a project that would traditionally require months and tens of millions of yuan; the team optimized prompts about four rounds, and Alibaba's video model reached more than 95 percent accuracy for smoke, debris, and other complex images, with Alibaba's model contributing more than half in the video generation stage. Actress Wang Luodan described staying up until 4 a.m. drawing AI-generated shots and finally getting the desired shot at 6 a.m., saying models still struggle to maintain a character's state, emotion, and temperament across a longer story.
Alibaba ATH business group vice president and Taotian Group chief scientist Zheng Bo said text, image, video, sound, and spatial capabilities are accelerating in coordination, and AI is moving from generating a single image, video, or music piece toward understanding intent and creating full immersive audiovisual experiences. He predicted a native full-modal unified generation model within three years, with future experiences no longer limited by modal boundaries. Alibaba also described a direction called Agentic Video, in which models understand creative goals, break down plots, make storyboards, organize characters and scenes, and complete generation from an idea or story. The next-generation video model will move from generating single shots toward understanding complete narratives.
At Intel Connection in Suzhou, Intel said agents are increasing the need for CPUs to handle scheduling, data processing, and execution. Intel data center group vice president and China general manager Chen Baoli used a news-summarization agent as an example: the agent must find web pages, PDFs, or videos, extract and process different formats, call models multiple times, and decide next steps. He described data centers as three cooperating parts: CPU servers or clusters for agent execution and orchestration, GPUs for model computation, and storage for task data, conversations, and generated documents. CPU work runs through all three. Chinese Academy of Engineering foreign academician and Tsinghua AIR founding dean Zhang Yaquin said agents can plan, act, trial, and iterate around a goal until a task is complete.
Chen said that in past months customers have asked both how many agents can run and how capable and reliable each agent is, especially when many sandboxes or agents run simultaneously and must meet response times. He introduced KV Cache optimization, including lossless compression using built-in processor capabilities and faster data transfer, as well as work with software partners on high-performance storage to reduce CPU and GPU idle time. Intel CEO Lip-Bu Tan said in a video speech that x86 remains a foundation of modern computing. QbitAI reported that Intel's stock rose 12 percent the previous day, which it said reflected market recognition of CPUs' importance in the agent era.
Intel vice president and China general manager for client computing and physical AI Gao Song described CPU, GPU, and NPU collaboration on personal devices: CPU handles planning and decisions requiring timely response, GPU handles matrix computation, and NPU supports long-running audio-video and computer vision tasks at lower power. He cited home AI, automotive AI, AI running on private data, and robots as device categories with different needs. Peking University's New Structural Economics Institute director Lin Yifu said AI's economic impact would be felt through industries using AI to raise efficiency, which for Intel means turning application demand into products and services customers will pay for.
The day's announcements point to two fronts in the agent race: Alibaba is trying to supply enterprise context and multimodal models so agents can act inside organizations, while Intel is arguing that agent workloads require more CPU-side execution, scheduling, and data movement alongside GPUs. Qwen Office said its enterprise agent platform connects with DingTalk, Feishu, and WeCom, and cited DingTalk's 800 million users, 27 million enterprise organizations, and Alibaba Cloud's 5 million customers as its moat. Intel said the exact CPU-to-GPU ratio still depends on tasks and system configurations.