Intel Says CPU Becomes Control Hub as AI Shifts From Training to Agents
Intel said at its Sept. 22 Connection event that agentic AI is making CPUs the control hub, as it pushes Xeon 6+ and third-generation Core Ultra for digital and physical AI.
The distinction matters because training and inference agents put different demands on hardware. Training is a standardized, parallel batch task focused on maximizing overall throughput. An agent is a serial state machine: each step depends on the output of the previous step, so small delays in any part of the chain are amplified along the task path. Stable, controllable, very low end-to-end latency therefore becomes a hard constraint for deploying agents at scale. Hardware loads also conflict. High-concurrency tool calls and high-bandwidth AI inference need to run together on a single chip, and their memory demands pull in different directions. Multi-tenant security isolation and dedicated resource needs further expose the limits of traditional CPU architecture, making architectural change and system redesign necessary.
Intel's answer is not to stack up individual parameters but to divide data center architecture at the system level. Chen Baoli, vice president of Intel's Data Center Group and general manager for China, said at the event that future data centers will be split into three types of clusters: CPU clusters for agent execution and task orchestration, GPU clusters for model inference, and storage clusters for continuously growing business data such as long conversations and documents. All data scheduling among these clusters will be handled by the CPU, he said.
On the product side, Intel introduced Xeon 6+ processors built on the Intel 18A process, with up to 288 efficiency cores. Paired with an agent suite, a single processor can deploy nearly 1,000 agents, according to the company. Intel said the claim is supported by test data: in SWE-bench, single-agent performance leads a competitor by 15 percent. In a cloud sandbox scenario that requires starting tens of thousands of sandboxes in 10 minutes with single-sandbox latency under 200 milliseconds, Xeon 6+ supports 350 concurrent sandboxes, about 1.46 times the competitor's performance, Intel said.
Intel also highlighted three specific areas where CPU value stands out. For heterogeneous data preprocessing, agents need to crawl web pages, video, and PDFs and extract text, tasks GPUs are not good at; Intel's AMX accelerator makes vectorization and reranking reach 2 times and 4 times the baseline, respectively. For storage scheduling, cost and latency move in opposite directions from HBM to DRAM to local SSD. As the core scheduler, CPU works with Yanrong Technology on a Xeon 6 and MRDIMM-based solution that took two global firsts in MLPerf Storage v3.0 for Llama3 checkpoint read/write and 3D-Unet training, eliminating storage bottlenecks, according to the report. For KV Cache management, a million-token context takes about 300GB. Intel's KV Cache intelligent acceleration suite uses QAT for lossless compression, reducing 300GB of raw data to 200GB, and uses the DSA engine to accelerate data movement by up to 9.66 times.
The report described the conclusion as follows: in the agent era, GPU carries the core inference computation for intelligence, while CPU must handle basic computing and take on global management. As business scale expands, the complexity of this scheduling and control work rises exponentially, which Intel presents as an area of long-term accumulated advantage.
For physical AI, Intel Vice President and General Manager of Client Computing and Physical AI for China Gao Song said at the Connection event that Intel hopes when people mention physical AI, the first company they think of is Intel. The harder problem is execution: making a robot dance is easy, but precise operation is difficult. The bottleneck is control, because real-time perception and precise execution in dynamic environments remain a major challenge for robots.
Third-generation Core Ultra addresses this with what Intel calls fusion of the big brain and the small brain. Using an XPU heterogeneous architecture, inference decisions and real-time motion execution are integrated on a single SoC. The CPU acts as the cerebellum, completing 4,000 motion adjustments per second at a 4kHz control frequency to ensure precise response from joint motors. The GPU handles high-level AI inference, with latency of 78 milliseconds at Pi0.5 precision, below the average 200 to 250 milliseconds of human visual reaction. The NPU is for visual perception, with YOLOv12n inference taking only 3.55 milliseconds and processing about 280 frames per second, far above the 60-frame output limit of ordinary cameras. Traditional designs need multiple independent units for perception, decision, and control, which are slower and prone to coordination errors. With a single chip, engineers can shift effort from low-level tuning to upper-layer intelligence development.
Cost is another pressure point. Token consumption for embodied intelligence grows linearly with user numbers. A founder of an embodied company said an engineer can burn hundreds of dollars in an afternoon of training. Intel's approach is to reduce underlying overhead through hardware load integration, unifying inference decisions and motion control on a single SoC's heterogeneous architecture, compressing compute redundancy and energy consumption from the hardware level, and bringing total cost of ownership into a commercially viable range.
In algorithm routes, the embodied field has no unified paradigm. World models, VLA, and other approaches are still evolving in parallel. Chip architecture design cycles last two to three years, so betting on the wrong route is extremely costly. Intel's differentiated strategy is not to follow a single algorithm route blindly but to use an open architecture and ecosystem to support parallel validation and rapid iteration of multiple algorithm paradigms. It has partnered with KeTong Industrial, Digital China, and others to launch an embodied intelligence development kit based on the third-generation Core Ultra platform, and it is working with partners on projects from prototypes to mass production in manufacturing, construction, medical, smart city, warehousing logistics, and hotels. Intel also announced that its embodied kit will launch on JD.com in November.
According to the Leiphone report, the shift means system design is moving toward the execution side. On digital AI, Intel's judgment is that inference and agent deployment expand CPU-side system work, and x86's existing advantages in control plane, ecosystem adaptation, and infrastructure software can meet that demand. On physical AI, it uses heterogeneous SoCs to shorten the perception-decision-control loop and relies on ecosystem breadth to hedge against algorithm uncertainty. The report said that not all compute may be used, but providing an optimal system-level choice may be the key to productivity gains as AI begins to solve real-world problems.