Nvidia and CoreWeave target CPU bottleneck in agentic AI infrastructure
Nvidia and CoreWeave are targeting CPU bottlenecks in agentic AI infrastructure, with CoreWeave planning standalone Vera CPU capacity and early Sandbox tests showing about 3x faster startup times, according to SiliconANGLE.
Agentic AI infrastructure must support the work between a model’s decisions as agents move from answering questions to executing tasks. Central processing units handle much of that execution, while graphics processing units power model reasoning. Those complementary roles shape Nvidia’s planned Vera CPU deployment with CoreWeave, said Hannah Coutand, director of product marketing for Vera CPU at Nvidia. Coutand said Nvidia’s approach to extreme co-design looks across the AI factory to find inefficient bottlenecks. “We recognized the CPU was becoming one of them,” she said. “We didn’t set out to say, ‘Let’s go and build CPUs.’ We set out to solve this bottleneck.”
As training and inference become increasingly intertwined, infrastructure must accommodate both model computation and the environments where agents act. Tool calls, application programming interface calls and SQL queries increase the execution work handled by CPUs, Coutand said. “I think the volume certainly plays a role,” she said. “That’s why having a CPU that handles those types of calls extremely well, with fast, beefy cores, high memory bandwidth and low latency, it handles both operating in this new world … and serves as a great foundation for simply agentic use cases [and] reinforcement learning.”
CoreWeave announced plans to offer standalone Vera CPU capacity alongside its existing infrastructure. Its Sandboxes offering provides isolated environments where agents can execute tasks during reinforcement learning and inference, said Harsh Banwait, senior director of product at CoreWeave. “Quite recently, we also tested that with Vera CPUs, and we were proud to share that we saw about a 3x improvement in performance in terms of Sandbox startup times,” he said. “For us, it’s a very important combination of getting the performance that we need from the silicon and getting our customers the experience that they expect from a product.”
Security also shapes agentic AI infrastructure. Nvidia’s Open Agent Safety Platform combines OpenShell runtime controls with Nvidia Sentry running on BlueField-4 data processing units to provide an independent layer of monitoring and enforcement. Coutand said the components sit in a single Vera Rubin tray that includes the Vera CPU as well as the DPU. In that tray, the Open Agent Safety Platform’s secure runtime, OpenShell, runs on the Vera CPU, while DOCA Sentry runs on the DPU part, she said.
Operating that infrastructure requires attention to power, cooling and automation. Networking, CPUs and storage also figure in CoreWeave’s engineering priorities, Banwait said. “All of that needs to be able to keep up,” he said. “We’re going to continue to focus on wherever the bottleneck is so that when the entire system is kind of advancing, it’s doing that in one cohesive way. Otherwise, the weakest link in the chain kind of holds it all back.” The interview was part of SiliconANGLE’s and theCUBE’s coverage of the Fully Connected event.