AI News Feed
Market watch
AI Chips & Compute

At AI Infra Summit, Tech Executives Push Faster AI Infrastructure as Memory Wall Looms

AI infrastructure firms detailed faster designs at the AI Infra Summit, targeting inference, memory and token efficiency.

Qualcomm Executive Vice President and General Manager of Datacenter and AI Tony Pialis said tokens per watt has become the new key metric in what he called the AI war. "We clearly cannot stay on this trend," Pialis said during a presentation. "We on the infrastructure side need to do better."

The exponential demand for AI has generated a diversity of compute in a short period. SiliconANGLE reported that training AI models required significant GPU power for much of 2025, while a focus on inference this year has led to skyrocketing token usage and demand for more heterogeneous infrastructure that includes central processing units. AWS Senior Vice President Peter DeSantis said the majority of compute is going to serve inference and called inference a massive workload.

AWS has built its own CPU portfolio over the years, and its Graviton family of 64-bit Arm-based CPUs is central to its infrastructure strategy. In June, AWS launched the Graviton5 CPU to support real-time AI reasoning and multistep task orchestration. DeSantis said the vast majority of workloads on AWS run cost effectively and faster on Graviton, adding that AI could unlock more specialization in software and a wave of general-purpose workloads.

Memory has become another critical constraint. SiliconANGLE analysts have documented that a typical AI server uses roughly eight times more memory than a traditional server, and AI server memory spending is projected to jump from $35 billion in 2025 to between $175 billion and $190 billion by 2027, about a fivefold increase. The growth has produced divergent views about how to meet AI's memory needs.

Qualcomm has shifted its focus from traditional high-bandwidth memory, or HBM, to a proprietary architecture it calls high-bandwidth compute, or HBC. Qualcomm estimates that HBC provides six times the bandwidth per watt versus HBM for large batch sizes and 200 times the capacity per watt versus static random-access memory, or SRAM. Pialis said HBC can overcome the memory wall, where compute has exceeded both memory and bandwidth. "The bottleneck is data movement, not arithmetic," he said. "The way to solve that is to bring the compute even closer to the memory. We've effectively moved the data into the same condo where the compute is. Everybody now has the same elevator, up and down. HBC delivers the benefit that the industry needs."

D-Matrix has pursued a different route based on dynamic random-access memory, or DRAM. The computing startup integrated higher-throughput 3D DRAM into its next-generation chip architecture, Raptor. The design stacks multiple layers of memory cells vertically, allowing higher storage density and improved performance. D-Matrix founder and CEO Sid Sheth said the company is much better than SRAM and does much better than HBM. "We've been preparing for a world of infinite inference for a long time," Sheth said in his AI Infra presentation.

Last week, d-Matrix unveiled a collaboration with Nvidia to incorporate Raptor into Nvidia's rack reference architecture, NVLink Fusion. The solution is designed for AI labs, hyperscalers and neoclouds to deploy ultra-low-latency premium-level token services. Sheth said his company decided to ride on that infrastructure.

Oracle, Broadcom and other companies also participated in the AI Infra Summit, where the common message was that infrastructure providers are under pressure to improve performance per watt and manage memory costs as AI adoption expands. The conference took place about an hour's drive south of Dreamforce, where the debate over AI's pace continued.