CoreWeave Expands AI Stack Beyond GPUs as Inference Demand Reshapes Token Economics
CoreWeave is expanding its AI stack beyond GPU compute as inference demand grows faster than training, according to theCUBE Research analysts.
Dave Vellante, chief analyst at theCUBE Research, said the mix of AI workloads is shifting. “Today they’re at 50/50, and they expect to be 10/90 by next year,” Vellante said during an exclusive broadcast on theCUBE, SiliconANGLE Media’s livestreaming studio. “Both curves are growing. If you look at CoreWeave’s numbers, they’re off the charts.” He spoke with Executive Analyst John Furrier.
The shift toward inference puts more emphasis on how efficiently an entire AI system can turn compute into usable output, favoring architectures optimized across silicon, networking, storage and software rather than around a single component, Furrier said. “If that shift happens, these data centers will look a lot different with the game still the same,” he said. “Pump out as many tokens per watt as possible. Get the intelligence shipping. You got the perfect storm on the supplier side with Dell, ecosystems booming and CoreWeave in pole position.”
Commercial terms are changing alongside the technical requirements. Vellante said CoreWeave customers are pushing for shorter contracts, spot pricing and on-demand access. Those structures carry higher prices than long-term take-or-pay agreements, suggesting flexibility itself is becoming part of the economics surrounding AI capacity. “The prices for those types of structures are much, much higher, and people are willing to pay,” Vellante said. “It’s going to be really interesting to see as that 98% starts to go down toward 50/50, how that’s going to sort of affect the market.”
GPU availability may bring customers to CoreWeave, but the company is broadening its stack to keep them there. Its strategy now includes networking, storage and a software layer built around observability, security and continuous improvement, Vellante said. “What you’re seeing CoreWeave do strategically is they’re expanding out beyond compute,” he said. “They’ve got networking; they’ve got storage. They announced [CoreWeave] Forge today … they’ve got this software layer, which is observability. They’ve got security in there. They’ve got this closed-loop system between evaluation, observation, runtime curating … that’s sort of how they’re saying their software layer is now intact.”
That broader systems approach also reinforces CoreWeave’s relationship with Nvidia Corp., according to the analysts. Nvidia silicon remains critical, but commercial advantage increasingly depends on how efficiently the full infrastructure stack turns those resources into useful tokens, Furrier said. “This is becoming increasingly a systems game; it’s a systems race,” he said. “Silicon matters enormously. What’s useful commercially is how quickly the entire system turns into useful tokens. At the end of the day, that’s the key.”
TheCUBE is a paid media partner for the Fully Connected event. Sponsors of theCUBE’s event coverage do not have editorial control over content on theCUBE or SiliconANGLE, according to the disclosure accompanying the interview.