AI News Feed
Market watch
AI Chips & Compute

Nvidia Exec Says AI Factory Economics Hinge on Tokens and Power Efficiency

Nvidia's Ian Buck says AI factory economics increasingly hinge on tokens per watt, system design and inference, not just GPUs.

Buck made the comments at the Fully Connected event in an interview with theCUBE Research's Dave Vellante and John Furrier, broadcast by SiliconANGLE Media's theCUBE. He said networking, storage, processors and software must operate together at scale while increasing the output generated from every unit of power.

"Instead of cars or devices or PCs, it's tokens," Buck said. "These assets are not IT; they're not cost. They're actually appreciating, revenue-generating, fungible, durable, productive parts of an economy."

Buck described inference as the source of an AI factory's commercial output, with deployed models processing requests and producing tokens. He said inference does not replace training because organizations continually update deployed models as data and market conditions change.

"It's not just fire and forget on all these services," he said. "As companies are using these models, they're refining them, they're aligning them, they're adding more data to them. Having them up to date and aware — that actually is a little bit of training. We're seeing the work in reinforcement learning and online alignment."

He also pointed to low latency as creating another economic tier for workloads where faster reasoning has greater value. Nvidia's Groq 3 LPX inference accelerator works with its Vera Rubin platform to increase per-user token rates for time-sensitive workloads, according to Buck.

"If there's value in those tokens to have the fastest possible thinking, LPX can be boosted on top of Vera Rubin to make that possible," he said. "We're seeing a lot of interest in areas like fintech and other areas where things are happening in real time."

Power capacity ultimately limits how much computing infrastructure a data center can deploy, Buck said. That constraint makes tokens per watt a central measure of AI factory economics and puts pressure on vendors to improve performance with each hardware generation.

"Data centers have a natural cap, and that cap is actually their power," Buck said. "With every generation of GPU, we make sure that our tokens per watt is upwards of 10 times more efficient. In fact, we saw that with Blackwell — we got, in the end, a 30x improvement in tokens per watt."

Buck said the shift changes the scale at which systems must be designed and operated. CoreWeave allows customers to choose configurations or use higher-level inference services that optimize the balance between throughput and token speed, he said.

"CoreWeave can do that for customers," he said. "They don't have to feel overwhelmed by all the choices. That's where our partner ecosystem is so important."

The interview was part of SiliconANGLE's and theCUBE's coverage of the Fully Connected event. TheCUBE is a paid media partner for the event; the article states that CoreWeave, the sponsor of theCUBE's coverage, and other sponsors do not have editorial control over content on theCUBE or SiliconANGLE.