Nebius CEO: AI Cloud Providers Shift from GPU Sales to Full-Stack Infrastructure
Nebius CEO says AI cloud moves from GPU sales to full-stack infrastructure, using large bare-metal contracts as financing tools.
A year earlier, Volozh was still answering who would buy so many GPUs; a year later, he said the real constraint is how much compute humanity can build. Leiphone compared his two interviews in “Spotlight On” in 2025 and 2026, seeing a mindset shift from GPU infrastructure supply to a next-generation AI cloud base.
Volozh once summarized the conditions for starting a new cloud as three legs: technology and talent, capital, and market sales. Nebius, spun out of the Dutch-listed parent of Yandex, inherited engineers who had long built large-scale machine learning infrastructure and a data center in Finland. The company had over $2 billion in cash and no debt on its balance sheet, but Volozh said it would continue raising capital. “You need good technology and people who understand it, and you also need a lot of capital to build AI infrastructure,” he said. Early customers were mainly concentrated in the U.S., with large tech companies and hyperscalers on one end and AI-native startups on the other, later expanding to broader enterprise markets such as finance, pharma, and robotics.
As AI customers have become more differentiated, new cloud products have begun to layer. For hyperscale companies like Microsoft and Meta, purchasing bare-metal capacity is sufficient because they have their own software stacks; those that truly rely on cloud vendors’ tools and platforms are AI-native startups and enterprise customers. Volozh explicitly said: “We do not treat these large bare-metal contracts as our main business, but as one of the financing methods to build our own cloud and serve other customers in the market.” According to Leiphone, Microsoft has signed a bare-metal contract worth over $17 billion, while Meta’s arrangement includes approximately $12 billion in bare-metal procurement and $15 billion in take-or-pay commitments, which are used to raise bank financing for new data centers and GPU capacity.
Competition in AI cloud now has two dimensions: scale and product depth. Downward, players need to control land, power, and data centers; upward, they need to provide IaaS, inference, and token services. Volozh emphasized vertical integration starting from land because outsourced systems are difficult to optimize as a whole, and every outsourced layer costs a margin. Early on, to launch quickly, Nebius used third-party colocation data centers, but as scale grows, it is building large sites of hundreds of megawatts to gigawatts.
The way compute is metered and who uses it are also changing. At the bare-metal layer, customers buy GPU hours; at the inference layer, compute is sold as tokens. Volozh sees a new generation of customers using the cloud in an “agentic” way, caring less about where compute is sent and not needing to stitch together a hundred tools, because agents optimize efficiency. For providers, making tokens faster and cheaper requires power, data centers, racks, and software to be tightly integrated.
Volozh does not view AI cloud as a red ocean where traditional vendors and new clouds fight, but a blue ocean opened by new users, new demands, and new services. “Facing these new demands and users, traditional hyperscale cloud vendors do not have the protective moat they had in their existing businesses. We are not trying to take away their existing businesses, but to build new ones together with them,” he said. It is not a winner-take-all consumer market, but a B2B market. He expects compute to remain a constraint for the next year, limiting how much infrastructure humanity can build.